I Hired Four AI Agents. Then I Became Their Middle Manager.
What 4.8 billion processed tokens taught me about Claude, Codex, and the cost of running a company with no shared memory.


One human manager carrying context between four isolated AI agents.
Hey, it’s Daniil.
Imagine hiring the smartest person you know.
Now erase their memory every morning, move their desk to another building, hide half the company files, and get annoyed when they ask what the product does.
Congratulations. You have built my AI stack.
Claude Code. Codex. Cowork. Hermes. Together, they have processed roughly 4.8 billion tokens for me.
That number looked impressive for about eleven seconds.
Then I realized it might also be a receipt for explaining the same company four times.
If you're not subscribed yet — this is the kind of thing we write about every week.
Subscribe now

My real Claude Code statistics: 782.9 million processed tokens, most of them cached context.
I cannot tell you what percentage of those tokens produced useful work and what percentage paid for agents to rediscover facts another agent already knew.
We broke down how Hermes and OpenClaw actually work in practice: I rebuilt my OpenClaw setup on Hermes + GPT-5.5
But I can tell you the pattern I kept seeing:
- Claude learns something important inside one project.
- Codex starts the next task without that lesson.
- Hermes knows the operational history but not the publication rules.
- I become the human API between all of them.
Very advanced technology. Very traditional middle management.
At first, I treated this as an annoying personal workflow problem.
Then I looked at it as a business owner.
The model subscriptions are the cheapest part. The expensive part is paying smart people to reconstruct company context, review contradictions, and fix work that was perfectly reasonable against the wrong version of reality.
My problem was not that I needed a smarter agent.
I had created a company where every employee was brilliant, fast, available 24/7—and permanently on their first day.
What “no context” actually costs
When an agent lacks context, it usually does not stop.
That would be convenient.
It fills the gap with the most plausible version of your company. The result can be polished, internally consistent, and completely wrong.
In practice, I see four kinds of waste.
1. You pay for the same discovery again
One agent maps the project, finds the relevant files, learns why a previous approach failed, and finally does the job.
A week later, another agent repeats the entire investigation.
This is not research anymore. It is an expensive onboarding ritual.
2. Two correct agents produce one wrong company
Claude can make a sensible product decision based on the product files it sees.
Codex can make a sensible review comment based on the diff it sees.
If neither sees the decision that connects the two, they can both be correct inside their own little universe—and still send the project backward.
Humans do this too. We call it a meeting.
3. Local corrections never become company knowledge
You tell Claude, “Never do X again.”
Claude apologizes. Claude fixes X. Everyone goes home happy.
Tomorrow, Codex does X with enormous confidence because nobody told Codex about yesterday.
The correction lived in a chat. The company learned nothing.
4. The human becomes the context router
This one is sneaky.
The agents look productive because they return work quickly. Meanwhile, you spend your day copying the brief into Claude, explaining the decision to Codex, forwarding the result to Hermes, and resolving the contradictions.
You did not remove coordination work.
You hired yourself as the dispatcher.
Share this with the person who keeps opening new AI chats and re-explaining the company from scratch.
Share
For a business, this is a very dumb way to buy intelligence
For one person, fragmented context is annoying.
For a company, it is financially stupid.
Here is deliberately simple, illustrative math:
10 people × 30 minutes per day rebuilding context = 5 hours per day
5 hours × 5 days = 25 hours per week
25 hours × $80 fully loaded cost = $2,000 per week
≈ $104,000 per year
That is before the Claude and Codex bills. More importantly, it is before the cost of a bad launch, a duplicated experiment, a wrong customer answer, or a senior person discovering that two teams implemented opposite decisions.
And 30 minutes is conservative. Context gets rebuilt in prompts, kickoff calls, Slack threads, reviews, corrections, and the little “quick clarification” messages that eat an afternoon one bite at a time.
The stupid part is not paying for AI seats.
The stupid part is buying intelligence and then making every instance start from zero.
Every new agent seat should make the company memory more valuable. Without shared context, it can do the opposite: more parallel work, more conflicting assumptions, and more humans paid to reconcile the result.
A giant context window is not company memory
The obvious solution is to give the agent everything.
Every document. Every chat. Every decision. Every Slack message from 2023, including the one where somebody wrote “let’s revisit this next quarter.”
Please don’t.
A large context window is a bigger desk. It is not an organized company.
If you dump old, private, contradictory, and irrelevant information into it, the agent now has a very large desk covered in garbage.
I think about context in three layers:

Company truth, project truth, and task state flow into one task packet shared by Claude and Codex.
The layers should not be treated equally.
Company truth changes slowly. Project truth changes whenever a decision is made. Task state is temporary and should usually die with the task.
Most bad agent setups mix all three in one immortal chat and hope the model works out which sentence still matters.
What shared projects actually look like inside Claude Team
If your company already pays for Claude Team or Enterprise, I would start here before buying another “AI memory platform.”
A shared Claude Project gives the team one reusable starting point:
- Project instructions — the role, rules, tone, and definition of good work.
- Project knowledge — uploaded documents, text, code, and other reference files.
- Visibility — private to invited members or available across the organization.
- Permissions — people can either use the project or edit its instructions and knowledge.

Claude’s real Team and Enterprise project creation screen with organization-wide and private visibility options.
This is already much better than ten employees maintaining ten slightly different “master prompts.” Everyone can begin with the same product brief, customer research, brand rules, and definition of done.

Claude’s real project-sharing panel with invited members, organization-wide access, and view or edit permissions.
But there is one very important limitation.
Claude shares the project, not everyone’s brain.
According to Anthropic’s current project-sharing documentation, chats inside a shared project remain private unless the user manually shares them.
So if Anna discovers why the Q4 campaign failed inside her chat, Boris does not automatically inherit that lesson tomorrow. Somebody still has to turn the lesson into a decision, update the project knowledge, and mark the previous answer obsolete.
That is shared context. It is not yet shared memory.
A sensible Claude setup is layered
I would not create one organization-wide project called Company Brain and upload everything into it.
That is how the pricing strategy, HR policies, old pitch deck, and customer-support macros all end up competing for the same answer—and possibly the same permissions.
A practical company setup could look like this:
- Organization instructions: universal rules Claude should follow everywhere—company identity, security constraints, language, citation requirements, and actions that always need approval. Claude now supports organization instructions for Team and Enterprise plans.
- Marketing shared project: brand voice, audience, approved claims, campaign history, performance insights, and visual rules.
- Customer-support shared project: current product documentation, refund policy, escalation rules, response templates, and known incidents.
- Product shared project: roadmap, customer research, decision log, current specifications, and rejected alternatives.
- Engineering shared project: architecture decisions, repository instructions, runbooks, testing rules, and incident lessons.
- Private chats: temporary thinking, drafts, and task-specific conversations that do not need to become company knowledge.
Each shared project needs an owner. Their job is not to write every document. Their job is to remove obsolete context, resolve contradictions, and decide which lessons deserve to become shared knowledge.
Claude now has a second layer: Ask Your Org
For Team and Enterprise accounts, Claude also provides a pre-configured enterprise-search project called Ask Your Org.
An owner chooses the company tools that can be connected. Individual employees authenticate with services such as Slack, Google, and Microsoft 365. Claude can then search across the sources that particular employee already has permission to see and return answers with citations.
This solves a different problem:
- Shared Projects answer: “Which context should this team reuse for this kind of work?”
- Ask Your Org answers: “Where does the company already have information about this question?”
Both are useful. Neither automatically turns yesterday’s correction into tomorrow’s company policy.
Search is a pull mechanism. Memory needs a write-back loop.
How to structure your agent stack so it doesn't collapse under its own weight: Your AI Agent Stack Is Spaghetti — It Should Be Lasagna
Shared memory is becoming its own software category
This space is confusing because vendors use enterprise search, company knowledge, context layer, and memory as if they mean the same thing.
They do not.
Here is how I would separate the useful options.
Onyx: open-source enterprise context
Onyx connects to systems such as Slack, Google Drive, GitHub, Confluence, Jira, SharePoint, Salesforce, and Zendesk, then keeps an internal search index updated.

A real Onyx connector configuration showing indexing scope and document-access controls.
The important enterprise feature is not the number of logos. It is permission syncing. On supported connectors, Onyx can mirror access-control lists from the source, so an employee should not receive a document they could not open in the original system. Onyx documents this as an Enterprise Edition feature.
I would look at Onyx when a company wants one searchable context layer across many tools, especially when self-hosting and infrastructure control matter.
Glean: enterprise search plus an organization graph
Glean builds a knowledge graph across content, people, and activity. It combines documents and messages with organizational relationships and existing permissions.
This is useful in a larger company where “who owns this?” can be as important as “where is the document?”
Dust Pods: a multiplayer workspace for humans and agents
Dust Pods are closer to collaborative project rooms. Conversations, files, tasks, humans, and agents live together around one goal. Agents can read the Pod’s shared context and write new files back into it.
Dust also distinguishes between live-synced Connections—for Notion, Google Drive, Confluence, and other sources—and a Pod’s own files, which can be written by humans or agents and become searchable immediately.
This is much closer to actual organizational memory because agents are not limited to retrieval. They can leave useful artifacts behind for the next person or agent.
Notion Enterprise Search: sensible if Notion already runs the company
Notion Enterprise Search searches the workspace plus connected apps such as Slack, Google Drive, Microsoft Teams, and Jira, and links answers back to their sources.
If the company already maintains decisions, plans, and processes in Notion, this can be the path of least resistance. But it will not rescue a workspace full of untitled pages, contradictory policies, and abandoned databases. AI search is not a substitute for ownership.
Zep: memory infrastructure for agents you are building
Zep belongs in a different category. It stores conversation history and business data in temporal knowledge graphs, then retrieves relevant context for custom agents. Its MCP server can let Claude, ChatGPT, Cursor, and internal agents work against the same governed memory.
I would treat Zep as developer infrastructure, not an employee-facing company search product. It is useful when you are building agents that must remember users, events, relationships, and changing facts across sessions.
What I would actually choose
- A small team already using Claude: Shared Projects, Ask Your Org, and one explicitly maintained decision log.
- A company with knowledge fragmented across many SaaS tools: Onyx or Glean.
- A team that wants humans and agents working in the same project space: Dust Pods.
- A Notion-first company: Notion Enterprise Search before adding another knowledge layer.
- Custom agents that must remember over time: Zep or a domain-specific memory layer such as ShopClaw’s Store Brain.
Before buying any of them, I would ask seven boring questions:
- Which systems are indexed, and how quickly do updates appear?
- Are source permissions enforced per user?
- Does every answer cite the original evidence?
- Can an agent write a verified result back, or only retrieve?
- How are old facts expired or marked as superseded?
- Who approves changes to company truth?
- Can we audit which context produced an action?
If a vendor has a beautiful answer to search and no answer to write-back, it is a knowledge browser—not company memory.
The small context system I actually use
For Creators AI, the useful context is not one magical brain.md file.
It is a few boring files with clear jobs:
STYLE_GUIDE.md
CONTENT_INSIGHTS.md
strategy/CURRENT_STATE_RESEARCH_AUDIT.md
strategy/CREATORS_AI_GROWTH_RETENTION_PLAN_2026.md
drafts/
This is not only my homemade convention. Claude Code’s own memory documentation says every session starts with a fresh context window, then uses CLAUDE.md and auto memory to carry selected knowledge across sessions.

Claude Code’s official documentation showing CLAUDE.md and auto memory as its two cross-session memory mechanisms.
Codex uses a different door into the same idea: AGENTS.md. The official open-source Codex repository has a real one at its root, full of concrete instructions about tools, tests, architecture, and what the agent must not touch.

The public OpenAI Codex repository with its real AGENTS.md project instructions.
The style guide tells an agent how we sound. The performance file tells it what readers historically respond to. Strategy files explain what matters now. The current draft contains the task state.
The important trick is that I do not load all of this into every job.
I start with a context-discovery request:
Read the current task and the project instructions first.
Before doing any work, tell me:
1. which company facts you need;
2. which project decisions you need;
3. which files are the source of truth for each;
4. what is still missing or contradictory.
Do not edit yet.
Do not read unrelated folders “just in case.”
This does two useful things.
First, it forces the agent to show me its map before it starts driving.
Second, it exposes facts that exist only in my head. If the agent asks, “What language should the publication use?” the fix is not to answer that question forever. The fix is to put the answer in the publication context.
The task packet: onboarding in 60 seconds
Every serious agent task now gets a small packet.
Not a 40-page prompt. Not the complete autobiography of the project. Just the contract the agent needs to do this job.
GOAL
What must be true when this task is finished?
WHY NOW
What changed, or what problem are we fixing?
SOURCE OF TRUTH
Which files, pages, or data are authoritative?
DECISIONS ALREADY MADE
What should not be reopened, and why?
CONSTRAINTS
What must remain private, unchanged, reversible, or in my voice?
PROOF
Which checks or visible states demonstrate completion?
STOP POINTS
Which actions require my approval?
Here is what that looks like for this article:
GOAL
Turn the post into a practical argument about why Claude and Codex
become inefficient without shared company context.
SOURCE OF TRUTH
STYLE_GUIDE.md, CONTENT_INSIGHTS.md, the current draft,
and the real usage screenshots already approved for publication.
DECISIONS ALREADY MADE
Do not write a token-counting methodology.
Do not frame 4.8B as unusually high or as a leaderboard.
CONSTRAINTS
English only. No private tabs, customer data, messages, account details,
or invented interface screenshots.
PROOF
The Substack-style preview loads, every image resolves,
and each screenshot proves the paragraph around it.
That packet is far more valuable than “use your best model.”
ShopClaw is my shared context framework for eCommerce
eCommerce is almost designed to create context failures.
The truth is fragmented across Shopify, GA4, Search Console, lifecycle marketing, the theme repository, the product catalog, campaign briefs, support messages, and whatever somebody remembered to put in Slack.
Connect an agent to one of those surfaces and it can become locally smart and globally stupid.
That is why I built ShopClaw as a shared context framework for eCommerce, not another chat window with a Shopify integration.
Every store gets a git-backed Store Brain containing its facts, tasks, decisions, experiments, playbooks, and measured outcomes. Specialists in strategy, growth, frontend, design, and QA do not inherit one enormous immortal conversation. They receive structured handoffs from the same store context, then write the result back for the next run.

The real ShopClaw Memory screen: accumulated knowledge across decisions, audits, research, playbooks, measurements, and reports—not one agent’s private chat history.
One real example: on August 5, ShopClaw found an active 15% discount with zero uses after four days.
The discount code worked. The storefront did not explain the offer, and the campaign button did not apply it.
Without shared context, the analytics agent reports zero conversions. The developer confirms that the code is valid. Marketing assumes the promotion is live. Three locally correct answers, one broken campaign.
ShopClaw connected the campaign state, storefront behavior, analytics signal, prior store decisions, and approval rules. It traced the failure, prepared a reversible theme fix, and captured desktop and mobile evidence for QA.
That is the point of shared context. It does not make one model magically smarter. It lets several specialists act like they work for the same company.
Claude builds. Codex reviews the contract, not Claude’s life story
Shared context does not mean identical context.
When Claude is building inside a project, it needs implementation history, relevant files, annoying edge cases, and the checks it must run.
When Codex reviews the result, I do not want it to inherit Claude’s entire conversation. That makes the reviewer too sympathetic to the route the builder chose.
But Codex still needs the company truth and the task contract.
My review handoff looks like this:
Review this change independently.
Shared context:
- the task packet;
- current company and project rules;
- decisions explicitly marked as closed.
Evidence:
- the diff;
- the required checks and their output;
- the final rendered state if the work is visual.
Do not repeat the builder’s reasoning.
Do not reopen closed decisions without new evidence.
Return findings by severity, with file, line, impact, and proof.

A real Claude Code session changing files and running typecheck and lint.
Visible tests are useful. But tests only prove the implementation against the checks you wrote.
They cannot prove that the agent understood the company.
That is why the reviewer needs both: technical evidence and the correct contract.
The write-back is where the company gets smarter
The task is not finished when the output ships.
It is finished when the next agent can avoid paying for the same lesson.

A five-step write-back loop that turns completed work into company memory for the next task.
At the end of important work, I ask:
Before closing this task, list:
1. new facts we verified;
2. decisions we made and why;
3. approaches that failed;
4. rules or checks that should prevent repetition;
5. context that is now obsolete.
For each item, propose the exact source-of-truth file to update.
Do not write back raw chat history.
The “raw chat history” line matters.
Company memory should be the useful residue of work, not a landfill of every sentence the company has ever produced.
I want decisions with dates and reasons. I want rules with scope. I want links to evidence. I want obsolete instructions removed.
One correction in one chat is customer support.
One correction written back into the system is organizational learning.
What I measure now instead of tokens
Token totals are fun in the same way step counters are fun. They prove movement. They do not prove you went anywhere useful.
The better questions are:
- How often did I have to explain a known fact again?
- How many decisions did a second agent accidentally reopen?
- How much work was duplicated before anyone noticed?
- How many review cycles came from missing context rather than poor execution?
- Did the task leave the next agent with a better starting point?
If I add three agents and spend more time routing information between them, I did not build an AI-native company.
I built a group chat.
The one-hour version you can copy
You do not need a vector database, a “second brain,” or a knowledge-management transformation to start.
Pick one project where you repeatedly use Claude or Codex.
Create four small files:
NOW.md — current state and priorities
DECISIONS.md — decisions, dates, reasons, rejected alternatives
RULES.md — stable constraints and non-negotiables
WORKFLOW.md — task packet, proof, approval gates, write-back step
Then do one real task:
- Ask the agent which context it needs.
- Give it only the relevant sources.
- Make the task contract explicit.
- Review the result against that contract.
- Write one useful lesson back into the files.
Run the same workflow with a second agent next week.
If it starts smarter than the first one did, your company has begun to remember.
That is the actual advantage.
How to give an agent the right context without dumping the whole company into its prompt: How Context Engineering Can 10x Your AI Model's Performance
Want to build the system behind the prompts?
If you are using Claude, Codex, or a growing team of agents and want practical workflows that survive the next model release, subscribe to Creators AI.
Continue with these field guides:
- How Context Engineering Can 10x Your AI Model’s Performance — how to give an agent the right information without dumping the whole company into its prompt.
- Skills, Plugins, Swarm Mode: Practical Tips for Claude — how to move reusable rules and workflows out of individual chats.
- Your AI Agent Stack Is Spaghetti—It Should Be Lasagna — a practical three-layer architecture for keeping tools and agents replaceable.
Models will change. Claude will improve. Codex will improve. A new tool will appear and make both of them look ancient for three weeks.
But the context—the decisions, mistakes, taste, constraints, and reasons—belongs to you.
One agent with the right company memory is more useful than five brilliant strangers.
Where does your company remember why it does things the way it does?
Tell me in the comments. And share this with the person who keeps opening new AI chats and re-explaining the company from scratch.
This article was first published in the Creators AI newsletter. View the original edition.

