TOKENOMICS: You don't have a token problem. You have a staffing problem.

2026-06-17

The CFO pulls up an AI invoice, turns the laptop around to show her team, and makes the face I used to make when I’d open a power bill after a month of my kids leaving every light on. They’d signed off on buying the AI tools, and they’d sat through the company off-site and heard the vague directive to the team to “start using AI more.” But they had no idea how quickly the bill would skyrocket. And how unpredictable it would be. They’re not alone. A KPMG survey this year found that only about a quarter of companies say they have a clear view of what their AI actually costs. Half have some idea. The rest find out when the bill arrives. KPMG says it has clients who burned through a full year of token budget in a few months, and at least one whose usage jumped sixfold over the same time period. One analyst put it about as plainly as you can: a lot of finance chiefs are going to “see their Anthropic bill and freak out this quarter.”

On one call, with the CIO of a large midatlantic university, it became very clear. Normally it’s a dance to get closer to their problems where we can help. I didn’t even get a chance to even open with my introduction when he blurted out - “This isn’t going to increase my Token Usage - is it? You have no idea how this is my number one problem.” It couldn’t have been any clearer the pain he was under.

In our experience, part of the problem is that most companies are running their single smartest, most expensive model on every task, top to bottom. That is the equivalent of putting your highest-paid person on every job in the building, including the ones a new hire could do in half the time. It would be easy to think the answer is to a high AI bill is to find a cheaper AI tool. Actually though, we’re seeing that the companies pulling ahead on cost are sticking with the tools they like, but doing a better job organizing their teams around the tools.

Who is doing the work

I made this exact mistake with people first

When we started The Archer Group in 2003, we hired the best people we could afford in each discipline. A great designer, a sharp strategist, a developer who knew how to ship. For a while that worked, because there were only a handful of us and everyone touched everything. Then we grew. And somewhere around twenty or thirty people, the math stopped working.

Early Archer Days

I remember watching a senior creative director, someone we paid real money for judgment and taste, spend an afternoon resizing banner ads. The work got done. It just got done by the most expensive hands in the company, on a task a junior could have handled without breaking a sweat. Multiply that by a few dozen people and a few hundred jobs a month, and you are bleeding money.

At Archer Group, we ended up doing what every agency that survives eventually does. We tiered the work. Production went to juniors. The craft went to the people in the middle who were good and getting better. The senior people focused on the things that carried risk: the client call where the relationship was on the line or the new idea the whole quarter was riding on. We stopped paying senior rates for junior work.

Companies should use the same concept when creating AI agents.

The AI Agent Org Chart

The part that cost me real money

The Communication Tax

A group of researchers at Concordia published a study this year that took apart a multi-agent coding system token by token, just to see where the money actually went. You would assume the cost is in writing the code. It is not. Writing the initial code was one of the cheapest things the system did. The expensive part, by a wide margin, was the review loop: the agents passing the work back and forth to check and refine it. That one loop ate around sixty percent of the tokens. The agents werere-reading each other's context, over and over, every time the work changed hands. The researchers called it ‘the communication tax.’

I suspect many of you are reading this and laughing, because we’ve all seen it happen.. The expensive part should be the work itself, but instead, what really costs you are the meetings about the work, and the wasted effort that comes along with it. Like the senior person getting re-briefed for the third time because the handoff was sloppy. Routing a task to a cheaper resource saves you nothing if you then make four expensive people sit around and talk about it.

A pile of agents chattering at each other can cost you more than the one expensive model you were trying to replace.

So in addition to “use cheap models for cheap tasks,” we’re advising clients that they also need to build a strategy for the handoff between agents.Every time work passes between agents, the whole context gets passed with it, and you pay for it a second time. If you split a job into ten agents without thinking about how they talk to each other, you can end up spending more than you would have running one expensive model straight through.

Same idea, Different workers

What this changes

You and your competitor can buy the exact same models tomorrow morning, from the same three vendors, at the same prices. But the actual asset is the routing logic underneath, who does what, and the discipline of the handoff, how cleanly the work moves between them. That part is unique to your company

So before you approve another budget or buy another tool, do the unglamorous thing. Map your AI work the way you would draw an org chart. List the tasks. Mark which ones need the senior partner and which ones you have been overpaying for out of habit. Then find the handoffs, the places where work changes hands and the context gets passed again, and treat those as the real line item they are.

Most companies are staring at the token bill trying to decide which tool to cut. We’re helping clients understand what the bill is actually telling you:that you put a senior partner on every job in the building, and then sat them all in the same meeting to trade notes.

common business tasks