Volume Scaling: AI's First Killer App With Three Dimensions
AI

Volume Scaling: AI's First Killer App With Three Dimensions

Your AI pilot was not wrong. It cost more than it saved and the output needed checking twice. But it measured the one axis of agentic AI that does not compound, which is why AI agents look marginal right up until they stop being marginal.

11 min read
audio-thumbnail
AI pilots aren't measuring the volume of change
0:00
/31.76

Three things about AI agents are improving at the same time.

  1. What you let them touch
  2. The cost to achieve a task
  3. The world of possibilities

Three doublings don't add up to six, they multiply to eight. I don't believe most people are budgeting for that type of improvement scaling.


Volume Scaling, defined

Double the dimensions of a cube and the surface area grows by four and the volume grows by eight. Galileo published that relationship in 1638, and it is the most useful thing anyone has written about agentic AI.

Every computing era before this one scaled like surface area. Two dimensions, more or less, with humans pushing both. AI agents scale like volume. Three dimensions, and a bonus that systems do most of the work.

I call this Volume Scaling, and it is the killer app for AI. Not any single agent. Not any single task. The compounding itself is the product.

A killer app is not about being the best software on a platform. It is the core reason people buy into a platform. The term was coined in 1987 and applied backward to VisiCalc, which earned it the killer app title: buyers paid a hundred dollars for the spreadsheet, then two thousand more for the Apple II so they could run VisiCalc. People saw the software and the hardware was just overhead.

Ask a business user today what they would buy AI agents for and the answer they believe is time saved. That answer is why the pilot failed. Time saved is a measurement taken on a single day, and a single day is but one interval and there are no compounding effects there.

Axis One: the permissions ladder

Scope is earned, and once earned it does not drift back.

An agent's journey should begin with something small and reversible. It drafts, it does not send. It proposes, it does not commit. After you, the operator, have seen it perform many repetitions, you widen it's scope: it sends, it commits, it runs the whole sequence without asking. Each rung is granted on evidence from the rung below.

This looks like how you promote a person but it isn't

A person carries their judgment across contexts, which is why promoting them is risky in a way agents are not. Humans will hit a situation nobody described and resolve it using values they absorbed rather than values you wrote down. Sometimes brilliantly, sometimes not.

In contrast, an agent working against a defined objective and a defined measure does not improvise its way into a new reading of the goal. It does what you specified. Probabilistic effort with deterministic outcome constraint is what makes the ladder safe to climb.

So the ladder only goes up.

Every rung an agent holds is permanent capacity, and permanent capacity is what compounds.

The operator who granted ten scopes last year did not spend the year saving time. He spent it building a floor he now starts from.

Axis Two: skill enablement

The same job gets cheaper and better while you do nothing.

This is the axis with the hardest evidence behind it. METR measures what they call the task-completion time horizon: the length of job, measured by how long a skilled human takes, that an agent finishes on its own with fifty percent reliability. That horizon has been doubling for six years. Their updated task suite puts the doubling time since 2024 at eighty-nine days.

The potential objection is pretty evident so here it is. This is software work, and labs optimize for software work. So METR ran the same method across nine benchmarks covering scientific reasoning, math, robotics, computer use, and self-driving. The rates diverge wildly. Competition math doubling rate is about three months, self-driving is much longer at about twenty.

METR is careful to say the data outside software are noisier and any single benchmark should be read loosely.

Read that as a partial defeat if you like. It isn't one, because of what did not show up:

Not one domain came back sub-exponential, and not one came back flat. The speed is domain-specific. The shape is not.

You didn't have to anything to receive that value. You didn't retrain anything. You didn't rewrite a workflow. It is as simple as the job you delegated 12+ months ago is better today and it will be done better again next quarter.

Axis Three: capability overhang

The first two axes describe an agent getting more scope and getting better at it. The third is a different kind of thing. The possibility of use is expanding in ways we don't even understand yet. Imagine the example of a non-obvious emergent behavior like language translation. Now, you can speak to any customer in their native tongue. An agent scoped to one narrow job inherits capabilities nobody scoped it for, because the model beneath improved.

The TAM for Intelligence is Infinity
Every major technology wave of the last 50 years gave humans a better tool. Intelligence is different in kind, not just degree — because for the first time, the tool begins to direct itself.

Which means a delegation made in the past can be a major compounding contributor to the future in unexpected ways. The gap between what these systems can do and what anyone has asked them to do is the largest unclaimed resource in business right now, and it grows faster than the requests do.

What three axes look like inside a real AI stack

I have run this on my own operations long enough to watch it happen.

It started with custom GPTs. Those grew into single skills doing narrow jobs. Then integrated into external systems with guardrails on both sides. Finally, a refined tuple of prompt, context and good/bad output examples ensure alignment with my goals. Those earned their way into orchestrators that own whole pipelines and dispatch sub-skills underneath them, which is how I've chosen to implement the permissions ladder.

The cost (time, resources, money or combination of all three) of new work fell. Adding a whole distribution channel now takes one sub-skill file and one row in a registry, where it used to take a rebuild. And model upgrades keep making existing skills do more than I ever specified, without my editing a line of them. It needs regular updating which I believe will become a new business discipline in the near future.

I bet on the compounding instead of the output, which is a different wager than most operators are making.

A concrete example

I assembled my newsletter by hand dozens of times before I automated any part of it. That was not caution. You cannot decompose a job you have not done, and the decomposition is the entire product.

What you are hunting for is the small problems buried inside the big one that resolve to a clean yes or no. Does this post belong in this issue. Is this link internal or external. Are appropriate tracking tags added. Has this subject line run before. Those answers are deterministic, and a machine should own every one of them.

Around them sits the part that is not: what is worth saying this week. The skill exists to hold those two apart, so the machine takes every yes-or-no and I keep the judgment.

Bundle enough of those small answers together and you have a skill. An orchestrator then calls them in the right order, and the assembly stops being something I do.

That didn't just buy my time back. It increased my focus and allowed my attention to move to the next repeatable problem instead turning the crank of the last one.

The piece you are reading went through that pipeline. Interview, research, draft gates, editorial gates, formatting, attribution, push, repurpose. I approved decisions at the gates and wrote none of the repeat behind-the-scenes work which gives me more time on creating the actual content.

Which is also why your pilot probably showed marginal utility.

A single skill, measured alone may not look useful, because alone it might not be. The return does not live in the skill itself. I've seen the highest value accretion at the orchestration step, when the deterministic pieces get sequenced and the whole thing runs end to end without you in it. Most pilots stop one move short of that, then report accurately that they found nothing. I have argued the adoption version of this before: one project produces a local outcome, and a local outcome is the wrong unit for measuring a capability you are trying to build.

Stop Searching for the Signal. Start Four AI Projects.
AI adoption is the biggest challenge for most businesses. Instead of chasing one use case, run four AI projects across a simple risk and implementation framework to build capability, reduce downside, and accelerate the real problem, adoption.

The control group

My position is that AI is volume scaling while most-to-all previous technologies were only surface area scaling.

Here is the objection that deserves the most respect: prior eras were multi-dimensional too. The internet compounded across reach and cost and speed simultaneously. So did mobile. So did cloud. Naming three axes and drawing a cube might be a cheap writer's trick.

So I did an audit.

  • ENIAC was built to compute artillery firing tables, and UNIVAC I went to the Census Bureau. One axis only: faster math. The machine got quicker at the thing it was built for and did nothing else until a human rewired it.
  • The personal computer added a second. VisiCalc made the machine useful to someone who could not program, and the software ecosystem meant capability arrived without new hardware. Two axes, both pushed by people writing code.
  • Email and then web browsers are the same type of compounding story. Reach, cost, and speed all improved together, and manual labor laying cable and shipping browsers did the work.
  • Cloud collapsed the capital requirement and made capacity elastic, which was an incredible platform for software builders. Two genuine dimensions of growth. Engineers still provisioned, architected, and decided what ran.
  • Mobile is the strongest case for previous volumetric scale. The App Store's move was not one killer app but access to an ecosystem, so every user found their own. Distribution, discovery, and capability all compounded, and the ecosystem grew without Apple building any of it.

Every one of those is real but every one of them also has the same signature: humans supplied the improvement on every axis, which caps the rate at how fast humans work. Developers ship more apps, engineers lay more cable. The growth can be steep but it is still linear.

Agentic AI breaks that signature on two of three axes.

  1. Skill improves without anyone touching the deployment.
  2. Domain expands without anyone requesting it.

Only the permission ladders still require a human, and granting technology scope is fundamentally a human decision, not a build. That is the whole difference, and it is why this curve needs a log axis to display and the others never did.

The contrarian position

The strongest objection is not that agents are expensive. It is that permissions, skill, and domain are the same variable counted three times.

Permissions widen because reliability improved, and reliability improved because the models improved. Domain expanded for the same reason. Strip the framing away and what is left may be one variable, model capability, with three downstream effects I have counted separately to manufacture a cube. If that is right, eight collapses toward two.

Why it is incorrect

The first is that the axes have different clocks and different owners. Permissions move on an operator's decisions and organizational trust, not on release schedules. I have watched a lab ship a capability I did not grant scope for, and watched myself grant scope on a capability that had been available for six months. If they were one variable they would move together. They do not.

The second is that they interact rather than merely coexist. A capability gain becomes a permission grant which becomes a wider surface for the next capability gain to land on. That is a multiplication benefit, not addition, and it is precisely what volume scaling describes.

Now some hard evidence that goes against volume scaling. METR ran a randomized controlled trial in early 2025 and found experienced open-source developers took nineteen percent longer with AI tools while believing they had been twenty percent faster. Not a survey but a controlled trial, from the same researchers I have been citing in this article.

It is the best available proof that operators cannot trust their own read on this, in either direction.

Roughly a year later METR went back to it. For the returning developers they now estimate an eighteen percent speedup rather than a slowdown. Rather than declaring victory METR are rebuilding the study, because too many developers refused to join an experiment that required them to work without AI.

So the contamination in the data is that the subjects would not give the tools up. Even so, this is not a clean win for me, and METR says plainly that the opt-outs bias the estimate. It is the second axis, showing up in the wild, measured by the people best positioned to report the opposite.

The timing cost of a missing a rung

Twelve months late is not twelve months of work to make up.

The person who started granting scope last year is not ahead on hours. They are standing on a rung already, and the rungs below him were paid for with reliability evidence that only piles up from here. You cannot buy that. Rung eight was granted on what rung seven proved, which is the whole point of a ladder and the reason there is no shortcut up one.

No shortcut is the cost. What you get for doing basic process audit and automation work into the world of AI is the part I would not have believed could be done in just a few years. You can now wake up to work that finished overnight, at a standard you did not have to check, on a job you defined upfront.

That is the whole prize, and it is available to anyone willing to start granting scope this quarter.

None of which requires the one-person billion-dollar company to arrive on anyone's schedule. Sam Altman described a betting pool among tech CEOs for the year it happens, and the claims that it has already happened do not survive much scrutiny. The direction, not the bet, is the useful part, because a one-person company worth a billion dollars is not a simple productivity story. It is how someone uses compounding arithmetic in more dimensions than one person can push by hand.

That is the shape of the thing worth watching. Not whether one founder gets to a billion, but that proof and capability have both stopped requiring capital and headcount to acquire.

Building a Company Is Nearly Free. And It's Coming for Venture Capital.
Proof of demand no longer requires capital. That breaks the oldest assumption in startup logic: that raising money is the next move after an idea. Here is what replaces it.

Which leaves a question I do not have a clean answer to. If permissions are the only axis still requiring a human, and every rung granted is permanent, then the operators pulling away are not the ones with better agents. They are the ones who have granted more scope, earlier, and have the evidence to keep granting it.

Is the scarce resource in agentic AI simply the willingness to hand something over and watch what happens?
Licensed under CC BY 4.0 .