- Match model capability and context to the complexity and value of the work.
- Measure the cost of completed outcomes, not raw model or token consumption.
- Give agentic workflows explicit budgets, boundaries and escalation points.
There was a point in the cloud transition when many organisations had technically moved to the cloud without really changing how they operated.
Servers that had lived in a data centre became virtual machines running in somebody else’s. Capacity was still provisioned for peaks. Infrastructure remained permanently available whether it was needed or not. The location and commercial model had changed, but much of the thinking had not.
The bigger shift came when organisations started designing around the characteristics of cloud itself.
Capacity could expand and contract with demand. Infrastructure could become ephemeral. Managed services could replace things businesses no longer needed to own. FinOps added another layer of discipline: measure consumption, understand unit economics and continually match resources to the workload.
Cloud taught us to make compute elastic.
AI may be approaching the same transition.
Cloud taught us to scale compute to the workload. AI will teach us to scale intelligence to the problem.
We started by maximising access
The first phase of enterprise generative AI has understandably been about adoption.
Give people access to capable models. Put copilots into existing tools. Encourage experimentation. Build agents. Remove friction. Measure usage.
Consumption-based AI makes a familiar cloud infrastructure dynamic more pronounced. Like logging, monitoring and telemetry, usage is easy to add and each increment can look inconsequential. But AI workflows can multiply that consumption through context, model calls, retries and reasoning steps — often without anyone explicitly deciding to spend more.
The economics are beginning to become visible.
Gartner describes an “inference paradox”: individual token economics improve while increasingly sophisticated AI workflows consume enough additional inference that the overall cost per agentic workflow rises. Gartner predicts those costs will increase by more than fivefold through 2028. 1
This is not simply a question of models becoming expensive. It is what happens when cheaper access enables us to consume much more intelligence.
The pattern should feel familiar.
The biggest model is the equivalent of the oversized server
Early cloud architectures frequently provisioned infrastructure for the largest conceivable workload.
Enterprise AI often does something similar with intelligence.
A powerful model is selected, connected to an application and becomes the default destination for requests regardless of complexity. But extracting an invoice number, classifying a support request and reasoning through a complex commercial decision do not require the same capability.
Treating them as if they do is effectively over-provisioning intelligence.
The alternative is an architecture in which capability becomes dynamic:
This is already becoming production infrastructure rather than theory. Amazon Bedrock’s Intelligent Prompt Routing selects between models according to predicted response quality and cost. AWS says routing can reduce costs by up to 30% without compromising accuracy for supported model families. 2
McKinsey reaches a similar conclusion from an enterprise economics perspective: organisations should match model capability to the work rather than assuming every task warrants the most capable model available. Its 2026 Enterprise AI FinOps research found that active optimisation is already producing material savings for some organisations. 3
The objective is not to use less AI. It is to stop using expensive intelligence where cheaper intelligence will do.
Model choice is only the beginning
There is a danger of reducing this to a new version of cloud right-sizing: replace the expensive model with a cheaper one and declare victory.
AI economics are more complicated than that.
A model may repeatedly receive thousands of tokens of context it has already seen. An agent may call the same tools several times. Conversation histories grow. Failed reasoning paths create retries. Multiple agents can delegate work to one another. A seemingly inexpensive interaction can therefore become an expensive workflow.
The FinOps Foundation warns about the effect of growing context windows and repeated token consumption when evaluating generative AI unit economics. 4
Caching repeated context, reducing unnecessary prompt content, batching appropriate workloads and designing retrieval carefully can therefore matter as much as headline model price. AWS says prompt caching can reduce input-token costs by up to 90% for supported workloads. 5
So the architectural question changes from “Which AI model should we use?” to “What is the least amount of intelligence and context required to achieve this outcome reliably?”
That is a much more useful question.
Elastic does not mean infinite
There is another lesson from cloud architecture worth remembering.
We never really meant infinite when we talked about elasticity.
A well-designed cloud platform has limits: budgets, quotas, rate limits, scaling policies, timeouts and circuit breakers. If a service comes under pathological demand, automatically provisioning infrastructure forever to satisfy every request would be a peculiar interpretation of resilience. At some point, the correct architectural behaviour is to stop scaling.
AI systems need the same principle.
An agent receiving a poor result from a tool might reason again. Then retry. Add context. Call another model. Delegate to a sub-agent. Retry once more. Each individual decision can look reasonable while the overall run becomes irrational.
Microsoft’s TokenOps work describes this failure mode: an agent can make hundreds of individually unremarkable model calls before the aggregate cost becomes apparent. Its response is a shared, run-scoped token budget capable of stopping a runaway workflow while it is executing. 6
Intelligence should scale with the problem — but it should also have a ceiling.
That ceiling does not have to be purely financial. It might be a token budget, number of reasoning steps, permitted tool calls, latency threshold, confidence score, risk classification or a point at which the system asks a human.
The mature architecture therefore is not simply harder problem → more AI. It is harder problem → proportionately more capability → evaluate → escalate if justified → stop when the threshold is reached.
From adoption to architecture
None of this suggests the AI opportunity is becoming smaller. The opposite is true.
As intelligence becomes cheaper, more capable and easier to embed into software, its use will expand dramatically. Gartner’s inference paradox exists precisely because falling unit costs make more ambitious applications economically possible. 1
The organisations that manage this well may therefore consume substantially more AI than they do today — but they will consume it differently.
They will route work according to complexity. Cache what does not need to be recomputed. Use deterministic software where determinism is valuable. Give agents explicit budgets and boundaries. Escalate intelligently. Measure the cost of completed outcomes rather than celebrating raw token consumption.
Cloud went through a similar maturation. Moving infrastructure off-premises was only the first step. The larger benefit came when architecture, operations and commercial management changed to exploit the characteristics of the new platform.
Enterprise AI access was the first step too.
The first phase of enterprise AI was about access. The next will be about architecture.
And perhaps the most important architectural question will not be how much intelligence can we use?
It will be how much intelligence does this problem actually warrant?
- Gartner, Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028, 17 August 2026. https://www.gartner.com/en/newsroom/press-releases/2026-08-17-gartner-predicts-ai-inference-costs-per-agentic-workflow-will-increase-more-than-fivefold-through-2028
- AWS, Amazon Bedrock Intelligent Prompt Routing. https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-routing.html
- McKinsey & Company, The cost of intelligence: How CIOs can manage AI demand at scale, 2026. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-cost-of-intelligence-how-cios-can-manage-ai-demand-at-scale
- FinOps Foundation, guidance on generative AI unit economics and token-based consumption. https://www.finops.org/
- AWS, Cost optimization for Amazon Bedrock. https://docs.aws.amazon.com/bedrock/latest/userguide/cost-optimization.html
- Microsoft, TokenOps: real-time, run-scoped cost control for AI agents, August 2026. https://commandline.microsoft.com/tokenops-real-time-run-scoped-cost-control-ai-agents/

