machine minds

The Collection of Models Is the Model

September 14, 202618 min read

Why the next AI PC may win by routing local intelligence, not by running the cloud frontier models.

Benchmark tables have a blind spot big enough to hold an entire product category. They tell us which model scored highest, what it costs, how quickly it answers, and whether this quarter’s champion gained three points on last quarter’s champion.

What they rarely ask is more awkward: what would happen if every question went to whichever model happened to be best at that particular question?

In June 2026, Bradley Fowler and ten coauthors put a number on that blind spot in an arXiv preprint. They evaluated 21 language models across 16 benchmarks, then constructed a Capability Frontier: the best result available at each cost level if an idealized router could choose the right model for each query. At matched cost, correcting for single-model evaluation reduced the average error rate by 54 percent relative to each benchmark’s top-performing model.

Article content

There is an important catch. The idealized router does not exist. It chooses with hindsight, which is a lovely feature if you can convince engineering to ship it. When it could also select among multiple generations, the paper reports an 82 percent error reduction and state-of-the-art accuracy at 85 percent lower cost. The work is a preprint in arXiv, not a deployed router, but one simulation matters here. As query-topic entropy rises, the gap between idealized routing and the best single model almost continuously. A knowledge worker’s day is unusually heterogeneous, which makes that result relevant even though it is not yet a product.

That should make us reconsider the question the AI PC market keeps asking: when will my computer be able to run the frontier model? A more useful standard is useful equivalence. If the machine can produce an outcome that is good enough to be indistinguishable for the task at hand, the user may not care whether a single local model ever matched the frontier architecture that produced the cloud answer. Those two questions sound similar, but they lead to very different architectures, economics, and businesses.

I made a related argument in Build AI Like You Expect the Ground to Move: the durable architecture is not local versus cloud, but one that can decide where intelligence should run without being rebuilt around each environment.

The PC Does Not Need to Catch the Frontier.

For most of the generative AI era, the personal computer has been a very well-dressed terminal, or, if we want the less flattering systems term, an AI client. You type. Your prompt leaves the building. An enormous remote system thinks on your behalf. The answer comes back, and your laptop takes credit for having good wifi.

Article content

On-device AI has started to change that arrangement, but we still tend to describe local models as smaller versions of cloud models. The implicit race is vertical: more parameters, more memory, more accelerator capacity, until one day the PC catches whatever counts as frontier intelligence at that moment. There are two problems with the race. The frontier keeps moving, and users do not actually need architectural equivalence. They need useful equivalence.

If a workstation can produce the same accepted code review, research synthesis, migration plan, contract analysis, or financial recommendation that a frontier cloud system would have produced, most people will not care that it arrived through a different path. They may care about ten extra seconds, privacy, reliability, cost, or whether the machine works on an airplane, but they will not care that no single local model could have produced the answer by itself. That creates room for a different kind of AI system.

Jagged is not necessarily broken.

AI capability is not smooth. Fabrizio Dell’Acqua and colleagues gave the shape a useful name: the jagged technological frontier. Their preregistered experiment, published in Organization Science in 2026, involved 758 Boston Consulting Group consultants. Across 18 realistic tasks deliberately placed inside GPT-4’s capability frontier, AI users completed more work, faster, at higher quality. On one deliberately selected task outside the frontier, the AI-assisted groups were 19 percentage points less likely to reach the correct answer. The authors themselves note the asymmetry: eighteen inside-frontier tasks and one outside-frontier task. (https://doi.org/10.1287/orsc.2025.21838)

Article content

That study is about humans using AI, not model routing, and stretching it into proof of a multi-model architecture would be citation gymnastics. Its contribution here is simpler: capability is lumpy. Tasks that look equally difficult to us are not equally difficult to a model, and anyone who works with these systems already knows this in a less scientific way. The model that elegantly restructures a codebase can give a shallow answer about distributed failure. Another writes a thoughtful strategic memo and then becomes strangely confident while doing arithmetic.

A programming specialist does not need to be a gifted historian, a mathematical reasoner does not need graceful prose, and a domain model tuned for insurance regulation does not need frontier-level expertise in cinematography.

We usually treat that unevenness as something the next training run should smooth away. But jaggedness can also be a design surface. If their strengths are complementary, the weakness of one model does not have to become the weakness of the system. In that sense, the collection of models begins to act like the model. That makes verification more important, not less; as I argued in I Think, Therefore I Might Be True, trust has to become a property of the workflow rather than something inferred from the fluency of any single answer.

One question can require several kinds of intelligence, but never all kinds.

Consider a familiar executive problem: should we replace a twenty-year-old insurance platform, modernize it incrementally, or build an intermediary layer and postpone the harder decision? It looks like one question because it fits in one sentence. It actually contains software architecture, security, data migration, operating cost, regulatory exposure, delivery risk, organizational capacity, vendor dependency, and the uncomfortable fact that every transformation sponsor believes their program will be the exception to the historical record.

Article content

As an architectural illustration, a composite system could decompose that request. A software specialist could examine architecture, a security specialist could look for threat and compliance issues, a finance-oriented specialist or tool-assisted workflow could frame the economics, a modernization specialist could evaluate migration patterns, and a critic could attack the first round of recommendations before a general model synthesized the result. None of that is a claim that today’s products reliably do this. In fact, the handoffs and verification steps in exactly this kind of workflow are where current multi-agent systems often fail, which is the tension the architecture has to solve rather than wave away.

I will call the broader pattern composite inference: independently deployable models cooperating at inference time to produce one outcome. I am deliberately avoiding the ‘federated’ label because readers understandably associate it with federated learning, which concerns decentralized training, and this is also different from mixture-of-experts routing inside one neural network. Here the models can have separate weights, vendors, lifecycles, and jobs. The user never needs six chat windows. From the outside, the system should behave like one intelligence.

Somebody already built the shelf.

No major vendor has announced the complete architecture I am describing. That distinction matters. There is enough speculative AI writing without volunteering to make it worse. What several vendors have done is build much of the infrastructure needed to make this possible, and those pieces are beginning to converge.

Microsoft’s Foundry Local is the clearest example of the shelf itself. It provides an on-device runtime, a curated model catalog, automatic hardware acceleration, local caching, lifecycle management, and OpenAI-compatible interfaces. Microsoft says prompts and outputs can be processed entirely on-device, no Azure subscription is required, and local inference carries no per-token cost. Strip away the branding and this begins to look like a package manager for cognitive capabilities. (https://learn.microsoft.com/en-us/azure/foundry-local/what-is-foundry-local)

Article content

Apple is approaching the same problem from the platform side. Its 2026 Foundation Models framework uses a common LanguageModel protocol across the on-device model, Private Cloud Compute, Core AI, MLX, and community-provided models. Dynamic Profiles can change the active model, tools, instructions, and reasoning configuration while preserving history. Apple also introduced an Evaluations framework for measuring intelligent features and catching regressions. Model abstraction and model evaluation arriving together is a stronger signal than either feature alone.

Qualcomm’s AI Hub advertises 300-plus optimized models, while GenieX provides an on-device runtime spanning NPU, GPU, and CPU with OpenAI-compatible APIs on Windows. NVIDIA and Microsoft are attacking the hardware and security side: RTX Spark systems are announced with up to 128 GB of unified memory and vendor-stated support for running 120-billion-parameter models locally, while OpenShell adds policy controls for agents on a user’s primary PC. Initial systems are due from six major PC makers, with more to follow.

These companies are not following one coordinated roadmap, which is exactly why the convergence is worth noticing. Model catalogs, abstraction layers, hardware-aware runtimes, orchestration tooling, evaluation, security primitives, and larger unified-memory machines are appearing independently. The plumbing is arriving before the product category has a settled name.

The missing layer is measured judgment.

A shelf full of models is not intelligence any more than a shelf full of cookbooks is a restaurant. The missing layer is judgment. A router needs to know what each specialist is good at, where it degrades, how much memory it needs, which tools and data it may access, and how trustworthy its confidence tends to be.

Call that a capability scorecard if a name is useful, but the important word is measured. A model’s declared competence is marketing metadata. Demonstrated competence on your actual workloads is routing data. That description has to be continuously earned through evaluation because models change, workloads drift, and yesterday’s strong specialist can become today’s subtle regression after an update.

Article content

This makes the evaluation harness a strategic asset. Whoever measures the specialists gets to route them, and whoever routes them shapes what the overall system can do. Apple’s decision to ship Evaluations beside its model-abstraction layer is therefore notable. The immediate action for technology leaders is not to buy a shelf of models. It is to build a representative task suite for high-value workflows, with acceptance criteria, red-team cases, human review, and enough telemetry to measure rework as well as first-pass quality.

The uncomfortable gap is that no general-purpose deployable router has demonstrated anything close to the idealized router in the Capability Frontier paper. Production routing faces distribution shift, privacy policies, cold-start costs, model updates, permissions, and changing workloads. Someone also has to own model-admission criteria, evaluation policy, exception handling, and the auditability of routing decisions, making the router as much a governance system as an engineering one.

That governance problem has a close cousin in Invisible Teammates, where I looked at how seemingly technical choices about prompts, tools, permissions, and autonomy quietly determine who or what gets authority inside a system.

Model-to-model communication deserves the same pragmatic treatment. Permissions, authentication, provenance, resource limits, state handles, tool invocation, and downstream actions belong in a structured control plane. The reasoning artifact can remain richer: a specialist can explain what it chose, where it is uncertain, and what it wants another model to challenge. The useful pattern is a hybrid: structure where determinism is required, language-rich artifacts where judgment and disagreement have to survive the hand-off.

I explored the same continuity problem from the memory side in AI’s Goldfish Problem: Reboot with CMI: once intelligence is spread across steps, preserving useful context stops being a convenience and becomes part of the architecture.

Memory is a scheduling problem: Load the Right Model at the Right Time.

The obvious objection is hardware memory. If a machine has a generalist, a coding model, a mathematics specialist, a vision model, a critic, and a synthesizer, they do not all need to sit in VRAM or unified memory at once. Some problems benefit from parallel work, while others are naturally serial. Frequently used models can remain resident. A rarer specialist can load for one stage, produce an artifact, and leave so the next model can take its place. Foundry Local already treats acquisition, caching, loading, inference, and lifecycle management as normal parts of local execution.

Article content

The trade is latency, and storage bandwidth becomes part of the inference architecture. Swapping multi-gigabyte models is only useful if the storage path can feed memory fast enough, and model state has to be preserved without quietly flattening the reasoning that matters. Still, latency is contextual. Hundreds of milliseconds matter in live voice and checkout interactions. They matter much less when an architect is reviewing a migration plan, a lawyer is examining a contract, or an engineer is debugging a problem that has already consumed two days of human effort.

For serious knowledge work, time to acceptable outcome may matter more than time to first token. The elapsed time until the user can accept the result without another round of repair. A slower local system that uses two specialists, a critic, and a verification pass can be more valuable than a fast first draft that creates forty minutes of cleanup. The machine does not need enough memory to contain every intelligence it owns simultaneously. It needs enough memory, storage bandwidth, and scheduling discipline to assemble the right intelligence when needed.

The idealized router is not a product.

This is where the attractive architecture has to survive contact with reality. The Capability Frontier demonstrates theoretical headroom from perfect selection. It does not show how to build the selector. A real router must recognize task type before knowing whether its classification was right, preserve information across hand-offs, verify outputs without repeating the same mistake, and adapt when workloads move underneath it.

Article content

The strongest counter-evidence comes from Mert Cemri and colleagues at NeurIPS 2025. Their MAST-Data contains more than 1,600 annotated traces across seven popular multi-agent frameworks, and their taxonomy identifies 14 failure modes grouped around system design, inter-agent misalignment, and task verification. Their opening observation is uncomfortable for anyone selling multi-agent systems as an automatic upgrade since gains over single-agent systems are often minimal. The systems repeat work, lose state, misalign roles, and fail to verify or terminate correctly.

The public version of the same lesson appeared during GPT-5’s 2025 launch. OpenAI introduced GPT-5 as a “unified” system with a real-time router between a fast model and deeper reasoning. On launch day, OpenAI acknowledged that the auto-switcher was out of commission for part of the day and that the system therefore appeared much less capable. That happened inside one vendor’s tightly controlled cloud stack, which makes the point narrower, not weaker. Routing quality can determine the apparent quality of an otherwise excellent collection.

So two well-supported facts pull in opposite directions. Idealized router selection across heterogeneous models shows enormous headroom, while deployed multi-model and multi-agent systems routinely leave much of that headroom on the floor. The distance between those two facts is not a caveat. It is the product.

There is an even more important unknown for the local thesis. Fowler’s study used 21 competitive models, most of them large, and its entropy result is evidence at that scale. It has not established that a workstation full of smaller, quantized specialists will fail in complementary ways. Their failures may be correlated because of shared training data, shared teachers, distillation practices, or, as a testable hypothesis, aggressive quantization. If they fail on the same questions, the composite collapses toward its best member. That experiment could strengthen this thesis or kill it cleanly, and both outcomes would be useful.

The economics reward second opinions.

Gartner recently gave the economic side of this problem a wonderfully paradoxical name: the Inference Paradox. Its August 2026 forecast says inference cost per agentic workflow will rise more than fivefold through 2028 even as model economics improve, because richer workflows reason, negotiate, route, call tools, and take more steps. Gartner’s conclusion is explicitly multi-model: competitive products will need inference tiering, routing, and orchestration rather than one economical model for every job.

Gartner also supplies the strongest counterargument to a simplistic local-cost story. In a separate 2026 forecast, it expects provider inference on a one-trillion-parameter model to cost more than 90 percent less in 2030 than in 2025. If cloud intelligence becomes dramatically cheaper, local AI cannot rest its case on “tokens are expensive”. Fortunately, the more interesting local argument is not primarily about token price.

Article content

The techniques that move a real system closer to the idealized router are redundant by nature: ask a second model, generate several candidates, run a critic, verify the answer, retry the weak sub-problem, and sometimes throw work away. In a metered environment, every discarded generation is still on the bill. On hardware you already own, those internal attempts have a different marginal cost structure. Local compute is not free. Electricity, hardware depreciation, support, security, and engineering all have an excellent habit of returning to budget meetings, but owned hardware can afford to be intellectually wasteful in ways metered services may have to ration.

That leaves a second set of advantages that do not disappear when cloud inference gets cheaper. Sensitive prompts and outputs can remain local by design, systems can keep working through connectivity failures, organizations can pin model versions and audit behavior, and teams can iterate without watching a token meter. None of those guarantees privacy by itself. Telemetry settings, retrieval sources, tools, downloads, and OS policy still matter. For regulated enterprises and governments, however, those controls can be architecture requirements rather than lifestyle features.

One catalog has an advantage. It also has power.

There is a reason the first compelling versions of this architecture may come from a single vendor. A unified catalog can standardize packaging, memory behavior, capability evaluation, versioning, permissions, safety policies, hand-off conventions, and update discipline. When a specialist changes, the same platform can rerun evaluations and adjust the routing policy rather than leaving customers to discover the regression during a quarterly close.

Article content

That vertical integration is a substantial engineering advantage, but it is also a governance decision disguised as convenience. The catalog can become the gatekeeper that decides which models are installable, which capabilities are trusted, which providers can participate, and what evidence qualifies a model for a particular task. An open ecosystem offers more choice but pushes far more integration and evaluation burden onto the operator. A curated ecosystem reduces that burden while concentrating authority.

The organizational consequence is easy to miss. Someone has to own the evaluation policy, model-admission criteria, audit trail, escalation rules, and the process for appealing a bad route. The router does not merely allocate compute. It allocates cognitive authority. That is a surprisingly consequential job for something that still looks, on most architecture diagrams, like middleware.

The cloud does not empty out.

None of this predicts abandoned data centers. Frontier training, enormous shared contexts, scientific workloads, global services, organization-scale coordination, proprietary data services, and anything requiring extreme parallelism remain natural cloud problems. Cheaper intelligence will also create demand that does not exist today, which is why local growth and cloud growth are not mutually exclusive.

Article content

What can change is the composition of demand. Here is simple arithmetic, not a forecast: if total AI consumption grows tenfold while the cloud’s share of inference falls to 30 percent, absolute cloud workload is still three times today’s level. Falling share and rising volume can coexist perfectly well. The strategic exposure is not empty capacity. It is the commoditization of routine intelligence and the migration of more marginal inference economics onto devices the cloud provider does not operate.

The local machine then becomes the default place where intelligence is assembled, while the cloud becomes an escalation path called when its incremental value justifies the cost, latency, privacy boundary, or capability gap. Today the PC mostly asks the cloud to think. A mature composite system would decide when the cloud is worth asking, which is a reversal of control rather than the disappearance of infrastructure.

The benchmark nobody is watching.

The Capability Frontier started this article by exposing a blind spot in single-model benchmarks. The deeper blind spot is that there is still no widely adopted public benchmark for the complete system: routing accuracy, hand-off fidelity, context retention, critique effectiveness, synthesis reliability, privacy boundaries, energy, latency, and accepted final outcome. We keep awarding trophies to individual models while the potentially valuable unit is becoming the composition.

Article content

Organizations do not need to wait for the industry to invent that benchmark. A representative task suite is the private version: measure accepted outcome rather than model score, track rework and time to acceptable outcome, preserve disagreement, and record whether a critic or second model actually changed the result. Whoever publishes a credible system-of-models benchmark will change the public scoreboard.

Whoever builds a good internal one first gains a practical routing advantage today.

This is the same reason I keep returning to accepted outcomes rather than impressive demos: in When the Vendor Only Gets Paid for Outcomes, the Demo Changes!, the measurement system becomes valuable precisely because it forces everyone to agree on what “worked” before celebrating the technology.

The milestone worth watching is therefore not the day a personal computer runs whatever model sits at the top of the leader board, because that target moves every quarter. It is the day a reasonably priced machine can assemble enough complementary intelligence, and enough measured judgment about how to use it, that the user stops caring where the answer came from. At that point the PC stops behaving mainly as an AI client and becomes an intelligence platform.

The title of this article is deliberately a little ahead of the evidence. The collection of models may become the model, but only if something can tell those models apart, route work intelligently, preserve what matters between them, and know when the cloud is actually better. If that layer becomes trustworthy, the most valuable component in the next AI PC may not be the smartest model on the machine. It may be the part that knows which one to ask.

Article content


Further Reading


Disclaimer: The perspectives shared in this article are our own and do not represent those of our employer or any affiliated organizations. All company names, product names, logos, and brands mentioned are the property of their respective owners and are used for identification and illustrative purposes only. No endorsement, sponsorship, or affiliation is intended or implied. References to specific companies or case studies are based on publicly available information and are used solely for educational and discussion purposes.