machine minds

Your AI Pipeline Has Too Many Geniuses!

September 29, 202616 min read

Stop paying for prose when all you need is a yes or no.

Somewhere in a large company right now, a model that could write a passable sonnet about actuarial tables is being asked whether a form field is empty. It will answer correctly, after a pause, at a price, and with a tidy paragraph of reasoning that nobody will ever read. Repeat that a few million times a month and you have one of the quietest cost centers in enterprise technology.

Article content

The habit is easy to understand. We took a system built to generate language, discovered it could do almost anything we asked, and installed it at every junction of our workflows. Then we started asking it questions whose honest answers are approve, reject, retry, route, escalate, or stop. It works, which is exactly why the habit deserves a hard look.

A workable architecture can harden into an expensive default long before anyone asks whether every step needed generation at all.

The most expensive traffic light in the building

Picture a city that found one brilliant traffic engineer and decided to station her at every intersection. She reads the flow of cars, anticipates the bus, and could write a thoughtful memo about each decision if you asked. She is also paid by the word, takes a second or two to think at every light, and has been posted to twelve thousand intersections where the only question is red or green. Nobody would design a city this way on purpose, yet a remarkable number of AI pipelines look exactly like it once you stop admiring the demo and trace a real request through the system.

Article content

The demo was a single, glorious exchange in which a document went in, an impressive answer came out, and the room nodded. Production looks nothing like that. A model reads a claim, a contract, a transaction, a support ticket, or a maintenance report, and then something has to decide whether the extraction is good enough. Something else picks the path the case should follow, another component checks the result, uncertainty may trigger a second attempt or a stronger model, and eventually something decides whether a person needs to get involved.

Those small decisions are scattered throughout the pipeline rather than parked at the end. Let’s call them decision gates. A decision gate is a bounded point in an orchestration where the system chooses among known actions instead of composing an answer. It needs good judgment, but it has no use for eloquence. Most agentic architecture diagrams, underneath the arrows and the optimism, are long chains of decision gates with a little generation sprinkled at the edges.

Judgment required, eloquence optional

The useful thing about the phrase is that it describes a role rather than a technology. If the policy is explicit and stable, ordinary code or a rules engine may be the best gate you will ever deploy.

If historical patterns predict a bounded label, logistic regression, gradient boosting, embedding similarity, or a small fine-tuned classifier may be plenty.

If the decision still requires broad semantic judgment, a routed or tightly constrained LLM may earn its place, and when the case is truly murky, the right gate may be a person with context, authority, and a slightly alarming inbox.

Article content

Research reached this conclusion years before the current wave of enthusiasm. SetFit, published in 2022, showed that a small, prompt-free classifier trained on a handful of labeled examples could match much heavier few-shot methods on text classification while training an order of magnitude faster. PERFECT, from the same year, removed step-by-step text generation from bounded few-shot tasks and reported nearly 100 times faster training and inference than the prompt-based methods it was compared against. Neither paper was trying to win an argument about generative AI, which makes both more persuasive here.

The newest entrant makes the idea harder to ignore. TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev in early access this month as what it calls a “System One Model.” Jev takes application state plus a schema of typed questions defined in advance and returns decisions with probabilities rather than prose, which the company describes as “smart if-statements.” I will come back to its numbers with appropriate skepticism, because Jev is a vivid example of the pattern and a poor substitute for the argument.

Novelists filling in checkboxes

Insurance makes the pattern unusually easy to see, because insurance is a business built on turning messy evidence into bounded decisions. A multimodal model may be exactly what you want for reading adjuster notes, photographs, PDFs, repair estimates, and the other wonderfully inconsistent artifacts people produce when something goes wrong with a car. That work is hard, and generative models are good at it. Nobody sensible is proposing to replace that step with a regular expression.

Article content

Once those artifacts have become structured state, however, the claim runs into a series of much narrower gates. One gate checks whether the file is complete enough to continue, and another selects the coverage path. A third decides whether a fraud signal crosses a threshold, a fourth checks whether the payment falls within an adjuster’s authority, and a fifth decides whether the whole thing needs a specialist. Each has a small number of exits, and each has historically been handled by some mix of rules, scoring models, and experienced people.

Paying a system to reason in prose about a yes-or-no question is a bit like hiring a novelist to fill in checkboxes. The novelist will manage, the checkboxes will be fine, and the invoice will be memorable.

Generation earns its place again near the end, when someone needs a clear explanation for the file or a letter to the customer that does not read as if a compliance department assembled it at 4 a.m. Between intake and resolution, though, most gates want a classification, a score, or a threshold comparison.

Every industry is running the same loop

The same shape shows up almost everywhere once you start looking. A bank may use an LLM to interpret customer correspondence and a scoring model to route the transaction. A security operations team may use generation to explain an unfamiliar alert while much smaller mechanisms triage the millions of routine events that arrive before lunch, and a manufacturer may use an LLM to digest years of service history while predictive models decide which machines deserve attention this week.

Article content

Customer service, logistics, compliance, and healthcare administration all contain the same transitions. The loop is usually some version of understand, decide, act, verify, and decide again. Different verbs in that loop justify different kinds of intelligence, because understanding often needs a large model, deciding often does not, and verifying frequently needs something cheaper and more predictable than the thing it is checking.

Enterprise architects absorbed this lesson long ago for everything except AI. Nobody insists that every component of a modern system share one database, one programming language, or one storage model, and anyone who proposed it would be dropped gently from an architecture review meeting. With AI, we are drifting toward exactly the one-model-for-everything architecture. The reason is rarely conviction, because a single general-purpose model is simply the easiest thing to wire up against a deadline.

The economics are not subtle

The research already offers useful numbers, provided nobody blends them into one magical “AI is 73 times cheaper” statistic for a board slide. Each figure comes from a specific setup, and each deserves to be read that way. DistilBERT, published in 2019, produced a model 40% smaller and 60% faster than BERT that kept 97% of its measured language-understanding performance, which shows that fitting model capacity to the task was sound economics long before anyone paid per token.

Article content

Routing research tells the same story from inside the LLM world, where every destination is still generative. RouteLLM, from the LMSYS team at Berkeley, trained routers that send easy queries to a cheaper model and hard ones to a stronger one. It reported cost reductions of over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K compared with sending everything to GPT-4, while still reaching 95% of GPT-4’s performance. FrugalGPT from Stanford took the cascade approach and matched the best individual model with up to 98% cost reduction in its experiments.

The spread in those results is the real lesson. RouteLLM’s routers performed at close to random on MMLU (Massive Multitask Language Understanding) until they were given a small amount of in-distribution training data, and FrugalGPT’s headline figure comes from particular experiments rather than a general law. None of these numbers is a forecast for your pipeline, and treating one as such is a fine way to embarrass yourself in a quarterly review. What they establish, rather loudly, is that choosing the computational mechanism for each step is an optimization problem that most organizations are not yet solving.

The multipliers shrink as the evidence gets closer to home

Jev’s headline claims are spectacular, at 193.6 times faster and 444.6 times cheaper than frontier LLMs, with input priced at $0.042 per million tokens and output unmetered. To TypeSafe’s credit, its launch post is unusually candid about where those figures come from. Its own capabilities team built the benchmark workflows, the reference answers are the averaged predictions of two frontier models rather than ground truth, and the company says it expects its numbers to sit at the high end of real-world gains.

Article content

The early outside tests land well below the headline, and the drop is instructive. BrillMark, an agency that builds AI helpdesks, ran 100 synthetic support tickets through Jev and three frontier models and found Jev four to seven times faster and 31 to 65 times cheaper. A Vercel engineer told TechCrunch that swapping an LLM safety classifier for Jev produced results five to 18 times faster, with greater accuracy. Then Nikhil Mudholkar, CTO of Bryo AI, classified 1,565 business emails and found two Gemini models slightly more accurate than Jev, with Jev roughly 10 to 22 times cheaper.

Follow those numbers in order and the story becomes clearer. The multiplier falls from hundreds to dozens to roughly ten as the test moves from the vendor’s lab toward real work, and the last step introduces a small accuracy trade that nobody can see from the headline. A benchmark scored against other models’ opinions measures agreement with those models, which is useful but is not the same as being right about a claim or a patient. The architecture is more interesting than the benchmark, and the benchmark deserves validation on your own workload before it goes anywhere near a spending decision.

The token bill is the smallest number on the whiteboard

The executive arithmetic arrives quickly, and it is where the conventional argument quietly falls apart. At ten million automated decisions a month, saving a tenth of a cent per decision is worth $10,000 a month, and saving a full cent is worth $100,000. Those numbers are real, and at that scale they justify the engineering. Most companies, however, are not yet running ten million decisions a month through their AI pipelines, and the math at ordinary volumes tells a different story.

Article content

Take an insurer handling 10,000 claims a month, with eight gates per claim, and move three-quarters of those gates from a generative model at an assumed 1.2 cents per call to a decision model at a tenth of that. The direct inference saving comes to about $648 a month, which is roughly what the architecture review costs in catered sandwiches. Now suppose the better-calibrated gates safely reduce unnecessary human review from 20% of claims to 16%, at 12 minutes and $50 an hour per review. That saves $4,000 a month, more than six times the token saving, before you count faster cycle times or happier customers.

These inputs are my assumptions rather than research findings, and every company needs its own. The shape of the result, though, holds up under a wide range of inputs. At moderate volume, the model bill is rarely the prize. The larger lever is human capacity, because the gate that matters most is the one deciding which cases a person needs to see, and its value depends almost entirely on whether you can trust its confidence.

Latency is a tax you pay in series

Cost is what leaders argue about in meetings. Latency is what customers actually feel. The two behave differently, and the difference matters. Costs add up across a pipeline, while latency stacks up along its critical path, one gate after another.

Article content

If six serial gates each take half a second longer than they need to, the customer waits an extra three seconds, even though no individual call looks offensive on the architecture diagram. Each team reviews its own component, finds it reasonable, and moves on, which is how a system ends up with every part fast enough and the whole too slow. TypeSafe cites frontier response times ranging from a few seconds to several minutes and claims 70 to 500 milliseconds for Jev, a gap that matters little in a chat window and a great deal inside code.

Naming each decision point as a gate makes the critical path visible. You can assign each gate a latency budget and decide which ones deserve an expensive model and which deserve something that answers in tens of milliseconds. Get the gates right and the numbers follow: what you spend, how fast it answers, how much it handles, and how it feels to customers.

Confidence needs a policy

All of this points to the part of the pattern that matters most, which is what happens when the system is uncertain. In a prompt-driven pipeline, the answer usually lives in a sentence like “only approve if you are confident.” That sentence is a risk policy, but it is written in a form no risk committee can read, test, or sign off on. It sounds sensible right up until someone asks what confident means, and it shifts whenever an engineer tweaks the wording.

Article content

A decision gate can state the policy plainly. Above 0.95, the workflow continues, and between 0.75 and 0.95 it buys a second opinion from a stronger model. Below 0.75, the case goes to a person, and when a hard control condition fails, everything stops. The Bryo test shows why that matters, because Jev’s most confident 85.5% of answers had zero errors on that set, while nearly half of its answers below 70% confidence were wrong. Armin Ronacher, CTO of Earendil, put the trade plainly to TechCrunch, describing a 50% answer as a coin toss to disregard and a 95% answer as something you can act on.

A confidence score is not truth, of course. Guo and colleagues showed back in 2017 that modern neural networks are often poorly calibrated, so a stated 90% may not mean right nine times in ten. Thresholds have to be tested against real outcomes, gate by gate, and tested again when the data drifts. That dull discipline is what separates AI programs that scale from AI programs that demo.

The payoff is governance that regulators already expect. The NAIC’s model bulletin on insurer AI, adopted in 2023 for states to implement, calls for a written AI governance program and describes documentation that insurance departments may request. The NIST AI Risk Management Framework treats governance, measurement, and oversight as work that runs across the whole lifecycle. Bounded outputs do not make an AI system safe by themselves, but they give safety policy somewhere better to live than a sentence in a prompt.

A fast factory feeding a slow door

TypeSafe named Jev after William Stanley Jevons, the Victorian economist who observed that more efficient steam engines increased total coal consumption rather than reducing it. The company is betting that cheaper machine decisions will follow the same curve, and I suspect it is right whether or not Jev itself wins. When a decision costs a tiny fraction of a cent and returns in a fraction of a second, the list of decisions worth automating grows dramatically. The fraud check that ran on a sample starts running on every transaction, and the quality review that happened weekly starts happening on every document.

Article content

This is where the tidy cost story breaks down. The organization rarely banks the saving, because it spends the saving on far more decisions, each of which needs a threshold, an owner, an escalation path, and someone who notices when it drifts. Most enterprises approve a handful of AI use cases per quarter through a committee. That structure was never designed for a world where the unit of automation is a single gate that costs almost nothing to run and very little to add.

The human review desk feels the shift first. If automated gates clear most cases in milliseconds, the remaining cases arrive faster and in greater volume than before, and they are by definition the hard ones. Organizations that celebrate the automation rate without redesigning the review desk will discover they have built a very fast factory feeding a very slow door. Decision gates therefore deserve to be treated as governed assets, each with an inventory entry, a named business owner, a documented threshold, a calibration check, and a record of how often it escalates.

Boring on purpose

None of this means the answer is a different tool for every step. Google’s paper on hidden technical debt in machine learning systems warned a decade ago that ML components pile up maintenance costs through entangled dependencies, glue code, configuration sprawl, and hidden feedback loops. A pipeline with six kinds of gate has more places to accumulate that debt than a pipeline with one kind of model. Adding five model families to save twelve dollars is not architecture. It is indulgent rather than disciplined.

Article content

The sensible adoption path starts with whatever model makes the workflow work. An LLM is often the right gate at first, when the categories are still being discovered and the business cannot yet say precisely what it wants decided. Once a gate stabilizes, its decisions and the human corrections to them become labeled data, which is exactly what a smaller model needs to take over. Volume settles the rest, because a gate that fires ten thousand times a month rarely justifies a model of its own, while one that fires ten million times rarely justifies anything else.

Most teams that walk their workflows this way will find a handful of gates that account for most of the calls, most of the waiting, and most of the human review. Those are the gates worth moving to a smaller, faster, more predictable mechanism. The rest can stay exactly where they are, and that restraint is as much a part of good architecture as the optimization itself.

Circle the verbs

Generative AI taught the industry to treat language as a programmable interface to intelligence, and that opened an extraordinary amount of territory very quickly. As AI systems become operating infrastructure rather than impressive demonstrations, architecture starts to matter again, and specialization, routing, thresholds, escalation, unit economics, and deliberately boring code reclaim their places alongside the models.

There is no reason to assume language is the final abstraction for every machine decision we automate, any more than SQL was the final word on every data problem.

Article content

The mature AI stack will not be the one with the most intelligence crammed into every step. It will be the one that allocates intelligence deliberately, with generation where the problem is open-ended, decision gates where the exits are known, plain code where the policy is exact, and people where uncertainty or consequence demands them. It will also know who owns the threshold at each gate and what happens when the confidence number drops.

So here is a modest exercise for your next architecture review. Put your most important AI workflow on a whiteboard, circle every point where the system chooses among known actions, and write the verb next to each circle. Then ask three unglamorous questions of every gate: (1) what it needs to know, (2) what a mistake costs, and (3) what the cheapest reliable way to decide might be. The next breakthrough in your AI program may be a model that can do more, but it may just as easily be noticing all the places where you have been paying an LLM to do too much.

Article content


Further Reading


Disclaimer: The perspectives shared in this article are our own and do not represent those of our employer or any affiliated organizations. All company names, product names, logos, and brands mentioned are the property of their respective owners and are used for identification and illustrative purposes only. No endorsement, sponsorship, or affiliation is intended or implied. References to specific companies or case studies are based on publicly available information and are used solely for educational and discussion purposes.