The model is not the moat.

12 August 2026

Jannis Kearney Bott Jannis Kearney Bott Co-founder and CEO
A hand holding a clear lightbulb against a pink and teal sky

Large language models are exchangeable commodities, and any one can be swapped for another in an afternoon. Which model you run is the least defensible decision in your AI strategy. The value, the intellectual property, and the moat sit in the layer above it, and in the data inside that layer.

Everyone is talking about AI, so loudly that you can no longer hear it

Every keynote, every board pack, every vendor pitch. The word has stopped carrying information.

And underneath the noise, the announcements have become strangely identical. We have rolled out Claude. We have rolled out Gemini. We have rolled out ChatGPT. Delivered with the pride of a moon landing. My honest reaction: so what?

I sometimes struggle to understand the excitement. You have given your people a capable assistant. It summarises content, drafts documents, automates some routine skills. Useful, genuinely. But there is nothing novel in it. Stanford’s AI Index puts organisational AI adoption at 78 per cent, up from 55 per cent a year earlier. Rolling out a large language model does not put you ahead of the field. It puts you in the field. The excitement phase is over. What is left is a commodity, and commodities do not differentiate anyone.

Which raises the question that actually matters: if the model is not the innovation, where does the value sit?

The least defensible layer

Here is where I have landed after building AI into client businesses, and into my own: the large language model is the least defensible part of your AI strategy. That is not a criticism of the models. It is a description of the market.

Every frontier model can be swapped for another in an afternoon. The prompt that ran on one provider yesterday runs on a competitor today, often at a lower price and a comparable result. Stanford’s AI Index found the performance gap between open-weight and closed models collapsed from eight per cent to 1.7 per cent on key benchmarks in a single year. The thing enterprises are paying the most attention to, which model to pick, is the one decision that matters least, because it is the one decision you can reverse at almost no cost.

In my view the model is exactly what the name says: a large language model, exchangeable for the next one. Exchangeable things do not hold value. The value sits in what you build on top of them, and in the data that lives inside what you build. So does the intellectual property.

The economics of raw model use do not stack

I watch this play out every week. Hand a team a frontier model and nothing else, and two things happen. First, they spend an enormous amount of time prompting. Every task starts from a blank page: re-explaining context, re-supplying data, re-establishing constraints the organisation already knows but the model does not. Second, they burn tokens at a rate nobody budgeted for.

The unit economics are extraordinary. Stanford’s AI Index recorded a 280-fold drop in inference cost for GPT-3.5-class performance between November 2022 and October 2024, from roughly US$20 per million tokens to seven cents. Intelligence per token has never been cheaper.

And yet the bills are climbing. Agentic workloads consume tokens exponentially rather than incrementally: an agent reasoning through a task generates orders of magnitude more tokens than a person typing questions into a chat window. Cheaper tokens, multiplied by exponential consumption, still produce a number that makes a chief financial officer wince.

Here is the uncomfortable arithmetic: raw model use pairs the highest human effort with the highest token burn and the least durable output. The model delivers value. I am not in the camp that says it delivers none. But value delivered and cost-benefit achieved are different tests, and raw consumption fails the second one.

Where the value actually sits

MIT’s research on AI in business, reported in August 2025, put a number on what most executives already suspected: 95 per cent of enterprise generative AI pilots deliver no measurable impact on the profit and loss. The revealing part is not the failure rate. It is the cause. The researchers found the problem was not model quality. Generic tools stall because they cannot retain feedback, adapt to context, or improve over time. The five per cent that succeed embed AI deep in real workflows, with memory, with context, with learning loops.

That matches everything I see on the ground: the divide between AI that works and AI that does not is drawn at the application layer, not the model layer.

An application built on top of a model does three things a raw model cannot.

It carries the context, so nobody re-prompts the same organisational knowledge a thousand times. The prompting cost is paid once, engineered in, and amortised across every user and every run.

It constrains the work, so tokens are spent on the task rather than on the model wandering. Well-built applications are, among other things, token-efficiency machines.

It compounds. Every interaction enriches the data inside the application: the decisions, the corrections, the outcomes, the edge cases. That data is what makes the agents inside the application capable, and it is data your competitors do not have and a model vendor cannot replicate.

That last point is the one that matters most. The model knows what the internet knows. Your application knows what your business knows. One of those is available to everyone for cents per million tokens. The other is yours alone.

The exchangeable core, and what rising spend does to it

So treat the model as what I believe it is: interchangeable infrastructure underneath the layer that actually differentiates you. That framing has a practical consequence as spend climbs.

When total token consumption is modest, the convenience of a hosted frontier model wins easily. But agentic workloads change the curve. As consumption compounds, the case for hosting your own open-weight model strengthens: the capability gap has narrowed to low single digits, the workload is predictable enough to provision for, and the data stays inside your perimeter.

None of that is possible if your value is welded to one vendor’s model. All of it is possible if your value lives in the application layer, because an application built properly swaps its engine without the passengers noticing.

What a good harness looks like

Harness is the word I keep coming back to. The model is the engine. The harness is everything around it that turns raw capability into an outcome you would put your name to. A good one has four parts.

The ontologyA live, structured map of your business: the entities that matter (customers, contracts, orders, assets, obligations), how they relate to each other, and the rules that govern them. An agent wired into an ontology does not reason over a soup of documents. It reasons over your actual business objects, which makes its outputs grounded, checkable, and safe enough to automate. This is the difference between an assistant that summarises and an agent that acts. It also has a prerequisite, and my colleague Carlos has written about it properly: you cannot build an ontology on data nobody owns and nobody trusts.
The plumbingIntegrations into the systems of record where the work actually happens, so the agent reads live truth and writes results back, instead of operating in a sidecar nobody reconciles.
The guardrailsPermissions, audit trails, and escalation paths, so every automated decision can be defended to a board, a regulator, or a customer.
The peopleA harness is a change programme wearing an architecture diagram. Workflows change, roles change, and the definition of a day’s work changes, which is a shift our chief technology officer Shibu has written about from the delivery side.

One honest caveat. The model layer is not standing still. The vendors are building harness capabilities into the models themselves (memory, tool use, agentic orchestration) and moving up the stack, closer to the customer. Some of what you would have built by hand a year ago, you now inherit out of the box. That does not rescue raw consumption as a strategy. It sharpens the build decision: inherit what is becoming commodity, and concentrate your own investment on the parts no vendor can ship for you. Your ontology. Your data. The depth of your workflows.

How enterprises should adopt AI

If I were sitting on your executive team, these are the three decisions I would push for. They are the same three we hold ourselves to inside my own firm, because I do not believe in selling a transformation we have not run on ourselves first.

First, stop making the model the strategy. Model selection is an engineering decision, revisited quarterly, reversible by design. Architect every application so the model underneath can be replaced without rebuilding what sits above it.

Second, put the investment where the value compounds: applications that deliver a defined business outcome, wired into real workflows, with the ontology, memory, and context that the successful five per cent got right. And run them as change programmes with executive ownership, because the workflow and the operating model are what actually change. If a proposed initiative cannot name the outcome it delivers and the workflow it lives in, it is a pilot waiting to join the 95 per cent.

Third, treat the data inside those applications as the asset it is. It is the raw material that makes your agents more capable than anyone else’s, and it is the intellectual property that survives every model generation. Guard it accordingly, and be deliberate about which vendors it flows through.

The model wars will keep producing better and cheaper engines. Good. Let the vendors fight that battle with their capital. The enterprises that win will be the ones that spent the same years building the applications and the data that no engine swap can touch.

The question I now ask of every AI line item, in my own budget and in my clients’, is not which model it runs on. It is: what does this own that a competitor with the same model does not?

Jannis Kearney Bott

Written by

Jannis Kearney Bott

Co-founder and CEO

Jannis helps mid-market and enterprise leaders make confident decisions on data, AI, and business applications, and explains complex ideas simply. He built J4RVIS around client success, leads with clarity, and starts every engagement from the outcome.

More from Jannis Kearney Bott

Sources

  1. Stanford HAI, The 2025 AI Index Report, April 2025. Organisational adoption (78 per cent, up from 55 per cent) and the 280-fold inference cost decline: economy chapter. The open-weight to closed performance gap (eight per cent to 1.7 per cent): technical performance chapter. https://hai.stanford.edu/ai-index/2025-ai-index-report
  2. Fortune, "MIT report: 95% of generative AI pilots at companies are failing", 18 August 2025, on MIT NANDA, The GenAI Divide: State of AI in Business 2025. https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/

Want to talk it through?

Bring us the AI line item you cannot defend. Let the vendors fight the model war with their capital: we help you spend yours on the part no vendor can ship for you.

Tell us the decision in front of you.

Not a form that routes into a queue. A principal reads it, and you get a call and a written point of view whether or not there is a project in it for us.

Looking for a role? See open roles at J4RVIS