Applied AI
A position on how organisations should build with AI, and who should own what.
Since the first software ran on a stored-program computer in the late 1940s, building it was slow and costly. Because building was the constraint, the scarce and valuable activity was deciding what to build. Strategy, analysis and advice commanded a premium precisely because execution was expensive and uncertain. The consulting industry is, in large part, a monument to that constraint.
However, that constraint has shifted, significantly. A single practitioner, working with frontier models, can now stand up a working system in days that would once have taken a team a quarter. We have done this repeatedly, across defence, property and capital markets. The length of task an AI can complete without human help has roughly doubled every seven months for six years,1 and the benchmarks that measured the last decade of progress now saturate faster than new ones can be written.2
The consequence is stark. When the cost of building collapses, the premium once paid to decide what to build collapses with it. The scarce inputs are no longer analysis and headcount, which AI now supplies in abundance; they are judgement about what is worth building, and ownership of the result once it exists. An organisation that rents both will find it has bought motion without control.
When building is cheap, the value moves to judgement and ownership.
This paper sets out what we believe follows. That AI is an accelerant rather than a strategy: it widens the gap between what a customer will pay and what it costs to serve them, but it does not invent that gap. That the economics of the firms most organisations rely on are, by construction, poorly suited to it. And that durable advantage lies not in renting intelligence but in owning the stack it runs on.
An organisation that rents both will find it has bought motion without control.
The frontier is moving faster than any organisation can absorb
The pace is not anecdotal. Over six years, the length of task an AI agent can complete unaided has grown exponentially, doubling roughly every seven months.1 The established benchmarks that tracked progress, from MMLU to GSM8K, have saturated; researchers now build harder ones specifically because models exhaust the old ones.2 Driving all of it is sheer volume of release: capable models now ship almost every week, and no team can keep pace. Below is a fraction of what arrived in the first half of 2026 alone:
Frontier and notable open-weight models released between January and June 2026, listed by month. A partial view: hundreds more quantised, distilled and fine-tuned open models shipped alongside them. Source: public launch announcements.
No organisation can evaluate all of this, let alone adopt it. But three shifts sit beneath the noise, and each changes who can build and at what cost: capability is getting radically cheaper, radically smaller, and increasingly open. We quantify them below. The direction is unambiguous, and it is the premise of everything that follows: the frontier is a fast-moving, increasingly commoditised input, not a moat you can rent forever. The advantage does not lie in having access to it, because soon everyone will. It lies in what you build on top, and in whether you own it.
It lies in what you build on top, and in whether you own it.
Cheaper, smaller, and increasingly open
Three forces are pulling the same way. Capability that once cost a fortune, filled a data centre, and lived behind a single provider’s API is becoming cheap, small, and open, each on the trajectory below.
A 280-fold fall in about two years, from $20.00 to $0.07 per million tokens. The line tracks the cheapest price to reach GPT-3.5-class quality, month by month, as each new tier of models undercut the last.3 Independent analysis puts the broader trend near fifty times cheaper per year.4 Log scale.
Source: Covelent analysis of Stanford HAI and Epoch AI data.
The same capability, 142× smaller, in two years: PaLM at 540 billion parameters in 2022, Microsoft's Phi-3-mini at 3.8 billion by 2024.5 Capability that required a data centre now runs on commodity hardware. Log scale.
On coding, scores have soared and the gap has closed at the same time. The best open-weight model has climbed from the low teens to the low-70s on SWE-bench Verified in two years, against a closed frontier now in the high-80s, and the lead has narrowed from roughly thirty points to the mid-teens.6 Open models such as GLM-5.1 and DeepSeek V4 already top individual coding and reasoning benchmarks.7 Percentage of issues resolved.
Source: Covelent analysis; trajectory illustrative, on SWE-bench data.
Investors have started to price the shift
The clearest independent read on where value is moving is the market itself. Through the first half of 2026, investors sold down the enterprise-software and IT-services providers whose revenue AI most directly threatens, the firms whose business is selling seats, licences and billable delivery of exactly the work that is now automatable. These are not fringe names. They are the incumbents on a typical enterprise supplier list, the CRM, the ERP suite, the systems integrator, the ratings and data provider, so for most organisations the names below read as a fair approximation of their own vendor roster, and the multi-year contracts most hold are with exactly the firms under the most pressure.
Year-to-date share price change, 1 January to 1 July 2026. Source: Yahoo Finance. Names shown are widely held enterprise software, IT-services and data providers, a fair approximation of a typical enterprise vendor roster.
We do not read a single half-year as destiny, and share prices are noisy. But the pattern is consistent with the thesis of this paper: the market is repricing the model of renting people and software to do work that can increasingly be built once and owned.
The advice was a product of slow building
The model was never only about the advice.
The economics of a professional-services firm are straightforward. Profit comes from leverage, a wide base of junior staff billed out beneath a thin layer of seniors, and from utilisation, keeping those juniors on billable engagements for as long as possible.8 Compensation tracks the hours billed.9 For decades it has worked.
The defining programmes of the last two decades, digital transformation, cloud migration, ERP rollouts, were sold as three-to-five-year, multi-million-pound engagements. Their real product was rarely a single recommendation; it was a standing team of external consultants embedded inside the client. And with that team went the knowledge the work created, product knowledge, domain knowledge and infrastructure knowledge alike. When the engagement ends, the understanding of how the enterprise actually runs leaves with the supplier. The advice was outsourced, but so was the knowledge, and knowledge is the harder dependency to unwind.
That is the design of the model, not a flaw in it. A client that cannot see how its own systems work cannot easily change course, or change supplier; the next phase always feels necessary. Letting a client build the answer itself, in weeks, and keep the knowledge, is anathema to a business built on embedding and hours.
AI turns this from lucrative to untenable.
When work that took a team a quarter can be done by one person in a week, the billable-hour engine runs in reverse: the very tools that raise productivity shrink the revenue those hours produce.10 Some firms already report their own AI tools saving close to a third of a consultant’s time,11 yet compensation still tracks hours, and the model still depends on the client never owning what was built. This is not a criticism of the people. It is arithmetic.
It is worth being precise about what is under pressure. Strategy and advisory, genuine judgement about what to do and why, is a distinct and durable discipline: it can stand on its own, it is often the precursor to everything that follows, and it is work we do ourselves. What the economics above describe is something else, the multi-year transformation, migration and rollout programmes that are long by design and notoriously slow to deliver on their original promise. It is that model AI turns from lucrative to untenable, not the counsel that precedes it. Building with AI needs neither a standing team nor an open-ended programme: it needs a small number of senior people who carry a problem all the way to a working, owned system, knowledge included, and who are paid for the outcome rather than the hours. The same logic governs the software an enterprise rents to run itself, the deeper version of the same dependency, which we turn to next.
The advice was outsourced, but so was the knowledge, and knowledge is the harder dependency to unwind.
The technology has arrived, and almost nobody has scaled it
Everything so far describes the supply of capability, and it is plentiful, cheap and getting cheaper. The demand side tells a different story. To test it we ran a pulse survey of senior leaders at large enterprises, and the picture is consistent: the technology has arrived, and almost nobody has scaled it. The bottleneck has moved from building to adopting, and adoption has stalled.39
Q: “Where is your organisation on AI adoption?” Single select, n=113.
Only one enterprise in eight has scaled AI across operations. The rest are piloting, running isolated production use, or have not begun. Hover a band to hold it.
Source: Covelent AI Adoption Pulse 2026, n=113.39
The independent series put that in context. US Census Bureau data for the two weeks ending 3 May 2026 shows 19.8% of all US businesses using AI to produce goods or services, rising to 37% of firms with 250 or more employees.40 Eurostat records 19.95% of EU enterprises using AI in 2025, and 55% of large enterprises.41 Our sample sits at the large end of that distribution, so a high rate of some use is expected and is not the interesting number. Stanford’s AI Index reports that in 2025 AI agent deployment remained in the single digits across nearly all business functions. Whatever is happening between first use and operational scale, it is not conversion.
Duration makes the point sharper. Sixty-two per cent of the leaders we surveyed have had initiatives sitting at pilot or evaluation for six months or more. Work published by MIT’s NANDA initiative found the same shape from the other direction: enterprises above $100m in revenue lead on pilot count but record the lowest pilot-to-scale conversion, taking nine months or longer to production against ninety days for the best-performing mid-market firms.43
The bottleneck has moved from building to adopting, and adoption has stalled.
For seven in ten enterprises, the blocker is control, not the technology
Asked for the single biggest blocker to moving AI into production, the leaders did not point at the models. They pointed at control. Compliance, data governance, vendor lock-in and legal exposure, four faces of the same question of who governs the system, together account for roughly seven in ten of the biggest blockers named, far ahead of cost, model reliability or skills.39 What stands between a large enterprise and production AI is not whether the technology works, but whether the organisation can run it on terms its own risk, legal and compliance functions will accept.
Q: “What is the single biggest blocker to moving AI into production?” Single select, seven options, all shown. n=113, rounded to whole percentages.
Control, not capability, is the binding constraint. Compliance, data governance, vendor lock-in and legal exposure, all questions of who governs the system, account for roughly seven in ten of the biggest blockers named, far ahead of cost, skills and model reliability.
Source: Covelent AI Adoption Pulse 2026, n=113.39
The blockers are not separate. Lock-in feeds the rest.
Vendor lock-in does not sit apart from the governance concerns. It sits upstream of them. The software and services incumbents an enterprise already depends on have every reason to keep it inside their ecosystem, and they use the compliance conversation to do it, steering customers towards the approved, integrated, rented option and warning that open-weight or self-hosted models are risky, unsupported or non-compliant. The stakeholders advising on how to adopt AI safely are frequently the ones with most to lose if the enterprise owns it.
Independent data points the same way. Eurostat asked enterprises that had considered AI and not adopted it why: 52.5% cited lack of clarity about the legal consequences, and 48.8% cited data protection and privacy.41 The Bank of England and FCA found the largest perceived regulatory constraint on AI in UK financial services is data protection and privacy, with 33% of firms describing regulatory burden as the main constraint.42
Nobody else appears to have measured the friction itself.
Three quarters of the leaders we surveyed had an AI initiative paused or blocked by legal, compliance or risk in the last twelve months. We went looking for someone else’s figure to check ours against, and there is not one. No government statistical series, no regulator survey and no independent study we could find puts a number on how often governance stops an AI project, or how long it holds one up. Adoption is measured exhaustively. The friction that determines adoption is not measured at all, which is its own comment on how the problem is understood.
The scarce skill is making AI reliable, compliant and owned
Almost everything written about AI is about what AI can do, which is the easy part now, and not where the difficulty or the value sits. Every serious field has this divide. In physics, one half works out what is theoretically possible, and the applied half makes it hold up in the world, at cost, under real constraints, in something a person actually uses. These are different disciplines, and the second is where things get built.
Applied AI is our name for that second half. It is not a strategy deck, and not a pilot that never leaves the lab. It is the engineering work of taking what the frontier makes possible and turning it into systems a business runs on, connected to real data, inside real governance, doing real work, and owned by the organisation that depends on them. Knowing that a model could in principle automate a workflow is now common. Standing that workflow up so it is reliable, compliant, auditable and yours is the scarce skill, and most of the job.
Most of that work sits outside the model, in two places. One is the harness, the code around the model that gathers its context, calls its tools, checks its output and decides when to stop; it now moves measured performance more than the choice of model does, and it has a section of its own later. The other is the ontology, the connected layer underneath that gives an agent something reliable to reason over, and it is the centrepiece of what Build-to-Own hands you.
And building this way is no longer exceptional.
In the past twelve months the number of new AI projects created on the largest public code platform grew by 178 per cent year on year, close to seven hundred thousand of them, after a 98 per cent rise the year before. Around 85 per cent of developers now build with AI tools, on two independent surveys.44 The barrier to turning an idea into a working, owned system has fallen through the floor. What the rest of this paper sets out is why the thing you build should be owned, what owning costs against renting, and how we deliver it.
You cannot own what you rent
Building with frontier models is now within reach of any organisation. Depending on them is a different matter. Access to an external model is access on someone else’s terms, and those terms are set by the provider, the market, and increasingly by governments.
Renting is not one option of two. It is the default, overwhelmingly. In our own survey, 69 per cent of senior decision-makers said their AI runs all or mostly on third-party and cloud APIs, against 4 per cent mostly on models they own.39 The question is not whether to move from owning to renting. It is whether an enterprise ever intends to move the other way.
The risk is not hypothetical. In July 2026, ByteDance and Alibaba disabled user-created humanlike AI agents to comply with new Chinese rules on anthropomorphic AI.12 Capabilities that businesses had built on those platforms disappeared on a regulator’s timetable. External providers also retire models on their own schedule; forced migration is a routine feature of renting, not an exception.13
And the intervention is not only Beijing’s.
In June 2026 the US Commerce Department ordered Anthropic to suspend its two most capable models, Claude Fable 5 and Mythos 5, three days after their 9 June launch, under export-control powers; access went dark worldwide for eighteen days.14 They returned only once the directive was withdrawn, and then with a new blocking classifier that fenced off some of the very capabilities that had set the models apart.15 Now the pressure runs the other way as well. As this paper goes to press, the US administration is weighing restrictions on domestic firms’ use of Chinese open-weight models, through procurement rules and Entity List pressure rather than an outright ban.16 Since most leading open-weight models are now Chinese, that would push American enterprises back onto a handful of closed providers. A closed model can be switched off on a government’s timetable, and the open one closed off by the same hand.
The pattern is not new. It is the deeper version of the same dependency.
The core systems enterprises already rent show it. Oracle’s licensing and audit regime is built to keep customers renewing rather than leaving,17 and it can reprice access on those already locked in: when its Java terms moved to a per-employee model, analysts put the rise at two to five times the old cost.18 Salesforce customers sit inside an ecosystem projected to earn its partners over six dollars for every dollar Salesforce makes,19 a dependence measured in specialist developers as much as licences. Even cloud carries egress fees and switching barriers the UK competition regulator found to harm competition across a £9 billion market.20
This is why ownership matters, and why regulation is pushing the same way. Ninety per cent of organisations already regard data held locally as inherently safer,21 and obligations from the EU AI Act onward are moving sensitive workloads inside the perimeter.22 The frontier can be rented to build. What runs the business should be owned.
The frontier can be rented to build. What runs the business should be owned.
The dependency you are told is permanent is already fragmenting
The usual objection to owning your stack is that the frontier is too far ahead and too fast-moving to keep up with in-house. That is true today, and it is exactly why we rent it to build. But the frontier you can see is not the real one: labs never stop iterating, and anything released publicly implies something better already in training, so the public models are a lagging indicator of where the frontier actually is. And the gap that makes renting worthwhile is closing on three fronts at once.
The models are commoditising.
The best open-weight model trailed the best closed one by eight per cent in early 2024 and by under two per cent a year later;23 on a single capability index the best open models now lag the closed frontier by roughly four months, about one release cycle, a gap that is shrinking rather than widening.24 The frontier also leaks: in June 2026 Anthropic told the US Senate Banking Committee that operators tied to a Chinese lab had run some 28.8 million exchanges through about 25,000 fraudulent accounts to distil cheaper models from Claude’s own answers, even though the leading labs had themselves trained on the public internet without consent.25 What is frontier and rented today is open and local tomorrow.
And they are getting cheaper and smaller.
Capability is getting cheaper by roughly fifty times a year4 and smaller by more than a hundred times for the same benchmark,5 to the point where one-bit models now run efficiently on ordinary CPUs, no specialised accelerator required.26
The hardware is diversifying off a single vendor.
NVIDIA still supplies the overwhelming majority of AI accelerators, but the largest labs are deliberately spreading their bets. Anthropic states that it runs Claude across three chip platforms, Google TPUs, Amazon Trainium, and NVIDIA GPUs.27 OpenAI, NVIDIA’s largest customer, is designing its own accelerators with Broadcom and sourcing memory from Samsung and SK Hynix.28 When the suppliers of compute are themselves reducing their dependence, an enterprise’s dependence on a single rented stack looks less like prudence and more like risk.
Reported and estimated accelerator counts for notable training runs and clusters. In June 2026 the world's fastest supercomputer reached 2.2 exaflops with no GPUs at all, on 13.8 million CPU cores;29 NVIDIA still powers more than 400 of the world's 500 fastest machines. Log scale.
Source: Covelent analysis of reported and estimated cluster sizes.
Put together, these trends point one way. The rented layer thins; over time, even the build comes in-house. The one thing that does not change is ownership. Nothing rented is ever owned, so the durable position is to own what runs, and to have built it so that a change at the frontier is an upgrade rather than an outage.
AI is spending about $680 billion a year it has not yet earned
There is a deeper reason to own rather than rent, and it is the price itself. The rented price is not merely temporary, it is not even real. Behind today’s cheap compute sits a build-out with no precedent and an arithmetic that does not yet close. In 2026 the largest cloud providers are guiding to roughly $600 billion of capital expenditure between them, most of it for AI, up from about $230 billion three years earlier. A widely cited framework takes the chip run-rate, doubles it for the rest of a data centre, and doubles it again for a normal software gross margin, to estimate the annual revenue the build-out must eventually earn to pay for itself.49
Spending has multiplied; the revenue to justify it has not arrived. Big-technology AI infrastructure capital spending by year, against the annual revenue that build-out implies it must eventually earn: roughly $800bn, against about $120bn earned today. The distance between the two, close to $680bn a year, is what today’s cheap prices are borrowing from.
Source: Covelent analysis of company filings and venture-capital estimates.49
None of this predicts a collapse. It is a statement about price. Compute sold below cost is compute sold on someone else’s balance sheet, and balance sheets get repriced. The organisations most exposed are the ones that have built their operations on the assumption that today’s rented price is the real one.
The price you rent at is subsidised
A $200 plan hands power users thousands in compute.
There is one more reason the rented layer is not the safe default it looks like: the price is not real. The subscriptions and API rates that make frontier AI look cheap are being sold below cost, funded by investors buying market share. In 2025 OpenAI spent about $1.69 for every dollar of revenue it earned, and is on course to burn tens of billions more this year.30 That is not a stable price. It is a promotional one, and the two lines below cannot stay this far apart.
Illustrative. What a heavy user pays for frontier AI sits far below the cost of the compute they consume; independent analysis puts a $200 plan's true API value at $8,000 to 14,000 a month.31 As investor subsidies unwind, the price of anything rented rises to meet its cost.
Source: Covelent analysis of SemiAnalysis subscription economics.
Independent analysis of the leading consumer tiers found a $200-a-month plan delivered roughly $8,000 of API tokens on one provider and close to $14,000 on another: a subsidy of forty to seventy times the sticker price.31 Providers lose money on their heaviest users by design, to take the market before the reckoning comes.32
Almost nobody has priced what that unwinding would cost them.
Asked whether they knew what they would pay if their AI providers repriced or withdrew a model tomorrow, only 13 per cent of the senior decision-makers in our survey could answer fully. Fifty-one per cent said flatly that they could not, and a further 35 per cent only roughly.39 Seven in eight enterprises are carrying an exposure they have never sized, on a price the provider has already said is below cost.
The real cost is on the API meter, and it moves one way.
At list, the flagships already cost real money per million tokens: GPT-5.5 at $5 in and $30 out, Claude Opus 4.8 at $5 and $25, Gemini 3.1 Pro at $2 and $12.33 When the labs turn from land-grab to returns, and they are visibly preparing to,34 the levers are obvious: raise prices, meter usage, or cut the limits that make today’s plans feel unlimited. Everything rented reprices with them, on their timetable. A system that runs on models and hardware you own has a cost you set and can forecast; a workflow that rents its intelligence inherits whatever price the market needs once the subsidy ends.
Providers lose money on their heaviest users by design, to take the market before the reckoning comes.
Rent versus own: the cost over time
The unit economics are not close.
Every argument so far has a number attached, and for most boards it is the deciding one. Renting and owning are not just different levels of control; they are different cost structures. Renting is pure operating expense: per-seat software and per-token inference, billed every month, rising with use, and leaving nothing owned when the contract lapses. Build-to-Own converts most of that into a one-off capital cost, the build and the infrastructure it runs on, followed by a low and predictable running cost, and it leaves an asset on the balance sheet.
Rented, the flagships cost $5 to $30 per million tokens.33 The same inference on an owned open-weight model is a few cents of electricity once the hardware is bought;35 the real expense is the server, an eight-GPU machine at $200,000 to $400,000, or reserved cloud at a few dollars an hour.36 Past a threshold that published analyses put at around $20,000 of API spend a month, or a few hundred million tokens, owning is simply cheaper, and the saving compounds.37
Illustrative, on public unit costs. A workload renting frontier models and seats accumulates cost in a straight line that never flattens. The same capability built and owned costs more on day one and less within the first year, with the gap widening every month after. The figures scale with the workload; the shape does not.
Source: Covelent analysis, on public unit costs.
Relative cost to serve the same task; routing every task to the frontier = 100.
You do not pay the frontier price for work a smaller model can do. Relative cost to serve the same operational task. Sending only the hard residual to a frontier model cuts cost by about 85%, and a cheap-first cascade by up to 97%. Both still meter every tier. An owned stack takes the cheap tiers to zero and meters only the small residual that still goes out.
Source: Covelent analysis of published routing and fine-tuning results.48
Read the marks and the split is honest. Renting leads where a rented tool should, on raw breadth and on standing something up in minutes. Owning leads on everything you then run the business on. Select a row for the detail.
Renting buys motion; owning builds an asset.
The rented line is operating expense: it recurs, it rises, and it leaves nothing behind. The owned line is capital plus a small running cost, and what it buys can be depreciated, adapted and reused. A bespoke tool that retires a six- or seven-figure software estate pays for itself inside a year and then costs almost nothing.38 And because today’s rented prices are subsidised, as the previous section set out, the real crossover arrives sooner than list prices suggest and moves further towards owning as that support unwinds. This is not an argument to own everything: for spiky, low-volume or non-sensitive work, renting elastic capacity is cheaper, and we would advise exactly that. Build-to-Own is about the systems the business runs on every day, where the meter never stops.
A way of working, not a thing you buy
Applied AI is a discipline, not a product on a shelf: an organisation seeing what is now possible and building it into how it actually runs. Ours has a name, Build-to-Own, and it sits deliberately outside the two options an enterprise is usually offered. A consultancy advises and leaves, or worse, embeds a dependency it has every incentive to deepen. An internal team builds slowly, inside existing governance and rented software, and rarely ships. We do neither: we build fast on the frontier, and hand over something you own outright.
Closed models on scarce hardware, rented for speed. Build fast; nothing here is permanent.
We rent the frontier only to build, in a lab, for as long as closed models on scarce hardware are ahead. Then we deploy once. Everything that runs the business is owned outright and runs inside your perimeter, and the rented part shrinks over time.
The principle is in the name. You cannot own what you rent. We rent the frontier only to build: for as long as closed models on scarce hardware are ahead, renting them accelerates delivery, so we do our building in a lab. Then we deploy once. Everything that runs the business, the harness that drives the models, the ontology that connects your data, the infrastructure it runs on, the local models and agents that serve it, and the data itself, is owned outright and runs inside your perimeter.
And the rented part is shrinking. As open-weight models close the gap, compute diversifies off a single vendor, and inference gets cheaper, even the build moves in-house over time. Ownership is the part that does not change.
Connect the data, and the value surfaces
The ontology is where the value concentrates.
Solutions built in isolation stay isolated. What moves a team forward is access to the disparate data that underpins its work, connected. This is the layer Build-to-Own produces and hands over, and it is where most of the durable value lives. It has four parts, and models are only one of them.
A database knows a supplier and an order as two rows in two tables. The ontology knows that this supplier supplies that material, which feeds that plant, whose delay puts this customer's revenue at risk, and it knows what may be done about it. Four parts sit on the connected data; only one is a model.
Four parts sit on top of the connected data, and only one of them is a model. Data pipelines ingest continuously from every system, databases and applications, documents and spreadsheets, sensor and telemetry feeds, third-party APIs and real-time streams, and resolve the same real-world thing across all of them, so a customer in the CRM, the ledger and the support desk is understood as one entity rather than three records. An ontology then gives that data shape: the entities that matter, the typed relationships between them as a graph, and the actions that can change them. Learning loops, machine learning and human feedback, refine the system from its own outcomes every cycle rather than issuing one-off predictions. And governed agents carry a decision through to execution, permissioned, logged and reversible.
A database knows a supplier and an order as two rows in two tables. The ontology knows that this supplier supplies that material, which feeds that plant, whose delay puts this customer’s revenue at risk, and it knows what may be done about it. That is the difference between storing the business and modelling it: the same connected model can answer a question, surface a risk, and trigger the action that resolves it, because entities, relationships and actions live in one place rather than scattered across the systems that happen to hold them.
Three words get used for this, and they are not interchangeable.
The vocabulary here is muddled, often deliberately, because each vendor sells whichever of the three words is in fashion. They are worth separating, because they stack rather than compete.51 A semantic layer is the dictionary of the business: the agreed terms, what an active customer is, what counts as revenue, what a plant is, so that finance, operations and the commercial team mean the same thing by the same word. An ontology is the formal grammar written on top of that dictionary: the entities that matter, the typed relationships between them, the constraints that must hold, and the actions that may be taken, set down precisely enough for a machine to act on. A knowledge graph is the instance: the ontology filled with your actual data and held as a graph that answers queries and updates as the business moves.
The sequence matters in practice. A semantic layer with no ontology gives you agreed definitions and no structure a machine can use. An ontology with no semantic layer formalises terms the business never agreed. And a knowledge graph built without either is a large store of connected rows whose shape came from whatever the source systems happened to hold, which is the failure mode most graph projects hit. Dictionary, then grammar, then graph. That order is also why the layer transfers to your people: they can read the definitions and argue with them before a single query runs.
None of this is exotic, and none of it is ours to gate. The techniques, retrieval, low-rank tuning, quantisation, agentic orchestration, are bounded by mathematics and physics, not by the economics of an advisor. They are open to everyone and expand every week. What we bring is the judgement to assemble them into something that fits your business, and to leave it owned.
The harness now moves the numbers more than the model does
There is a second place the value has moved, and outside the labs almost nobody is watching it. A model on its own does nothing. What turns weights into work is the harness around them: how context is gathered and pruned before the model sees it, which tools it may call, how its actions are executed and checked, when it retries, when it escalates, and when it stops. Recent work decomposes that into six runtime jobs, observation, context, control, action, state and verification, and argues that the bottleneck in long-horizon work now sits at least as much there as in the model itself.52
Same weights, different harness, a different answer.
The evidence is uncomfortable for anyone who buys on benchmark scores. In a controlled experiment across three frontier models and three harness configurations, changing the harness moved results by 8.5 to 13 percentage points, while changing the model moved them by 2.5 to 5. Harness variance was 7.8 times model variance, and six of nine model comparisons reversed their ranking when the harness changed.53 Published results show the same effect: one frontier model scores 45.9 per cent on a hard software benchmark under a standardised research scaffold and 55.4 per cent under its own vendor’s harness; another scores 58.6 per cent under an open scaffold and 72 to 75 per cent under the lab’s.53 A benchmark number is a property of a model and a harness together, and published alone it says less than it appears to.
Per cent of tasks resolved. Bars run from zero to 70%.
Tasks resolved on the same benchmark subset, three frontier models run under three harness configurations. Each model gains 8.5 to 13 points from the harness alone, against the 2.5 to 5 points that separate the models inside any one configuration. The strongest model on a minimal harness (55.0) finishes below the weakest model on a full one (60.5), which is why a score published without its harness is close to meaningless.
Source: Y. Zhang et al., controlled factorial experiment, 2026.53
The gains are available without a bigger model or a bigger bill.
Because harness improvements are engineering rather than training, they land immediately and at almost no marginal cost. Adding prompt optimisation, middleware and a verification step lifted one terminal benchmark from 52.8 to 66.5 per cent; automated tuning of the harness alone lifted a frontier model from 69.7 to 77.0 per cent on the same benchmark, with the weights untouched.53 This is the cheapest performance in the stack, and it accrues to whoever owns the harness.
And a disclosed harness still tells you nothing about your work.
Correcting for the harness fixes the comparison, not the relevance. A public benchmark, however carefully run, is somebody else’s questions. It cannot tell you whether a system will reconcile your systems, follow your policies, survive your hard exceptions or hold up at your volumes. The only evidence that answers those questions is generated on a private set of representative tasks drawn from your own work, where you have first written down what a correct result is and which errors matter most. That evidence is also what identifies the smallest permitted system that clears your bar, a testable claim rather than an argument from parameter count.
Which is why you should build your own.
The obvious objection is that the labs’ harnesses are better, and today they often are. The agent frameworks from Anthropic, OpenAI and Microsoft ship more options than any single enterprise would write for itself, and we use them in the lab for exactly that reason. Three things argue against making one of them the thing you run on. First, the orchestration does not travel: tool integrations written to the open protocol survive a change of provider, but the workflows, permissions and routing built into a vendor’s harness do not, and support for that protocol varies between them.54 Second, the harness is where your business goes. Which systems an agent may touch, what a fourth-line escalation looks like, which check must pass before a payment run is released: that is your governance expressed as code, and it is the last thing to hand to a third party. Third, and most awkward for the vendors, harnesses are small. The team behind the standard software benchmark publishes an agent of roughly a hundred lines of Python that resolves over 74 per cent of it, and the leading open harness, MIT-licensed and model-agnostic, sits around 72 per cent.55 Most of the available gain is reachable with code a competent team can read in an afternoon.
So the harness is the one layer where owning is barely a trade at all. It is cheap to build, it is where domain knowledge accumulates, and it is what keeps the model swappable: when a better or cheaper model arrives, an owned harness is pointed at it in a day, while a rented one waits for the provider to support it. Rent the frontier to build. Own the harness that runs.
Where this argument could break
A position paper that only makes its own case is marketing. Here are the strongest arguments against ours, and where they change our advice.
What if the frontier keeps pulling away?
It is possible that closed models stay far enough ahead that renting them remains worthwhile indefinitely, and open weights never quite catch the frontier for the hardest work. If so, the “build” side of Build-to-Own stays rented for longer than we expect. It does not change the conclusion: you still run owned, because the reasons to own, data residency, continuity, and freedom from a provider’s terms, are independent of how good the rented model is.
What if owning the stack is simply more expensive?
For some workloads it is. Renting elastic cloud capacity for spiky, non-sensitive tasks can be cheaper than owning idle hardware, and we would advise exactly that where it applies. But today’s rental prices flatter the comparison, because they are subsidised, as the previous section sets out; as that support unwinds, the balance tips further towards owning. Build-to-Own is not a demand that everything move on-premise tomorrow; it is a principle about what must be owned, the data, the models and agents that touch it, and the systems the business depends on, and a recognition that the cost of ownership is falling as fast as anything else in this paper.
What if the regulatory picture reverses?
Rules may loosen as easily as they tighten. But the events of the last two years, models withdrawn by governments on both sides of the Pacific, point to more intervention, not less, and an owned system is robust either way. Ownership is the low-regret choice: it costs a little more optionality today and removes a large tail risk tomorrow.
We would rather state these openly than have a prospect find them for us. On balance, none of them recommend renting what runs your business.
An AI lab you own
We do not hand you a system and leave. The first build, one real problem solved end to end in days, becomes a permanent internal AI lab: one that runs inside your perimeter and that you own outright, not a service you subscribe to.
What makes it different is not a secret; it is a stance. The lab both ships, building working tools against real problems, and continuously tests the frontier against your own work rather than a public benchmark. So the question every organisation now faces, what should we run ourselves and what should we still rent, is answered by your own evidence, week after week, instead of taken on a provider’s word. That is the opposite of renting intelligence and hoping the roadmap happens to suit you.
Building used to be slow.
So advice was the product.
That constraint is gone. What stays scarce is judgement about what to build, and ownership of what you build. Rent the frontier. Own what runs.
The full argument, laid out as a designed document:
Download the full paper (PDF)References
- 1METR, Measuring AI Ability to Complete Long Tasks, 2025.
- 2Stanford HAI, AI Index 2025, Technical Performance.
- 3Stanford HAI, AI Index 2025 ($20.00 → $0.07 per million tokens).
- 4Epoch AI, LLM inference price trends, 2025.
- 5Stanford HAI, AI Index 2025 (PaLM 540B → Phi-3-mini 3.8B; ~142× parameter reduction).
- 6SWE-bench Verified, 2024 to 2026; best open-weight vs best closed, % resolved (illustrative). swebench.com.
- 7SWE-bench Verified and SWE-bench Pro leaderboards, 2026 (GLM-5.1 and DeepSeek V4 lead individual benchmarks).
- 8David Maister, Managing the Professional Service Firm, 1993.
- 9D. Duncan, T. Anderson and J. Saviano, AI Is Changing the Structure of Consulting Firms, Harvard Business Review, Sep 2025.
- 10Thomson Reuters Institute, The $2,000-hour problem, 2025.
- 11Consultancy.uk, Is AI about to kill the billable hour?, Jan 2026.
- 12South China Morning Post, 5 Jul 2026 (China Interim Measures on anthropomorphic AI).
- 13OpenAI model deprecations. developers.openai.com/api/docs/deprecations
- 14Forbes, 16 Jun 2026; National Law Review, Jun 2026 (Commerce Dept export-control directive; Fable 5 and Mythos 5 launched 9 Jun, suspended 12 Jun).
- 15Axios, 27 Jun 2026; TechCrunch, 1 Jul 2026 (directive withdrawn 30 Jun; redeployed 1 Jul with a new blocking classifier).
- 16TechCrunch and CNBC, 24 Jul 2026 (US weighs restrictions on Chinese open-weight models; Nvidia, Microsoft and Meta warn against premature restrictions); US NTIA, Dual-Use Foundation Models with Widely Available Weights, 2024.
- 17The Register, Oracle ULA audits are a license to bill, May 2024.
- 18The Register, Oracle’s new Java license terms, Jul 2023 (Gartner estimate).
- 19IDC (Salesforce-commissioned), The Salesforce Economy, 2021.
- 20Competition and Markets Authority, cloud services market investigation, final decision summary, 31 Jul 2025 (egress fees and technical barriers found to harm competition in a £9bn market).
- 21Cisco, 2025 Data Privacy Benchmark Study.
- 22European Commission, EU AI Act, GPAI obligations (from 2 Aug 2025).
- 23Stanford HAI, AI Index 2025 (Chatbot Arena, 8.04% → 1.70%).
- 24Epoch AI, Capabilities Index, 2026 (open models roughly four months, about one release cycle, behind the closed frontier).
- 25Anthropic, letter to the US Senate Banking Committee, 10 Jun 2026; Forbes and Tom’s Hardware, 25 Jun 2026 (~28.8m exchanges via ~25,000 accounts, late Apr to early Jun 2026).
- 26Ma et al., Microsoft Research, The Era of 1-bit LLMs, arXiv:2402.17764.
- 27Anthropic, Expanding our use of Google Cloud TPUs, 23 Oct 2025.
- 28OpenAI and Broadcom, 10 GW collaboration, 13 Oct 2025; Samsung/SK Hynix memory, Oct 2025.
- 29TOP500, 67th list, Jun 2026 (LineShine, 2.198 exaflops HPL, 13.79m CPU cores, CPU-only); Al Jazeera and Network World, 24 Jun 2026.
- 30The Information; SemiAnalysis, 2025 to 2026 (OpenAI spend vs revenue; multi-billion cash burn).
- 31SemiAnalysis, consumer AI subscription economics, 2026 ($200 plan ≈ $8,000 to 14,000 API value; 40 to 70× subsidy).
- 32SemiAnalysis, 2026 (providers loss-making on heavy users at low utilisation).
- 33Published API list prices, Jul 2026, per 1M tokens (GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro). pricepertoken.com.
- 34The Information, 14 Apr 2026 (Anthropic Enterprise moved to $20/seat plus API pricing); OpenAI Codex pricing aligned to API token usage, 2 Apr 2026, extended to all Enterprise plans 23 Apr 2026.
- 35Marginal inference cost of a self-hosted open-weight model (energy and compute per 1M tokens).
- 36NVIDIA H100 8-GPU server, $200k to 400k; reserved cloud ≈ $1.5 to 10 / GPU-hour, 2026.
- 37Self-host vs API break-even ≈ $20k/mo API spend, or ~100 to 256M tokens/mo, 2026.
- 38Salesforce published list pricing, 2026 (per user, per month); saving depends on scope and seats.
- 39Covelent AI Adoption Pulse 2026. Online survey, fielded 1 Jun to 6 Jul 2026. n=113 senior decision-makers (CIO, CTO, CDO, CISO, General Counsel and PE operating partners) at enterprises of $200m to $80bn revenue; United States (46), Europe (35), United Kingdom (24), Middle East (8). Unweighted.
- 40US Census Bureau, Business Trends and Outlook Survey, reference period ending 3 May 2026 (19.8% of US businesses; 37% of firms with 250+ employees).
- 41Eurostat, Use of artificial intelligence in enterprises, 2025 (19.95% of EU enterprises; 55% of large enterprises; barriers among non-adopters).
- 42Bank of England and FCA, Artificial intelligence in UK financial services 2024, Nov 2024 (n=118 firms).
- 43MIT NANDA, The GenAI Divide: State of AI in Business 2025, Jul 2025. Preliminary findings; survey component n=153 recruited at industry conferences, a convenience sample, not peer-reviewed.
- 44GitHub Octoverse, 2024 and 2025 (+98% then +178% year on year; ~694,000 new AI projects in twelve months); Stack Overflow and JetBrains developer surveys, 2025.
- 45F. Dell’Acqua et al., Navigating the Jagged Technological Frontier, Harvard Business School Working Paper 24-013, Sep 2023 (field experiment, n=758).
- 46METR, controlled study of experienced developers, Jul 2025 (~19% slower on their own repositories). arXiv:2507.09089.
- 47Epoch AI, AI data-centre cost breakdown, 2026 (~60% servers, ~30% facility, ~7% energy); SemiAnalysis, AI server bill-of-materials.
- 48RouteLLM (UC Berkeley, 2025); FrugalGPT (Stanford); open-model fine-tuning results, 2024 to 2025.
- 49D. Cahn, Sequoia Capital, AI’s $600B Question, Jun 2024; NVIDIA FY2026 data-centre segment revenue; aggregate hyperscaler capital expenditure guidance, 2023 to 2026.
- 50IBM, Cost of a Data Breach 2025 (shadow-AI breach premium ~$670,000).
- 51C. Hunt, Semantic model vs ontology vs knowledge graph, 2025; W3C, OWL 2 Web Ontology Language and RDF 1.1 recommendations.
- 52J. Guo et al., From Question Answering to Task Completion: A Survey on Agent System and Harness Design, Jun 2026. arXiv:2606.20683 (harness decomposed into observation, context, control, action, state and verification).
- 53Y. Zhang et al. (Tulane, Rutgers, Virginia Tech), Stop Comparing LLM Agents Without Disclosing the Harness, May 2026. arXiv:2605.23950 (controlled factorial experiment on a SWE-bench Verified subset; harness variance 7.8× model variance; SWE-bench Pro and TerminalBench 2.0 figures).
- 54Model Context Protocol specification, 2025 to 2026; OpenAI Agents SDK portable sandbox manifests, Apr 2026; Microsoft Agent Framework, Build 2026. Tool integrations port; orchestration, permissions and routing do not.
- 55SWE-agent team (Princeton and Stanford), mini-swe-agent, 2025 to 2026 (~100 lines of Python, >74% SWE-bench Verified); All Hands AI, OpenHands, MIT licence (~72% SWE-bench Verified, model-agnostic).