Field Notes · Issue 001 · 29 July 2026

AI came up on 337 earnings calls last quarter. Most leadership teams still can't draw the stack.

It is the highest count FactSet has recorded in a decade, and the spending behind it is not subtle: the four largest US hyperscalers are committed to roughly $725 billion of capital expenditure this year alone. Yet sit in almost any boardroom and ask for the map — what sits under what, what is genuinely deployed, what is still a slide — and the room goes quiet.

That silence is expensive. It is how vendors sell a model as a platform, how pilots get funded twice, and how a capable executive ends up nodding at a word they would rather have questioned. This is the map, drawn for the people who have to decide. Three layers, two rings around them, thirty-seven terms — every claim carrying a date and a source.

I'm reading this as Written for the people who have to decide, not the people who have to build.
Draw me the AI stack thumbnail
337
S&P 500 earnings calls citing AI — a 10-year high
$725B
Big-four hyperscaler capex committed for 2026
39%
Report any EBIT impact from it

The map

Three layers, not one word cloud

Most AI diagrams fail because they flatten three different questions into a single list. How it is built, what it can do, and where it is deployed are separate questions with separate answers — and confusing them is exactly how a foundation model gets sold as a business outcome. Every term sits in one of these three layers, or in one of the two rings around them — the infrastructure it runs on, and the governance that keeps it defensible. Tap any term to jump.

Applications

Layer 3 · where it is deployed

The only layer your customers, your staff and your P&L ever meet directly. If a business case does not land here, it is not a business case.

Capabilities

Layer 2 · what it can do

Ordered as the industry itself is sequencing them — perception, generation, reasoning, agency, and finally the physical world. Not all five are equally deployed, and the gap between them is where most budgets go missing.

Foundations

Layer 1 · how it is built

You will never buy this layer directly, and you should never be charged for it as though it were the product. Knowing the vocabulary is how you tell a genuine capability claim from a repackaged API call.

Infrastructure

Ring 1 · where it runs and how it connects

Underneath all three layers. Ignore it and the real constraints — cost, latency, jurisdiction, integration — only surface in year two.

Governance & control

Ring 2 · what keeps it defensible

Around everything. For people who have to decide rather than build, this ring is not a footnote — it is roughly half of what you are personally accountable for.

Read bottom-up to understand how it works. Read top-down to decide what to buy. Read the rings before you sign anything.

RING 2 — GOVERNANCE & CONTROL · 6 TERMS Regulation · Guardrails & prompt injection · Evals · Shadow AI · Data readiness · Decision intelligence RING 1 — INFRASTRUCTURE · 5 TERMS Applications Layer 3 — where it is deployed. The only layer that reaches your P&L. Copilots · Enterprise AI · Workflow automation · Predictive maintenance Vision · Conversational · Digital twin 7 TERMS Capabilities Layer 2 — what it can actually do, in the order the industry is sequencing it. Perception · Generative · Reasoning · AI agents & agentic · Multiagent Physical AI · Multimodal 7 TERMS Foundations Layer 1 — how it is built. You never buy this layer directly. Machine learning · Foundation model · LLM · Domain models · RAG · Fine-tuning Prompting · Hallucination 12 TERMS Compute & GPUs · Cloud, on-premise & sovereign · IoT & edge · MCP · Open vs closed weights Roughly half of what you are accountable for If a business case does not land here, it is not a business case Deployment falls off sharply down this list Priced at Layer 3 rates, this is plumbing you are paying twice for Read up to understand it Read down to decide
The 2026 AI stack — three layers, two rings, 37 terms. Field Notes 001.
What this means for you


Layer 1 of 3

Foundations — how it is built

Twelve terms. You are not buying any of these, and that is the point: when a proposal is priced against this layer rather than against an outcome, you are paying for plumbing. Every card opens with the plain-English version first.

Machine learningThe parent category

In plain EnglishSoftware that improves at a task by being shown examples, instead of being given rules by a programmer. Nobody writes "an invoice usually has a total in the bottom right" — the system infers it from ten thousand invoices.

How it actually worksYou supply labelled examples, the system makes a guess, measures how wrong it was, and adjusts itself slightly. Repeat several million times. The output is not a program you can read; it is a set of numbers that happen to produce good answers. This is why "explainability" is a genuine engineering problem rather than a documentation problem.

Why it matters: machine learning is the parent category — every other term on this page is a subset of it. When a vendor says "our AI learns from your data," this is what they mean, and the honest follow-up is: how much of our data, labelled by whom, and what happens when the process changes?

Where you'll meet it: credit scoring, demand forecasting, churn prediction, quality prediction on a production line. Mostly the unglamorous, high-return end of AI.

Deep learning & neural networksThe technique that unlocked the rest

In plain EnglishMachine learning arranged in many stacked layers, loosely inspired by neurons. "Deep" simply refers to the number of layers — the depth is what lets it learn subtle patterns nobody thought to describe.

How it actually worksEach layer transforms the data slightly and passes it on. Early layers in an image system detect edges; middle layers detect shapes; later layers detect "this is a cracked weld." No human specified those stages. They emerged from training, which is both the power and the governance headache.

Why it matters: deep learning is why AI stopped being a research curiosity around 2012 and why it needs so much compute. It is also the reason the infrastructure bill is so large — every one of the $725 billion of 2026 hyperscaler capex ultimately exists to run layers of arithmetic at scale.

Where you'll meet it: under the bonnet of every capability in Layer 2. You will rarely buy it as a named product.

Foundation modelThe expensive part

In plain EnglishOne very large model trained on a very broad body of data, then adapted to many specific jobs. The opposite of the old approach, where you trained a separate small model for each task.

How it actually worksTraining happens once, at enormous cost, by a handful of organisations worldwide. Everyone else adapts what exists — through prompting, retrieval, fine-tuning, or a thin application layer. This is the single most important structural fact about the current market: the expensive part has already been paid for by someone else.

Why it matters: it explains the strange economics you are seeing. A five-person company can ship a credible AI product because it is renting the foundation. It also means "we built our own AI" is almost never literally true — and asking which foundation model sits underneath is a fair, revealing question.

Where you'll meet it: named in the fine print of your vendor contract. Ask which one, which version, and where it runs.

Large language model (LLM)The famous one

In plain EnglishA foundation model specialised in language. It predicts what text should come next, which — at sufficient scale — turns out to look a great deal like writing, summarising, translating and reasoning.

How it actually worksIt has no database of facts and performs no lookup. It generates a plausible continuation, one fragment at a time, from statistical structure learned during training. Everything else — accuracy, citation, tool use — is engineering built around that core behaviour, not a property of the model itself.

Why it matters: this single design choice explains nearly every failure mode you have heard about. It also explains the workhorse use cases: Gallup's Q2 2026 data shows workplace AI use clusters heavily around drafting, revising and finding information — language work, which is precisely what these models do.

Where you'll meet it: ChatGPT, Claude, Gemini, Copilot, and the chat box now embedded in roughly every enterprise product you own.

Small & domain-specific modelsThe 2026 shift

In plain EnglishDeliberately smaller models, trained or tuned for one domain — legal, clinical, industrial maintenance, a single language. Less general knowledge, far lower cost, and able to run on your own hardware.

How it actually worksYou trade breadth for economics and control. A model that will never need to discuss Renaissance art can be a fraction of the size while outperforming a general model on your contracts, because it has seen far more of them.

Why it matters: Gartner named domain-specific language models one of its top ten strategic technology trends for 2026. For any organisation with data it cannot send outside the building — defence, pharma, banking, government, engineering IP — this is the line item that makes AI legally possible rather than merely desirable.

Where you'll meet it: on-premise and air-gapped deployments, sovereign-AI tenders, and any procurement conversation where "where does the data go?" is the first question asked.

Training vs inferenceTwo completely different bills

In plain EnglishTraining is teaching the model — once, at vast expense. Inference is using it — every single time, forever. Confusing the two is the most common costing error in enterprise AI.

How it actually worksTraining is a capital event measured in months and hundreds of millions. Inference is an operating expense measured per request. Your AI bill is almost entirely inference, and unlike seat-based software it scales with usage, not headcount — which is exactly why it surprises finance teams.

Why it matters: this is where budgets break. Industry surveys through 2025–26 found the large majority of enterprises overshooting AI infrastructure forecasts, typically by more than 25%, because usage-based pricing behaves nothing like the SaaS contracts finance is used to modelling.

Where you'll meet it: the second-year invoice. Ask for a modelled inference cost at three times your pilot volume before you sign.

Tokens & context windowThe unit you are billed in

In plain EnglishA token is a fragment of text — very roughly three-quarters of a word. The context window is how much the model can hold in mind at once: your prompt, your document, and its own answer, all counted together.

How it actually worksYou are billed per token in and per token out. When a conversation or a document exceeds the window, the earliest material silently drops away — the model does not announce that it has forgotten, it simply answers as though the missing part never existed.

Why it matters: it is the meter on your bill and the most common cause of "it worked in the demo." A vendor demonstrating on a two-page sample and deploying against your three-hundred-page tender specification is not running the same test.

Where you'll meet it: pricing pages, API contracts, and every complaint that begins "it forgot what I told it earlier."

Parameters & weightsThe size claim

In plain EnglishThe internal numbers a model adjusts during training. "70 billion parameters" is a size claim — roughly analogous to engine capacity, and about as reliable a guide to actual performance.

How it actually worksMore parameters generally means more capability and always means more cost to run. But training data quality, tuning method and how the model is deployed matter at least as much. Since 2025 the frontier has been moving toward smaller models that are better trained rather than simply larger ones.

Why it matters: parameter count is the specification most often quoted at executives and the least useful for deciding anything. Ask for accuracy on your documents, not size. A demonstration on your own data settles in an afternoon what a specification sheet cannot settle at all.

Where you'll meet it: vendor slides, model release announcements, and any conversation where someone is trying to sound technical.

Fine-tuningOften the wrong answer

In plain EnglishTaking an existing model and training it further on your own examples, so it adopts your tone, format or specialist judgement.

How it actually worksYou supply a curated set of input–output pairs and run a much smaller training job. It shifts behaviour and style reliably. It is a poor and expensive way to teach the model facts, because every time your facts change you must do it again.

Why it matters: fine-tuning is proposed far more often than it is needed, and it is usually the most expensive line in the quotation. If the requirement is "know our current products and policies," the answer is retrieval, not fine-tuning — and knowing that distinction is worth real money in a negotiation.

Where you'll meet it: in a proposal, priced at six figures. Ask what specifically cannot be solved with retrieval first.

RAG — retrieval-augmented generationThe workhorse pattern

In plain EnglishBefore answering, the system searches your own documents and hands the relevant passages to the model. The model then answers using that material rather than from memory.

How it actually worksYour documents are indexed and stored. A question triggers a search, the best-matching passages are pasted invisibly into the prompt, and the model composes an answer grounded in them — usually with a citation back to the source page. Update a document and the answer updates immediately, with no retraining.

Why it matters: this single pattern underpins the majority of genuinely useful enterprise deployments — policy assistants, tender search, standards lookup, technical support. It is also the honest answer to hallucination, because a system that must cite a page can be checked by a human in seconds.

Where you'll meet it: anything described as "trained on your documents." It almost certainly is not trained on them — it retrieves them, which is better.

Prompt & context engineeringThe one skill you can acquire this week

In plain EnglishThe craft of telling the model precisely what you want and giving it the right material to work from. It is the only item on this entire page that you personally can get better at, without buying anything.

How it actually worksOutput quality is governed far more by the instruction and the supplied context than by which model was chosen. "Context engineering" is the 2026 term for the discipline this matured into — deciding deliberately what goes into the context window, in what order, and just as importantly what is left out of it.

Why it matters: this is the real explanation for the gap between the colleague who says AI is useless and the one who says it saved them a day. They are using the same tool. Gallup's 45%-to-90% productivity spread by breadth of use is, underneath, a skill distribution — not a licensing one.

Where you'll meet it: every single day. It is also the cheapest intervention available to you: a half-day of structured practice moves an entire team.

HallucinationA feature of the design

In plain EnglishWhen a model states something false with complete confidence. Not a bug that will be patched — a direct consequence of a system built to produce plausible continuations rather than to look things up.

How it actually worksThe model has no internal sense of certainty about facts. Fluency and accuracy are produced by the same mechanism, which is why a wrong answer reads exactly as convincingly as a right one. Grounding, citation and human review reduce the rate substantially; nothing available today eliminates it.

Why it matters: McKinsey's survey work found inaccuracy to be the most frequently reported negative consequence of AI use, ahead of every other risk category. Any deployment without a named human check on the output is not a deployment, it is an exposure.

Where you'll meet it: in the one output nobody verified. Design the review step before the pilot, not after the incident.

What this means for you


Layer 2 of 3

Capabilities — what it can actually do

Seven terms, in the order the industry is sequencing them. Perception is mature and boring. Generation is everywhere. Reasoning arrived recently. Agency is the loud one, and the least deployed. Physical AI is the one your operations team should be watching. Deployment falls off sharply as you move down the list — that drop-off is the single most useful thing on this page.

Perception AIWave 1 · mature, proven, unglamorous

In plain EnglishMachines recognising things — images, speech, defects, faces, signatures, sounds. The oldest capability on this list and, for industrial businesses, still the one with the clearest return.

How it actually worksA model trained on many labelled examples learns to classify new ones. It does not understand what it sees; it matches learned patterns extremely well. Accuracy on a narrow, well-defined task now routinely exceeds a tired human on the third shift.

Why it matters: while attention has moved to chat, perception quietly delivers. It runs at the edge, needs no internet connection, produces a measurable defect-rate number, and does not hallucinate. If you need an AI win with a defensible ROI inside two quarters, it is usually here.

Where you'll meet it: visual quality inspection, ANPR at your gate, safety-PPE detection, document scanning, voice transcription.

Generative AIWave 2 · everywhere, and now normal

In plain EnglishProducing new content — text, images, code, audio, video, designs — rather than classifying existing content. The capability that made AI a boardroom subject in 2023.

How it actually worksThe model generates one fragment at a time, each conditioned on everything before it. Quality depends far more on the instruction and the material you supply than on which model you chose, which is why prompt discipline outperforms model-shopping in almost every organisation.

Why it matters: Gallup's Q2 2026 survey found 52% of US employees now use AI at work — the first time it has crossed half — with 15% using it daily. Access is no longer the constraint. The finding that should reshape your training plan sits deeper: productivity gains climbed from 45% among narrow users to 90% among those applying it across seven or more task types. Breadth of use, not access, is what pays.

Where you'll meet it: already inside your email, your CRM, your design tools and your developers' editors — frequently before anyone approved it.

Reasoning modelsWave 3 · the recent step change

In plain EnglishModels that work a problem through in steps before answering, rather than responding immediately. They spend more time — and more money — on hard questions, and get substantially better answers on them.

How it actually worksThe model generates an internal chain of intermediate steps, checks itself, and can revise before producing a final answer. This is why responses arrive more slowly and cost more per query. On multi-step analysis, engineering calculation and code, the improvement over standard generation is not marginal.

Why it matters: reasoning is what makes the agent conversation credible at all — you cannot delegate a multi-step task to something that cannot plan multiple steps. It also inverts the cost model: harder questions now cost meaningfully more to answer, so "unlimited AI for everyone" stops being a sensible procurement position.

Where you'll meet it: the "thinking" or "extended reasoning" mode in your existing tools. Worth reserving for the questions that justify it.

AI agents & agentic AIWave 4 · loudest, least deployed

In plain EnglishAn AI agent is the thing itself — a single system given a goal and the means to pursue it. Agentic is the property: how much initiative it is permitted. The two words are used interchangeably in the market and should not be. Systems that take actions on their own initiative — planning a sequence, using tools, calling other software, adapting when something fails — rather than only answering a question you asked.

How it actually worksA reasoning model is given a goal, a set of tools it may call, and permission to loop: plan, act, observe the result, re-plan. Autonomy is a dial, not a switch. Nearly every deployment that works in practice keeps a human approval gate on anything that spends money, contacts a customer or changes a record.

Why it matters: here is the gap between the noise and the ledger. Gartner's 2026 CIO survey found only 17% of organisations have actually deployed AI agents, though more than 60% expect to within two years — the steepest expected adoption curve of any emerging technology. Gartner also projects more than 40% of agentic AI projects will be cancelled by the end of 2027, on cost, unclear value and inadequate risk controls. Both numbers are true at once. Plan for the direction; do not budget for the brochure.

Where you'll meet it: in every vendor deck this year. The question that separates real from repackaged: what can it do without asking a human first?

Multiagent systemsWave 4b · emerging

In plain EnglishSeveral specialised agents working together on one process — one gathers data, one drafts, one checks, one executes — each narrow and therefore more reliable than a single agent attempting everything.

How it actually worksAn orchestrator hands work between agents with defined roles and handover rules. It mirrors how you would structure a competent team, which is both the appeal and the difficulty: errors compound across handovers, and diagnosing a failure across four agents is materially harder than across one.

Why it matters: Gartner lists multiagent systems among its top ten strategic technology trends for 2026, and supply chain is one of the first places it is being applied seriously. Treat it as the 2027 conversation you should be able to discuss competently now, not the 2026 purchase.

Where you'll meet it: multi-step back-office processes — order-to-cash, claims handling, tender response assembly.

Physical AI & roboticsWave 5 · the one operations should watch

In plain EnglishAI that senses and acts in the real world — robots, drones, autonomous handling equipment, machines that adapt rather than repeat a fixed program.

How it actually worksIt combines perception (cameras and sensors), a model that decides, and actuators that move — increasingly trained in simulation before touching a real machine. Gartner describes it as AI models combined with IoT sensors, robotics and automation systems to enable real-time sensing, analysis and execution.

Why it matters: Gartner names physical AI a top strategic trend for 2026 and a leading supply chain trend. For manufacturing, warehousing and logistics this is where AI stops being a software line item and starts appearing in the capex plan — with safety, insurance and workforce consequences that arrive alongside it.

Where you'll meet it: autonomous mobile robots in warehouses, drone stock-take, adaptive pick-and-place, inspection robots in hazardous areas.

MultimodalCross-cutting

In plain EnglishOne model that handles several kinds of input together — text, images, audio, video, diagrams — instead of a separate system for each.

How it actually worksDifferent input types are converted into a shared internal representation, so the model can reason across them at once: read a drawing and the specification beside it, or watch a video and answer a question about what happened at 4:12.

Why it matters: this is what makes AI useful in engineering and manufacturing rather than only in marketing. Most industrial knowledge is not text — it is drawings, P&IDs, photographs, inspection footage and scanned annotations. Multimodal capability is the difference between a tool that reads your reports and one that reads your actual work.

Where you'll meet it: drawing interrogation, damage assessment from photographs, video-based safety audit, scanned-document extraction.

The five waves — and where deployment actually stops Attention travels left to right. Deployment does not. The gap between wave 2 and wave 4 is where most 2026 budgets go missing. MATURE 88% STANDARD 17% EMERGING 1 2 3 4 5 Perception Generative Reasoning Agentic Physical AI Proven and unglamorous.Still the clearestindustrial return. Use AI in at least onefunction. Access is nolonger the constraint. Now standard in frontiertools. Used selectively,because it costs more. Have actually deployedagents — against 60%+expecting to by 2028. A top 2026 trend, movinginto the capex planrather than software.
Sourced figures: 88% McKinsey State of AI · 17% and 60% Gartner CIO Survey 2026. Waves 1, 3 and 5 are positioned qualitatively — the shape of the fall-off is the point, not the exact height.
What this means for you


Layer 3 of 3

Applications — where it is actually deployed

Seven terms. This is the only layer that appears in your P&L, and the only one worth arguing about in a budget meeting. A proposal that cannot name which of these it delivers is a proposal for Layer 1 wearing a Layer 3 price tag.

Copilots & assistantsHighest adoption, hardest to measure

In plain EnglishAn AI helper sitting inside software your people already use — drafting, summarising, searching, suggesting. It advises; the human still acts.

How it actually worksA generative model is wired into the application's own context — your document, your inbox, your ticket — so its suggestions are relevant without anyone pasting anything anywhere. Adoption is easy precisely because nothing new has to be learned.

Why it matters: this is where most organisations' AI spend actually lands, and where the measurement problem is worst. Microsoft's 2026 Work Trend Index found only 16% of AI users have genuinely redesigned their workflows around it — the rest are doing the same work slightly faster, which produces a real but modest and hard-to-attribute gain.

Where you'll meet it: Microsoft 365 Copilot, Gemini in Workspace, the assistant in your CRM, ERP and helpdesk.

Enterprise AIA category, not a product

In plain EnglishAI deployed against your own institutional knowledge — contracts, standards, drawings, policies, project history — under your own access controls.

How it actually worksAlmost always retrieval over your document estate, with permissions honoured so a user only ever sees answers drawn from material they were already entitled to read. The hard part is never the model. It is the state of your documents and the accuracy of your permissions.

Why it matters: "enterprise AI" appears on every vendor site and means almost nothing on its own. The useful question is which of your document estates it indexes and who controls where that index lives — which is precisely where on-premise and sovereign deployment become commercial rather than ideological questions.

Where you'll meet it: tender libraries, standards lookup, HR policy assistants, engineering document search.

Workflow automationWhere agents actually earn

In plain EnglishA defined multi-step process handled end to end — intake, extract, validate, route, draft, escalate — with humans holding the exceptions rather than the whole queue.

How it actually worksTraditional automation handles the deterministic steps; AI handles the judgement ones — reading an unstructured email, deciding whether an invoice matches a purchase order, classifying an unclear complaint. The combination outperforms either alone.

Why it matters: this is where the returns are, and where the discipline is. McKinsey's data is consistent on the point: the small group of AI high performers are markedly more likely to have redesigned the workflow rather than layered AI on top of the existing one. Automating a broken process makes it break faster.

Where you'll meet it: accounts payable, order processing, claims, onboarding, tender qualification.

Predictive maintenanceThe industrial classic

In plain EnglishUsing sensor data to predict that a machine is heading for failure, and intervening before it stops — rather than on a fixed calendar or after the breakdown.

How it actually worksVibration, temperature, current and acoustic data are monitored continuously against learned normal behaviour. Deviations that historically preceded failure raise a flag with lead time attached. It needs sensors, historical failure data and someone empowered to act on the alert — the third is the one projects usually miss.

Why it matters: it predates the current AI cycle by years, has a defensible business case, and speaks the language finance already accepts: unplanned downtime avoided, per line, per month. For manufacturing it is frequently the most credible first AI investment precisely because nobody has to believe anything new to approve it.

Where you'll meet it: rotating equipment, compressors, pumps, CNC spindles, HVAC, fleet.

Vision inspectionPerception, applied

In plain EnglishCameras checking quality, completeness or safety compliance in real time, at line speed, on every unit rather than on a sample.

How it actually worksA perception model trained on your defects runs on a small computer at the line — no cloud, no latency, no data leaving the plant. Modern systems learn new defect types from a modest number of examples rather than thousands.

Why it matters: it converts sampling into total inspection, which changes the quality conversation with your customer rather than merely reducing cost. It is also the easiest AI project to prove: run it in parallel with your existing QC for one month and compare escape rates. No belief required.

Where you'll meet it: surface defects, assembly verification, label and print checks, PPE and safety-zone monitoring.

Conversational AINow genuinely different

In plain EnglishCustomer-facing systems that hold an actual conversation — across chat, voice and messaging — and increasingly resolve the issue rather than routing it.

How it actually worksA language model grounded in your policies and product data, connected to the systems that can actually do something: check an order, issue a credit, book a slot. The connection to those systems, not the conversation quality, is what separates a resolution from a well-worded apology.

Why it matters: the old chatbot deserved its reputation; this generation does not. But McKinsey's own framing is the useful caution — when nearly everyone has deployed something, having a bot stops being a differentiator. Customers reward first-contact resolution, not conversational fluency.

Where you'll meet it: support, order status, appointment booking, inbound qualification, and increasingly voice.

Digital twinRehearsal, not advice

In plain EnglishA live virtual copy of a physical thing — a machine, a production line, a building, a supply chain — fed by real sensor data, used to test a change before you commit to it in the real world.

How it actually worksSensors stream the asset's actual state into a model that mirrors it. You then run the proposed change against the twin: a new schedule, a different feed rate, a rerouted supply. What the twin tells you is only as good as the sensor coverage feeding it, which is why this sits on top of your IoT investment rather than replacing it.

Why it matters: it moves AI from giving advice to letting you rehearse, which is a materially different risk conversation. With Gartner naming physical AI a top 2026 trend and the industrial IoT base already in place, the twin is the layer that connects the two — and for capital-intensive operations it is often the easier board approval of the pair.

Where you'll meet it: plant layout and line balancing, energy optimisation, maintenance scheduling, and increasingly as the simulation environment where physical AI is trained before it touches a real machine.

What this means for you

Field Notes · Issue 002 is in production

One long-read a month. Nothing else in your inbox.

Each Field Note is a working page like this one — researched, sourced, interactive, and paired with a two-minute film. Subscribers get it the day it goes live, before it is posted anywhere else.

No pitching. Unsubscribe in one click.


The ring around all three layers

Infrastructure — where it physically runs

Five terms that cut across every layer above. Ignore them and you will sign a contract whose real constraints — cost, latency, jurisdiction — only surface in year two.

Compute, GPUs & data centresThe bill underneath everything

In plain EnglishThe specialised processors and buildings that make all of this possible. GPUs — originally built for graphics — turn out to be ideally suited to the arithmetic that neural networks require.

How it actually worksTraining and inference both run on large clusters of these chips, which consume substantial power and require serious cooling. Availability, not ambition, has become the binding constraint: capacity is being pre-sold years ahead, and power connection timelines now shape what can be built and when.

Why it matters: the four largest US hyperscalers are guiding to roughly $725 billion of 2026 capital expenditure, up around 77% on 2025's ~$410 billion, with analysts projecting beyond a trillion in 2027. Someone has to earn a return on that. It will arrive in your renewal pricing, which is a good reason to fix multi-year terms now rather than later.

Where you'll meet it: price rises framed as "capacity constraints," and lead times on any dedicated deployment.

Cloud, on-premise & sovereign AIThe question procurement asks first

In plain EnglishWhere the model runs and where your data goes while it does. Cloud means someone else's data centre. On-premise means your building. Sovereign means inside your country's legal jurisdiction, under its rules.

How it actually worksCloud gives you the largest models with no capital cost and a usage-based bill. On-premise means an appliance or server you own — higher upfront cost, no per-token meter, and nothing crossing your firewall. Smaller domain models have made the on-premise option genuinely viable for the first time, because you no longer need a frontier-scale model to do useful work on your own documents.

Why it matters: Gartner's 2026 trend list includes both confidential computing and geopatriation — the deliberate repatriation of workloads into controlled jurisdictions. For regulated industries, defence, engineering IP and government work, this has moved from a philosophical preference to a procurement precondition.

Where you'll meet it: the data-residency clause, your DPO's review, and any tender with a national-security or IP-sensitivity dimension.

IoT & edge computingWhere AI meets the plant

In plain EnglishIoT is the network of connected sensors and machines generating data. Edge computing means processing that data where it is created — on the machine, in the plant — rather than shipping it to a data centre first.

How it actually worksA sensor produces a continuous stream; a small model at the edge decides what matters and acts in milliseconds; only exceptions and summaries travel upstream. This is what makes real-time inspection and safety intervention possible at all — a round trip to the cloud is far too slow for a moving production line.

Why it matters: IoT Analytics counted 21.1 billion connected IoT devices at the end of 2025, up 14% year on year, with about 45% of them enterprise connections — heading toward 39 billion by 2030. That installed base is the raw material for industrial AI. If your machines are already instrumented, you are further along than most organisations debating chatbots.

Where you'll meet it: your SCADA and MES data, condition-monitoring sensors, and any AI vendor asking what your machines already log.

MCP — Model Context ProtocolThe wiring that makes agents possible

In plain EnglishAn open standard for connecting an AI model to your tools and data. The metaphor that stuck is USB-C for AI: one universal connector, instead of a different custom cable for every system you own.

How it actually worksBefore it, every connection was hand-built — one adapter for your ERP, another for SharePoint, another for your CRM — and all of it rewritten whenever you changed model. With MCP, each of your systems exposes a small "server" that describes what it can do and what it will permit. Any compliant model can then discover and call it. Change the model, keep the integrations.

Why it matters: released by Anthropic in November 2024 and since adopted by OpenAI, Google DeepMind, Microsoft, IBM and Amazon. In December 2025 it was donated to the Agentic AI Foundation under the Linux Foundation, making it vendor-neutral. One enterprise tracker puts Fortune 500 implementation near 28% inside eighteen months, and the NSA published security design guidance in May 2026 describing it as the de facto standard. A revised specification ships today, 29 July 2026. An agent with no standard way to reach your systems is just a chatbot — this is the difference.

Where you'll meet it: every agent proposal from here on. The question worth asking: does it speak MCP, or is this bespoke wiring we will pay to rebuild the next time we change vendor? And separately — see guardrails — who audits what those connections are permitted to do?

Open vs closed weightsThe fork that decides on-premise

In plain EnglishClosed means you rent access through an API and never possess the model. Open weights means you can download it and run it on your own machines, permanently.

How it actually worksClosed models generally hold the frontier on raw capability, cost nothing upfront, and leave you dependent on a supplier's pricing, availability and terms. Open-weight models you host yourself: higher setup effort, no per-token meter, complete control of where data sits — and the full compliance and security burden lands on you rather than the vendor.

Why it matters: combined with domain-specific models, this is what has made genuinely air-gapped AI practical rather than aspirational. It is also your exit risk. If a supplier trebles pricing or withdraws a model version, a closed deployment has no fallback — which is why it belongs in a contract review, not only a technical one.

Where you'll meet it: any tender with a data-residency clause, any defence or government-adjacent work, and the moment your renewal quote arrives.

What this means for you


The second ring · around all three layers

Governance & control — what keeps it defensible

Six terms, and the most under-discussed part of any AI map. Everything above this point is about what the technology can do. This is the part you are personally accountable for — and the part a regulator, a customer's due-diligence team or your own board will ask about first.

Regulation & the AI ActAlready binding, not coming

In plain EnglishThe rules you must follow when you deploy AI. The EU AI Act is the most consequential, phasing obligations in by risk tier. India's Digital Personal Data Protection Act governs the personal data that feeds these systems. Sector regulators — financial, medical, safety — apply their own rules on top.

How it actually worksObligations scale with the risk of the use case, not the sophistication of the technology. A simple model making hiring or credit decisions carries far heavier duties than an advanced one drafting marketing copy. The EU framework also reaches beyond Europe: if your output touches the EU market, it can apply to you from Pune or Delaware.

Why it matters: most organisations discover their obligations during a customer's due-diligence questionnaire rather than their own review — which is the expensive way to find out. Classifying your AI use cases by risk tier is a half-day exercise that materially changes what you can safely bid for.

Where you'll meet it: enterprise contracts, tender prequalification, your DPO's review. This page is a briefing, not legal advice — take the classification to counsel before you rely on it.

Guardrails, prompt injection & tool poisoningThe risk that arrived with agents

In plain EnglishGuardrails are the checks placed around a model — what it may see, say and do. Prompt injection is when hidden instructions buried inside content the model reads hijack its behaviour. Tool poisoning is when something it is connected to has been tampered with.

How it actually worksA model cannot reliably tell the difference between instructions from you and text it happens to be reading. So a booby-trapped email, invoice, CV or web page can carry commands. While the system only answers questions, that produces a wrong answer. Once it has permissions to act, the same trick produces an action.

Why it matters: this stopped being theoretical in 2026. Security researchers filed more than thirty CVEs against the MCP ecosystem in January and February alone, and real incidents followed — including a cross-tenant data leak and tool-poisoning attacks against widely-used open-source connectors. The NSA issued formal security design guidance in May 2026. Every permission you grant an agent is a permission an attacker may borrow.

Where you'll meet it: the moment anyone proposes giving an agent write access to a real system. The right question is not "is it secure" but "what is the blast radius if this is hijacked once?"

Evals & benchmarksHow you prove it before you sign

In plain EnglishAn eval is a fixed set of your own real cases, with known-correct answers, that you score every model and every version against. A benchmark is the public equivalent — and it is largely marketing.

How it actually worksYou assemble perhaps a hundred genuine examples from your own work — real tender clauses, real invoices, real support tickets — and record what a good answer looks like for each. Then any vendor, model or upgrade is scored against that same set. Suddenly the comparison is arithmetic rather than opinion.

Why it matters: public leaderboard scores tell you almost nothing about performance on your documents, and vendors demonstrate on material chosen to succeed. A hundred-case eval set costs you a week once and then settles every procurement argument you will have for the next three years. It also catches silent degradation when a supplier updates a model underneath you.

Where you'll meet it: build one before your next vendor demonstration, not after. Ask any serious supplier to run against it — the reluctant ones tell you something useful.

Shadow AIAlready happening in your organisation

In plain EnglishStaff using AI tools nobody approved, on personal accounts, usually with company material pasted into them. Sometimes called BYOAI.

How it actually worksSomeone has a deadline, a consumer tool solves it in four minutes, and a customer contract or a draft board paper gets pasted into a system with no logging, no data-processing agreement and no retention control. There is no malice anywhere in that sequence, which is exactly why policy alone does not stop it.

Why it matters: survey work through 2026 consistently finds that most workers sourced their AI tools themselves, and around four in ten say their employer gave them no preparation at all. Prohibition moves the behaviour underground rather than ending it. The organisations that regain control do it by providing a sanctioned tool that is genuinely good enough to use.

Where you'll meet it: in your own team this week. The honest diagnostic is to ask, without consequences attached, what people are already using.

Data readinessThe real project

In plain EnglishWhether your documents and records are in a state AI can actually use. Not whether you have data — everyone has data — but whether it is findable, current, machine-readable and correctly permissioned.

How it actually worksThe obstacles are mundane and universal: scanned drawings with no extractable text, three conflicting versions of the same standard with no way to tell which is live, permissions that stopped reflecting the organisation chart four reorganisations ago. None of it is glamorous and all of it is decisive.

Why it matters: when researchers examine failed AI projects, the causes cluster around data and integration gaps rather than model quality. This is the single most reliable predictor of whether your deployment works — and the line item most often cut from the budget to make the business case look cleaner.

Where you'll meet it: in the first four weeks of every deployment, whether or not you planned for it. Budget it explicitly and it becomes a project. Ignore it and it becomes the reason for a delay.

Decision intelligenceWhere the accountability sits

In plain EnglishDeciding — deliberately, and in writing — which decisions AI may make alone, which it may only recommend, and which stay entirely human. Then recording what was decided and why.

How it actually worksEach decision type is classified by consequence and reversibility. Low-consequence and easily reversed can be automated. High-consequence or irreversible keeps a named human accountable. The classification is written down, reviewed periodically and audited — which is what turns "we use AI responsibly" into something a regulator or a customer can actually inspect.

Why it matters: McKinsey found nearly two-thirds of respondents citing security and risk as the main barrier to scaling agentic AI — well ahead of regulatory uncertainty or technical limits. The constraint on scaling is confidence, not capability. This is the work that produces the confidence.

Where you'll meet it: your AI policy, your risk register, and the board question you would rather have answered before it was asked.

Decision rights — the grid that makes AI defensible Classify by consequence and reversibility, never by how advanced the technology is. A simple model making credit decisions carries heavier duties than a frontier one drafting copy. CONSEQUENCE → EASILY REVERSED IRREVERSIBLE REVIEW AI drafts, a human signs Draft a customer email · Prepare a quotation Propose a schedule change · Summarise a claim Named reviewer, every time. No exceptions by volume. HUMAN ONLY AI does not touch this Issue a credit or refund · Change a safety setpoint Submit a bid · Terminate anything · Sign anything This is the row a regulator will ask about first. AUTOMATE Let it run Categorise an inbound ticket · Suggest a routing Draft an internal summary · Extract fields for review Log everything. Sample the output weekly. RECOMMEND Propose, never execute Post to an internal channel · Archive a closed record Send a routine notification · Write an audit entry Small in isolation. Unwindable only in aggregate. Write the classification down, review it quarterly, and name the accountable human for everything in the upper half. That document is what turns we use AI responsibly into something a customer or a regulator can actually inspect.
The decision-rights grid — the artefact that answers the board question before it is asked. Field Notes 001.
What this means for you


Reality check

Adoption is not impact

Four numbers that belong together. Read individually, each supports a different agenda. Read together, they describe where the market actually stands in July 2026 — and they are the most useful slide you can take into a budget conversation.

88%
of organisations report using AI in at least one business function
McKinsey, State of AI
39%
report any EBIT impact at enterprise level — and most of those say it is under 5%
McKinsey, State of AI
17%
have actually deployed AI agents, against 60%+ who expect to within two years
Gartner CIO Survey 2026
16%
of AI users have genuinely redesigned their workflow around it
Microsoft Work Trend Index 2026

The pattern is consistent across every credible dataset: near-universal adoption, narrow value capture. MIT's Project NANDA study of enterprise deployments found the overwhelming majority producing no measurable P&L impact, and roughly two-thirds of organisations have not yet begun scaling AI beyond individual use cases. The differentiator is not the model anyone chose. It is whether the work itself was redesigned — the small group of high performers are consistently found to have rebuilt the process rather than bolted AI onto it.

Which makes the most actionable finding of 2026 this one: Gallup's Q2 data shows reported productivity gains climbing from 45% among employees using AI narrowly to 90% among those applying it across seven or more different task types. The constraint is no longer access. It is breadth — and breadth is a training and workflow problem, not a procurement one.

What this means for you


Next four quarters

The watch list

Six things worth tracking into 2027, with an honest label on each: whether it is ready to buy, ready to pilot, or still a conversation to be able to hold.

01

Domain-specific models go mainstream

Smaller, cheaper, specialised models that run on your own hardware. Named by Gartner as a top 2026 trend, and the single development that makes on-premise AI commercially sensible rather than merely compliant.

Ready to pilot
02

The agent cancellation wave

Gartner expects more than 40% of agentic AI projects to be cancelled by end-2027. Expect a visible correction in the narrative during 2027, and a buyer's market for anyone who waited.

Plan around it
03

Physical AI reaches the plant floor

AI models fused with IoT sensors, robotics and automation for real-time sensing and execution. Moves AI from the software budget into capex, with safety and insurance implications attached.

Watch closely
04

Geopatriation & confidential computing

Both on Gartner's 2026 trend list. Data residency is hardening from preference into precondition, particularly for regulated sectors and government-adjacent work.

Procurement impact now
05

Inference cost becomes a board metric

With capex heading past a trillion in 2027, the return has to come from somewhere. Usage-based AI spend is already overshooting forecasts by more than a quarter in most enterprises.

Budget for it
06

Breadth-of-use becomes the KPI

The Gallup finding — 45% to 90% productivity gain as task variety widens — reframes the internal measure from licences issued to number of distinct tasks each person applies AI to.

Change your dashboard
What this means for you


Read this before you quote any of it

Three honest caveats

Surveys measure claims, not audits

Adoption figures come from self-reported surveys. When 88% of organisations say they use AI, that is 88% saying so. Nobody inspected the deployments.

The definitions are contested

"Agent" in particular means very different things to different vendors, which is exactly why deployment figures vary so widely between reputable sources.

This page has a date on it

Every number here carries its source and period. Capex guidance in particular has been revised upward repeatedly through 2026. Check before citing.

Sources & data citations (FactSet, McKinsey, Gartner, Gallup, Microsoft, IoT Analytics, MIT, RAND, etc.)... Read More

Field Notes

Get the next one before anyone else

One in-depth, interactive Field Note a month, each with a two-minute film. Written for the people who have to decide.

No pitching. Unsubscribe in one click.