Cubic Pixel
GET IN TOUCH ↗
PRODUCTS → WORK → PLAYGROUND → BLOG → ABOUT → GET IN TOUCH ↗
HOME / BLOG / ARTICLE

What agentic AI actually is, and when you should want it

Everyone is selling "agents". Here is what the word means when it means something: how a model gets run in a loop, how much autonomy you are really buying, and the four kinds of software you are choosing between.

OCTOBER 6, 2026·47 MIN READ
What agentic AI actually is, and when you should want it
What agentic AI actually is — preview

Ask a language model to write you an essay and it does something no human writer does. It starts at the first word and goes straight through to the last, in order, with no outline, no second draft and no delete key. That is still how most people use these systems, and the odd thing is not that the results are patchy. It is that they are as good as they are.

Now ask the same model to work the way you would. Sketch an outline. Look a few things up. Write a rough draft, read it back, find the paragraph that is not pulling its weight, look one more thing up, rewrite. It takes longer and it costs more, and the finished piece is reliably, measurably better. The analogy belongs to Andrew Ng, who opens his Agentic AI course with it, and it is the whole of the idea in one image. Agentic AI is not a smarter model. It is a different way of running one.

The word has since been stretched over almost anything with a chat box. In June 2025 Gartner counted roughly 130 vendors whose products could act on their own in any real sense, among the thousands claiming the label, coined the phrase “agent washing” for the rest, and predicted that more than 40 per cent of agentic AI projects would be cancelled by the end of 2027. So this piece is for the reader who has been handed the word and would like to know what it refers to: what an agentic workflow actually consists of, how much autonomy you are really buying, and when a plain prompt — or plain old code — is the better tool.

It started from the opening lectures of Ng’s course and was set against the engineering guidance on building agents published by Anthropic, the measurements by METR of how long a task today’s agents can finish unaided, three adoption surveys that flatly disagree, and a Berkeley study of why these systems fail. Where the sources conflict, both numbers are shown rather than averaged.

IN BRIEF
  1. “Agentic” describes how a model is run, not what it is. In a loop — plan, act, check, revise — with tools and outside information, instead of answering once.
  2. The loop beat the bigger brain. On a standard coding test, an older model run in an agentic loop (95 per cent) beat a newer model answering in one pass (67 per cent).
  3. Autonomy is a dial, not a switch. The useful question is how much of “what happens next” the model decides at run time rather than the programmer in advance.
  4. The agent is the last resort, not the default. Four kinds of software compete for every task, and the agent wins only when the path cannot be scripted and the result can be checked.
  5. Reliability compounds the wrong way. A step that works 95 per cent of the time, repeated twenty times unchecked, works 36 per cent of the time. Evaluators exist to break that arithmetic.
  6. The frontier is measured in hours of human work — about five hours at 50 per cent reliability at the end of 2025 — with error bars of roughly a factor of two either way, and far shorter when you need to be right four times in five.
  7. When they fail it is rarely the model. In the Berkeley study, four failures in ten traced to how the task and roles were specified and one in five to nobody checking the result.
  8. Adoption depends on who is asking. A consultancy’s survey finds 23 per cent of organisations scaling agents in any function; a vendor’s finds 81 per cent.
PART I

The essay with no delete key

A large language model — the engine inside ChatGPT, Claude, Gemini and the rest — is a system trained on a vast quantity of text to predict what comes next, one token at a time. (A token is a word or a piece of one; “unbelievable” is three of them.) Trained hard enough, on enough text, next-word prediction turns out to produce fluent, mostly sensible writing about almost anything. What it does not produce on its own is any of the things a careful worker does around the writing: planning, checking, looking things up, trying again.

The ordinary way of using a model — a prompt in, an answer out, nothing in between — is what engineers call zero-shot. The model gets one pass. An agentic workflow gives it several. Ng’s definition is deliberately plain: an application built on a language model that executes multiple steps to complete a task. For the essay, the steps might be: draft an outline; decide whether any web research is needed and what to search for; call a search engine; read what came back; write a first draft; read the draft and decide which parts are weak; research those; revise. Somewhere in there you might add a step where the system can ask a human to check a fact before it goes on.

ONE PASS · HOW MOST OF US USE A MODELPromptModelEssay“Write me an essay on X”first word to last, no backspaceone draft, never re-readAGENTIC WORKFLOW · THE SAME MODEL, RUN IN A LOOPOutlineResearchDraftReviewReviseReportHuman check?STILL THIN? GO BACK AND RESEARCH AGAIN
Figure 1. The same model, used two ways. In the top lane it answers once and the arrows only go right. In the bottom lane a review step reads the draft, and an arrow goes back to research and revise. That backward arrow is the entire difference. Based on the opening of Andrew Ng’s Agentic AI course, DeepLearning.AI, 2025.

Notice what changed. Three things, really. First, each step is a smaller, easier job than “write the essay”; a model that is mediocre at the whole is often very good at the parts. Second, the model gets to see its own work. A reviewing step reads the draft with fresh eyes, which is exactly what you do when you come back to something the next morning. Third, and most important, the system gets information from outside itself. The search results, the contents of a file, whether a piece of code ran: these are facts about the world rather than the model’s recollection of it, and feeding them back into the loop is what keeps the loop honest.

The running example in Ng’s course is a research agent. You type a topic (his was how to start a rocket company to compete with SpaceX) and the system plans what to look into, calls a search engine, downloads pages, synthesises and ranks what it found, drafts an outline, hands the draft to a second model acting as an editor, and finally writes a report with an introduction, background and findings. Ng notes, with some relief, that the report concluded this would be a hard company to start. The point is not the rocket. It is that the same shape of loop — plan, gather, draft, check, revise — turns up in legal research, in compliance review, in writing code and in handling a customer’s refund.

If one skill separates people who can build these systems from people who cannot, it is decomposition: taking a task a model does badly in one go and cutting it into steps it does well. Ng calls it tricky, and he is right. It is closer to designing a job than to writing a program.

PART II

Why the loop beats the bigger brain

The most quoted number in this field comes from a letter Ng wrote in March 2024. He pulled together results from several research teams on HumanEval, a standard test of 164 small programming problems, and lined them up. GPT-3.5, answering in one pass, solved 48.1 per cent. GPT-4, a newer and far larger model answering in one pass, solved 67.0 per cent. GPT-3.5 — the older, weaker model — wrapped in an agentic loop that let it write, run and revise its code reached 95.1 per cent.

SHARE OF HUMANEVAL CODING PROBLEMS SOLVED100%75%50%25%48.1%GPT-3.5answering in one pass67.0%GPT-4answering in one pass95.1%GPT-3.5the same model, run in a loop
Figure 2. The jump from one generation of model to the next was 19 points. The jump from running the older model once to running it in a loop was 47. Figures compiled by Andrew Ng from several research teams’ published results, each on its own method; The Batch, March 2024.

Two caveats, both of which Ng would accept. The 95.1 figure is the best result reported by one team on its own method, and HumanEval had been public for three years by then, long enough for its problems to leak into training data. But the direction has been confirmed so many times since, on so many kinds of task, that it is now simply how the field works: the gap between a model answering once and the same model running in a loop is often larger than the gap between one generation of models and the next.

In that letter Ng also set out four design patterns, and they remain the clearest map of the territory. Reflection: the model examines its own output and works out how to improve it. Tool use: the model is given functions it can call, such as a web search, a calculator, a code interpreter or a database, to gather information or take an action. Planning: the model writes out the steps before it starts, and may rewrite them as it goes. Multi-agent collaboration: several model instances, each with a role — a writer and an editor, say, or a coder and a tester — split the work and check each other.

None of this appeared from nowhere. The idea of a model that alternates between thinking out loud and acting — a line of reasoning, then a tool call, then a look at what came back — was written up by Princeton and Google researchers in October 2022 under the name ReAct. The reviewing step got its own paper, Reflexion, in early 2023. That summer the model vendors added function calling to their interfaces, a structured way for a model to say “run this function with these arguments” instead of hoping a programmer would fish it out of prose. In November 2024 Anthropic published the Model Context Protocol, a shared plug that lets any tool describe itself to any model; by December 2025 it had been handed to the Linux Foundation, with OpenAI, Google, Microsoft, Amazon and Block among the backers. Tool use went from a trick to a standard in about three years.

This matters for a reader trying to get their bearings because every product you have heard called an agent — a coding assistant that edits your files and runs your tests, a “deep research” button that returns a cited report, a support bot that can actually issue the refund — is one of these loops with a particular set of tools plugged into it. The branding varies. The anatomy does not.

PART III

Autonomy is a dial, not a switch

Ng is emphatic on one point, and it is the point most arguments about this topic miss: there is no line with “an agent” on one side and “not an agent” on the other. There are degrees. Anthropic’s engineers draw the most useful distinction in their guidance from December 2024. In a workflow, the model and its tools run along paths the programmer wrote in advance. In an agent, the model decides its own path — which tool to call, whether to carry on, when it is done — and keeps control over how the task gets finished. Both are agentic systems. They differ in who holds the steering wheel.

A ladder published by Hugging Face alongside its smolagents library makes the degrees concrete, and it is worth walking up it with a single example. Imagine a system that handles supplier invoices.

HOW MUCH OF “WHAT HAPPENS NEXT” DOES THE MODEL DECIDE?PROGRAMMER DECIDES THE PATHMODEL DECIDES THE PATHProcessorMODEL DECIDESnothingReads the invoice intofields; code does the restRouterMODEL DECIDESwhich branchPicks one of threefixed proceduresTool callerMODEL DECIDESwhich function to runChooses which lookupto run, and with whatMulti-step agentMODEL DECIDESwhether to go onChecks each result anddecides what comes nextMulti-agentMODEL DECIDESwhen to spawn a loopSpins up a specialistloop per odd invoice
Figure 3. Five rungs of autonomy. At the bottom the model only transforms data and code decides everything; at the top one loop can start another. What moves as you climb is who writes the next step of the program: the programmer in advance, or the model at run time. Adapted from the agency levels in Hugging Face’s smolagents guide, 2024.

At the bottom the model is a processor: it reads the scanned invoice and pulls out the supplier, the amount and the date, and ordinary code does everything after that. One rung up it is a router: the model looks at the invoice and chooses which of three fixed procedures it should go down: standard, needs a purchase order match, looks suspicious. Then a tool caller: the model decides which function to run and with what arguments — look up this purchase order, check this supplier’s payment terms. Then a multi-step agent: the model decides, after each result, whether to continue, what to do next and when the job is finished. At the top, one agentic loop can start another: a manager loop that spins up a specialist for each unusual invoice.

What moves as you climb is control flow — the question of what the program does next. At the bottom the programmer wrote every branch in advance; at the top the model is writing the branches at run time. That is the entire trade. More autonomy buys flexibility on tasks you could not have scripted. It costs predictability, because a path chosen at run time is a path nobody tested.

How much autonomy can you actually buy?

The honest answer comes from METR, a non-profit that measures what frontier models can do on their own. Its yardstick, introduced in March 2025, is disarmingly simple: take a set of real software and research tasks, time how long skilled humans need for each one, then ask how long a task — in human time — a model can complete half the time. They call it the 50 per cent time horizon. In early 2025 the best models succeeded almost always on tasks a person finishes in under four minutes, and less than one time in ten on tasks that take a person more than four hours. Claude 3.7 Sonnet, then the frontier, had a horizon of about an hour. Ask instead for four successes in five and the same model’s horizon dropped to roughly a quarter of an hour.

HUMAN-TIME LENGTH OF TASK A MODEL FINISHES HALF THE TIME · LOG SCALE95% CONFIDENCE RANGE1 MIN10 MIN1 HOUR8 HOURSGPT-4Mar 20233.5 minGPT-4 TurboNov 20233.6 minClaude 3.7 SonnetFeb 202560 mino3Apr 20252 h 1 minClaude Opus 4May 20251 h 41 minGPT-5Aug 20253 h 34 minClaude Opus 4.5Nov 20255 h 20 min
Figure 4. What a model can finish half the time, measured in the time a skilled person would need. The dots are METR’s estimates and the whiskers the ranges published with them: for the latest model the honest answer is “somewhere between three hours and twelve”. Point estimates and 95 per cent ranges from METR’s Time Horizon 1.1 release, January 2026; later 2026 models sit higher but at the edge of what the task set can measure.

The trend is the striking part. The horizon has doubled roughly every seven months since 2019, and METR’s revised task set in January 2026 found it doubling closer to every four months since 2023. By the end of 2025 the frontier sat at about five hours of human work at 50 per cent reliability. Extrapolated, that puts month-long tasks within reach around 2030 — a sentence METR itself surrounds with warnings.

Three of those warnings matter for anyone deciding how much autonomy to grant. First, the error bars: the five-hour figure for Claude Opus 4.5 came with a 95 per cent confidence range running from under three hours to about twelve, and METR’s researchers say a factor of two in each direction is normal. Second, the domain: on tasks that involve looking at a screen and operating a computer, horizons were 40 to 100 times shorter, and one informal test put a model’s horizon for making a cup of coffee at about two minutes. Third, and stated bluntly by METR, a horizon of X hours does not mean you can delegate X-hour tasks. Some jobs — anything where a mistake is expensive and hard to spot — need a 98 per cent success rate to be worth automating at all, and the 50 per cent number tells you nothing about that.

So the dial has a sensible setting for every task, and it is rarely the top. The question is not “can the model do this on its own” but “at what success rate, checked by what, and who approves the step that cannot be undone”.

PART IV

The anatomy of an agentic workflow

Strip the marketing away and most agentic systems are built from the same eight or nine parts. Not all of them are present in every system, and several can be played by the same model wearing different hats, but it helps to know their names, because the names are what vendors, frameworks and job adverts use.

GUARDRAILS · STEP LIMIT · SPEND CAP · SANDBOXA human sets the goal“Sort out this invoice”Plannerturns the goal into stepsMemoryshort-term + long-term storeOrchestratorthe loop: chooses the next stepToolssearch · code · databasesHuman checkpointapproves irreversible stepsEvaluatortests · rules · a second readResultreport, fix, refund, replyCALLRESULTDRAFTREVISEPASSNEEDS A SIGN-OFF?
Figure 5. The anatomy, laid out as a loop. A goal comes in from a person, a planner turns it into steps, an orchestrator runs the loop — calling tools, reading and writing memory — and an evaluator checks every draft, sending it back to be revised or on to the result. The dashed frame is the set of limits that stops the loop running away. Synthesised from Ng’s course and Anthropic’s “Building effective agents”.

The model is the reasoning engine, and there may be more than one: a small, cheap model for routing and a large one for planning is a common and sensible split. Everything else exists to feed the model good information and to catch it when it is wrong.

The orchestrator runs the loop. It holds the state of the task, decides what to call next, passes results back in and enforces the limits. In a workflow the orchestrator is your code and the model is a function it calls; in an agent the model takes over much of the orchestrator’s job and your code shrinks to a loop with a step counter. Anthropic names a pattern after it — orchestrator-workers — in which one model breaks a task into sub-tasks it could not have predicted in advance, hands them to worker models and stitches the results together.

The planner turns a goal into a sequence of steps and, in better systems, revises the plan when a step fails. Showing the plan to the user rather than hiding it is one of Anthropic’s three stated principles for agent design; a visible plan is the cheapest form of transparency there is.

Tools are the functions the model is allowed to call: a search engine, a code interpreter, a database query, a file system, a browser, the API of a payment or messaging system. Each comes with a description the model reads to decide when and how to use it, and that description is an interface in the full sense. When Anthropic built its coding agent for SWE-bench — a test made of real bugs from public code repositories — the team reported spending more time on the tools than on the prompt, and one fix was simply to require full file paths because the model kept getting lost in relative ones.

Memory comes in two kinds. Short-term memory is the context window — the finite stretch of text the model can see at once, into which the orchestrator keeps appending results. Long-term memory is anything the system can write to and search later: files, a database, a vector store that finds passages by meaning rather than by exact words. Retrieval, which means fetching the relevant passage and putting it in front of the model before it answers, is the most common augmentation of all, and often the only one a task needs.

The evaluator checks the work. It may be another model call — read this draft, list what is wrong with it — which is the reflection pattern in its simplest form. It is much stronger when it is deterministic: run the tests, validate the output against a schema, check the total against the purchase order. Anthropic’s evaluator-optimiser pattern pairs a generating call with an evaluating one in a loop until the evaluator is satisfied, and recommends it for exactly the tasks where a human reviewer could articulate what is wrong, literary translation and multi-round research among them.

Guardrails and stopping conditions are the parts beginners forget and veterans build first: a maximum number of iterations, a spending cap, an allow-list of tools, a sandbox in which nothing the agent does can touch the real system. Anthropic’s advice is to get “ground truth” from the environment at every step — the actual result of the tool call, not the model’s belief about it — and to test extensively in a sandbox before anything runs for real.

The human sits at the checkpoints. A system can pause for approval before any action that cannot be undone — sending the email, issuing the refund, deleting the file — or ask for help when it is stuck. The industry shorthand distinguishes a human in the loop, who approves each step, from a human on the loop, who watches and intervenes. Which you want depends entirely on the cost of a wrong step.

And finally evaluation — evals, in the jargon — which is not the evaluator inside the loop but the discipline outside it: logging every run, measuring end-to-end success and the success of each component, and finding out which step is the weak one. Ng devotes a full module of his course to it, Anthropic calls measurement the key to the whole enterprise, and it is the part that separates teams who ship agents from teams who demo them.

ComponentWhat it doesWho usually does it
The modelReasons, writes, decidesA language model — often more than one, sized to the job
OrchestratorRuns the loop, holds the state, enforces the limitsYour code in a workflow; mostly the model in an agent
PlannerTurns a goal into steps; revises them when a step failsThe model
ToolsAct on the world: search, run code, query, sendFunctions you write and describe to the model
MemoryShort-term context; long-term store and retrievalCode and a database; the model chooses what to keep
EvaluatorChecks the work before it moves onTests and rules where possible; a model as a reviewer where not
Guardrails and stop conditionsCaps on steps, spend and permissions; a sandboxCode, set before the first run
Human checkpointApproves irreversible steps; answers when the system is stuckA person
EvalsMeasures success per run and per componentYou, continuously

Six shapes the loop takes

Anthropic’s guidance catalogues the workflow patterns its customers actually use in production. They are composable; real systems combine two or three, and the last row is the one most people mean when they say “agent”.

PatternWhat it isFits when
Prompt chainingA fixed sequence: each call works on the last call’s output, with checks betweenThe steps are known and the task decomposes cleanly
RoutingA classifier sends each input to a specialised prompt or a cheaper modelInputs fall into distinct categories handled differently
ParallelisationIndependent sub-tasks run at once, or the same task runs several times and votesSpeed, or confidence from several independent opinions
Orchestrator-workersOne model breaks the task into sub-tasks it could not predict, delegates, synthesisesCode changes across many files; research across many sources
Evaluator-optimiserA generator and a critic loop until the critic is satisfiedClear criteria, and a task that improves with feedback
Autonomous agentThe model plans, acts with tools and judges its own progress in an open loopThe path cannot be scripted, results can be verified, failures are affordable
PART V

Four kinds of software, one job

Here is the question that matters most in practice, and it is better asked as a ladder of its own. For any task you want to automate, four kinds of software compete for it, and the right one is the lowest rung that works.

Ordinary code. If the rule can be written down — calculate the VAT, sort these by date, reject anything over the credit limit — write the rule. Code is fast, costs nothing per run, gives the same answer every time, can be tested exhaustively and can be audited line by line. No model matches any of those properties, and a surprising share of what gets pitched as “agentic” is a rule somebody did not want to write. Gartner’s analyst put it plainly: many of the use cases being sold as agentic do not need an agent.

A single model call. If the task is one transformation of language — summarise this, classify that, extract these fields, translate, draft a reply — a single call, perhaps with retrieved documents and a few worked examples in the prompt, is usually enough. Anthropic’s guidance says so explicitly: for many applications, optimising one call with retrieval and examples is all that is required, and the right move is to find the simplest solution and add complexity only when measurement demands it.

A workflow. If the task has several steps but you know what they are — extract, validate, look up, decide, write — build the sequence in code and use the model only for the steps code cannot do. This is the predictable middle: the model’s judgement where you need it, and a fixed, testable path everywhere else. Most production “agents” are, on inspection, workflows, and that is a compliment.

An agent. If the path really cannot be predicted in advance — the number of steps depends on what turns up, the next action depends on the last result, no fixed sequence covers the cases — then let the model choose. But Anthropic attaches conditions, and they are the whole decision: you must be able to check the result, you must be able to afford the failures, and you must trust the model’s judgement over many turns. The two domains where its customers found agents paid off, customer support and coding, share one property: success is verifiable. The ticket is resolved or it is not; the tests pass or they do not.

Can you write the rule down?Ordinary codefast, cheap, testable, the same every timeYESNODoes one read of the inputgive the answer?One model callwith retrieved documents and a few examplesYESNODo you know the steps in advance?A workflowcode runs the steps; the model does the fuzzy onesYESNOCan you check the result —and afford the failures?An agentthe model chooses the path; you watch the costYESNONot yetnarrow the task until you can check it
Figure 6. The decision, as a ladder. Each “no” moves you down to a more expensive, less predictable kind of software; the agent is the last rung, and it has conditions attached. Derived from Anthropic’s “Building effective agents” (December 2024) and Gartner’s 2025 guidance.

Run the invoice through all four. Totals, tax and due dates are code. Reading a scanned invoice into fields is one model call. Matching it to a purchase order, flagging exceptions and routing approvals is a workflow, with a model only at the fuzzy steps. Chasing down why one supplier’s invoices never match — reading the contract, querying three systems, drafting the email to procurement — is the one piece that might justify an agent, and even that piece ends with a human pressing send.

A cost argument belongs here too. A model is charged by the token, and an agent re-reads its entire history on every step, so the cost of a run grows faster than the number of steps. A ten-step agent is not ten times the price of one call; it is considerably more, and slower, because the steps happen one after another. Anthropic’s phrasing is that agentic systems trade latency and cost for task performance, and the trade only makes sense when the performance was not available any other way.

PART VI

The arithmetic that keeps agents honest

There is a piece of arithmetic every agent builder learns the hard way, and it is worth learning the easy way instead. Suppose each step in a chain succeeds 95 per cent of the time, which is a good day for a model. Over five steps the whole chain succeeds 77 per cent of the time. Over ten, 60 per cent. Over twenty, 36 per cent. At 99 per cent a step — which almost nothing achieves — twenty steps still lose nearly one task in five. Reliability engineers have called this Lusser’s law since the 1940s; in 2025 Utkarsh Kanwat, an engineer who builds these systems for a living, wrote the essay that made it the field’s favourite cold shower.

CHANCE THE WHOLE TASK SUCCEEDS, BY NUMBER OF UNCHECKED STEPS100%75%50%25%82%36% after 20 steps12%99% PER STEP95% PER STEP90% PER STEP110304020 STEPS50 STEPS
Figure 7. Multiply a per-step success rate by itself and it falls fast. The middle line is the one to remember: a step that works 95 times in 100 gives a task that works 36 times in 100 by step twenty, if nothing in between is checked. Simple multiplication, no data behind it beyond the per-step rates on the labels.

The maths is correct and incomplete, and the gap is exactly where good agents live. The multiplication assumes each step is a coin toss nobody looks at. The whole purpose of the evaluator is to look: a step that is checked and retried is no longer an independent failure, and a loop with a real verifier can turn 95 per cent components into a reliable system. So the arithmetic is not an argument against agents. It is an argument against agents without verification, and for keeping the number of unchecked steps as small as the task allows.

What verification looks like when it fails is documented in a study from UC Berkeley, published in 2025 under the title “Why do multi-agent LLM systems fail?”. The team went through more than 200 full traces from seven open-source agent frameworks — the first 150-odd by hand, the rest with a model taught to apply the same labels — and classified every failure. Fourteen distinct modes emerged, in three families.

WHERE 200+ MULTI-AGENT RUNS WENT WRONG · SHARE OF FAILURES42%SPECIFICATIONThe task or roles were badly written: aconstraint ignored, steps repeated, thethread lost, no idea it had finished37%AGENTS TALKING PAST EACH OTHERActing on a wrong assumption instead ofasking; withholding what another agentneeded; reasoning one way, acting another21%VERIFICATIONStopped early, checkednothing, or checked thewrong thing
Figure 8. Where the failures came from. Four in ten were in the specification — the task or the roles as written — and only one in five in verification, which is where most people assume the problem lies. Shares from Cemri et al., “Why Do Multi-Agent LLM Systems Fail?”, UC Berkeley, April 2025 revision, 200+ annotated traces.

Four failures in ten traced to specification — how the task and the roles were written: an agent ignoring a constraint it was given, repeating steps it had already done, losing track of the conversation, or not knowing it had finished. Nearly four in ten were agents talking past each other: proceeding on a wrong assumption instead of asking, withholding information another agent needed, reasoning one way and acting another. One in five was verification: stopping early, not checking at all, or checking the wrong thing. Their sharpest example is a system that produced a chess program which passed every review stage and then accepted illegal moves, because the reviewer checked that the code compiled and had comments, not that it played chess. Another retrieved ten songs from a playlist one at a time over ten rounds of conversation when a single call would have done, which is the kind of inefficiency that quietly multiplies a bill by ten.

Two of their findings are the ones to carry away. Systems with a dedicated verifier failed less often than systems without — but adding one extra check against the high-level goal, rather than the low-level code, improved one framework’s success by 15.6 points, which tells you how shallow the existing checks were. And their overall conclusion was that most of what went wrong was organisational design, not model intelligence: better models would not have fixed it. An organisation of competent people with a bad structure fails, and so does an organisation of competent models.

PART VII

Who is actually using this

Ask how many organisations are running agents in earnest and you get a number that depends almost entirely on who asked.

SHARE OF ORGANISATIONS SAID TO BE PAST THE PILOT STAGE, BY WHO ASKED0%25%50%75%100%McKinsey · mid-20251,993 respondents, 105 countriesscaling agents in at least one function23%Gartner · Jan 20253,412 webinar attendeesmaking a “significant investment”19%CrewAI · Feb 2026the vendor’s own survey · 500 executivesadoption scaled or actively expanding81%
Figure 9. Three surveys, one question, three answers. The two independent ones agree with each other and disagree with the vendor’s by a factor of four. Sources: McKinsey, The State of AI 2025 (survey June–July 2025); Gartner webinar poll, January 2025; CrewAI, 2026 State of Agentic AI, February 2026 — a survey run by a company that sells agent software.

The annual survey from McKinsey, run in the middle of 2025 across 1,993 respondents in 105 countries, found 62 per cent of organisations at least experimenting with agents and 23 per cent scaling them in at least one business function — and in no single function did more than one organisation in ten report scaling. Gartner’s January 2025 poll of 3,412 webinar attendees found 19 per cent making significant investments, 42 per cent investing cautiously and 31 per cent waiting to see. CrewAI, a company that sells an agent platform, surveyed 500 senior executives at large enterprises in early 2026 and reported that 65 per cent were already using agents, 81 per cent had scaled or were actively expanding them, and the average organisation had automated 31 per cent of its workflows.

The numbers are not reconcilable and were never meant to be. They ask different questions of different people — a webinar audience self-selects for interest, executives report ambition, a vendor’s survey frames adoption generously — and “using agents” can mean a scheduled summary email or a system that approves payments. Read together they say something more useful than any one of them: curiosity is nearly universal, production use in a single function is common, and production use across a business is still rare. Gartner’s forecast that more than 40 per cent of projects will be scrapped by the end of 2027 sits comfortably beside its own forecast that a third of enterprise applications will contain agentic features by 2028. Both can be true. A lot of what gets built will be the wrong rung of the ladder.

What this means if you are the one deciding

  1. Write the rule first. If you can state the logic, code it. The model is for the parts you cannot state.
  2. Start with one call, and measure it. Add the loop only when the measurement says the single call is not good enough — and keep measuring once you add the loop, because it is now the only way to know which step is failing.
  3. Choose the rung on purpose. Decide how much of the control flow you are handing to the model, write it down, and make everything else deterministic.
  4. Build the evaluator before the planner. A system that can check its own work can afford to plan badly. A system that plans beautifully and never checks will fail quietly.
  5. Count the steps. Every unchecked step you remove is a multiplicative gain in reliability; every one you add has to pay for itself.
  6. Put a human before anything irreversible, cap the iterations and the spend, and run it in a sandbox until you have watched it fail.
  7. Pick tasks where success is checkable. Tests pass, tickets close, totals match. The domains where agents have worked all share this; the ones where they have not, mostly do not.

What would change this picture

Methods and limits

The HumanEval comparison is a synthesis Ng assembled from several teams’ published results, each on its own method, on a test old enough to have leaked into training data; it shows a direction, not a precise effect. METR’s horizons are point estimates with wide, stated confidence ranges, measured on software and research tasks that may not resemble yours, and its authors are the first to say so; the figure here uses the January 2026 revision, and METR has since revised some estimates again as it adjusted its model. The Berkeley failure study examined open-source frameworks running 2024-era models, and its category shares shifted between versions of the paper as more traces were annotated; the shares here are from the April 2025 revision. The three adoption surveys measure different things of different populations and are shown together to make that visible, not to be averaged. Gartner’s figures are forecasts and a poll of a self-selected audience. Nothing here is my own measurement. Where I give an opinion — on which rung to choose, on building the evaluator first — it is the opinion of someone who has sat in the rooms where these projects get approved, and it is offered as such.

Sources

Ghassan
WRITTEN BY
Ghassan

Designer & builder behind Cubic Pixel. I write about digital systems, product craft, and what I learn along the way.

ABOUT ME →

Comments

NO COMMENTS YET
Be kind — comments are moderated.
✓
Thanks — your comment is awaiting moderation.
No comments yet.

Be the first to share your thoughts — I read and reply to every comment.

NEXT → The Open Model Economy: A Mid-2026 Market Study
← BACK TO THE BLOG