Ask a language model to write you an essay and it does something no human writer does. It starts at the first word and goes straight through to the last, in order, with no outline, no second draft and no delete key. That is still how most people use these systems, and the odd thing is not that the results are patchy. It is that they are as good as they are.
Now ask the same model to work the way you would. Sketch an outline. Look a few things up. Write a rough draft, read it back, find the paragraph that is not pulling its weight, look one more thing up, rewrite. It takes longer and it costs more, and the finished piece is reliably, measurably better. The analogy belongs to Andrew Ng, who opens his Agentic AI course with it, and it is the whole of the idea in one image. Agentic AI is not a smarter model. It is a different way of running one.
The word has since been stretched over almost anything with a chat box. In June 2025 Gartner counted roughly 130 vendors whose products could act on their own in any real sense, among the thousands claiming the label, coined the phrase “agent washing” for the rest, and predicted that more than 40 per cent of agentic AI projects would be cancelled by the end of 2027. So this piece is for the reader who has been handed the word and would like to know what it refers to: what an agentic workflow actually consists of, how much autonomy you are really buying, and when a plain prompt — or plain old code — is the better tool.
It started from the opening lectures of Ng’s course and was set against the engineering guidance on building agents published by Anthropic, the measurements by METR of how long a task today’s agents can finish unaided, three adoption surveys that flatly disagree, and a Berkeley study of why these systems fail. Where the sources conflict, both numbers are shown rather than averaged.
- “Agentic” describes how a model is run, not what it is. In a loop — plan, act, check, revise — with tools and outside information, instead of answering once.
- The loop beat the bigger brain. On a standard coding test, an older model run in an agentic loop (95 per cent) beat a newer model answering in one pass (67 per cent).
- Autonomy is a dial, not a switch. The useful question is how much of “what happens next” the model decides at run time rather than the programmer in advance.
- The agent is the last resort, not the default. Four kinds of software compete for every task, and the agent wins only when the path cannot be scripted and the result can be checked.
- Reliability compounds the wrong way. A step that works 95 per cent of the time, repeated twenty times unchecked, works 36 per cent of the time. Evaluators exist to break that arithmetic.
- The frontier is measured in hours of human work — about five hours at 50 per cent reliability at the end of 2025 — with error bars of roughly a factor of two either way, and far shorter when you need to be right four times in five.
- When they fail it is rarely the model. In the Berkeley study, four failures in ten traced to how the task and roles were specified and one in five to nobody checking the result.
- Adoption depends on who is asking. A consultancy’s survey finds 23 per cent of organisations scaling agents in any function; a vendor’s finds 81 per cent.
The essay with no delete key
A large language model — the engine inside ChatGPT, Claude, Gemini and the rest — is a system trained on a vast quantity of text to predict what comes next, one token at a time. (A token is a word or a piece of one; “unbelievable” is three of them.) Trained hard enough, on enough text, next-word prediction turns out to produce fluent, mostly sensible writing about almost anything. What it does not produce on its own is any of the things a careful worker does around the writing: planning, checking, looking things up, trying again.
The ordinary way of using a model — a prompt in, an answer out, nothing in between — is what engineers call zero-shot. The model gets one pass. An agentic workflow gives it several. Ng’s definition is deliberately plain: an application built on a language model that executes multiple steps to complete a task. For the essay, the steps might be: draft an outline; decide whether any web research is needed and what to search for; call a search engine; read what came back; write a first draft; read the draft and decide which parts are weak; research those; revise. Somewhere in there you might add a step where the system can ask a human to check a fact before it goes on.
Notice what changed. Three things, really. First, each step is a smaller, easier job than “write the essay”; a model that is mediocre at the whole is often very good at the parts. Second, the model gets to see its own work. A reviewing step reads the draft with fresh eyes, which is exactly what you do when you come back to something the next morning. Third, and most important, the system gets information from outside itself. The search results, the contents of a file, whether a piece of code ran: these are facts about the world rather than the model’s recollection of it, and feeding them back into the loop is what keeps the loop honest.
The running example in Ng’s course is a research agent. You type a topic (his was how to start a rocket company to compete with SpaceX) and the system plans what to look into, calls a search engine, downloads pages, synthesises and ranks what it found, drafts an outline, hands the draft to a second model acting as an editor, and finally writes a report with an introduction, background and findings. Ng notes, with some relief, that the report concluded this would be a hard company to start. The point is not the rocket. It is that the same shape of loop — plan, gather, draft, check, revise — turns up in legal research, in compliance review, in writing code and in handling a customer’s refund.
If one skill separates people who can build these systems from people who cannot, it is decomposition: taking a task a model does badly in one go and cutting it into steps it does well. Ng calls it tricky, and he is right. It is closer to designing a job than to writing a program.
Why the loop beats the bigger brain
The most quoted number in this field comes from a letter Ng wrote in March 2024. He pulled together results from several research teams on HumanEval, a standard test of 164 small programming problems, and lined them up. GPT-3.5, answering in one pass, solved 48.1 per cent. GPT-4, a newer and far larger model answering in one pass, solved 67.0 per cent. GPT-3.5 — the older, weaker model — wrapped in an agentic loop that let it write, run and revise its code reached 95.1 per cent.
Two caveats, both of which Ng would accept. The 95.1 figure is the best result reported by one team on its own method, and HumanEval had been public for three years by then, long enough for its problems to leak into training data. But the direction has been confirmed so many times since, on so many kinds of task, that it is now simply how the field works: the gap between a model answering once and the same model running in a loop is often larger than the gap between one generation of models and the next.
In that letter Ng also set out four design patterns, and they remain the clearest map of the territory. Reflection: the model examines its own output and works out how to improve it. Tool use: the model is given functions it can call, such as a web search, a calculator, a code interpreter or a database, to gather information or take an action. Planning: the model writes out the steps before it starts, and may rewrite them as it goes. Multi-agent collaboration: several model instances, each with a role — a writer and an editor, say, or a coder and a tester — split the work and check each other.
None of this appeared from nowhere. The idea of a model that alternates between thinking out loud and acting — a line of reasoning, then a tool call, then a look at what came back — was written up by Princeton and Google researchers in October 2022 under the name ReAct. The reviewing step got its own paper, Reflexion, in early 2023. That summer the model vendors added function calling to their interfaces, a structured way for a model to say “run this function with these arguments” instead of hoping a programmer would fish it out of prose. In November 2024 Anthropic published the Model Context Protocol, a shared plug that lets any tool describe itself to any model; by December 2025 it had been handed to the Linux Foundation, with OpenAI, Google, Microsoft, Amazon and Block among the backers. Tool use went from a trick to a standard in about three years.
This matters for a reader trying to get their bearings because every product you have heard called an agent — a coding assistant that edits your files and runs your tests, a “deep research” button that returns a cited report, a support bot that can actually issue the refund — is one of these loops with a particular set of tools plugged into it. The branding varies. The anatomy does not.
Autonomy is a dial, not a switch
Ng is emphatic on one point, and it is the point most arguments about this topic miss: there is no line with “an agent” on one side and “not an agent” on the other. There are degrees. Anthropic’s engineers draw the most useful distinction in their guidance from December 2024. In a workflow, the model and its tools run along paths the programmer wrote in advance. In an agent, the model decides its own path — which tool to call, whether to carry on, when it is done — and keeps control over how the task gets finished. Both are agentic systems. They differ in who holds the steering wheel.
A ladder published by Hugging Face alongside its smolagents library makes the degrees concrete, and it is worth walking up it with a single example. Imagine a system that handles supplier invoices.
At the bottom the model is a processor: it reads the scanned invoice and pulls out the supplier, the amount and the date, and ordinary code does everything after that. One rung up it is a router: the model looks at the invoice and chooses which of three fixed procedures it should go down: standard, needs a purchase order match, looks suspicious. Then a tool caller: the model decides which function to run and with what arguments — look up this purchase order, check this supplier’s payment terms. Then a multi-step agent: the model decides, after each result, whether to continue, what to do next and when the job is finished. At the top, one agentic loop can start another: a manager loop that spins up a specialist for each unusual invoice.
What moves as you climb is control flow — the question of what the program does next. At the bottom the programmer wrote every branch in advance; at the top the model is writing the branches at run time. That is the entire trade. More autonomy buys flexibility on tasks you could not have scripted. It costs predictability, because a path chosen at run time is a path nobody tested.
How much autonomy can you actually buy?
The honest answer comes from METR, a non-profit that measures what frontier models can do on their own. Its yardstick, introduced in March 2025, is disarmingly simple: take a set of real software and research tasks, time how long skilled humans need for each one, then ask how long a task — in human time — a model can complete half the time. They call it the 50 per cent time horizon. In early 2025 the best models succeeded almost always on tasks a person finishes in under four minutes, and less than one time in ten on tasks that take a person more than four hours. Claude 3.7 Sonnet, then the frontier, had a horizon of about an hour. Ask instead for four successes in five and the same model’s horizon dropped to roughly a quarter of an hour.
The trend is the striking part. The horizon has doubled roughly every seven months since 2019, and METR’s revised task set in January 2026 found it doubling closer to every four months since 2023. By the end of 2025 the frontier sat at about five hours of human work at 50 per cent reliability. Extrapolated, that puts month-long tasks within reach around 2030 — a sentence METR itself surrounds with warnings.
Three of those warnings matter for anyone deciding how much autonomy to grant. First, the error bars: the five-hour figure for Claude Opus 4.5 came with a 95 per cent confidence range running from under three hours to about twelve, and METR’s researchers say a factor of two in each direction is normal. Second, the domain: on tasks that involve looking at a screen and operating a computer, horizons were 40 to 100 times shorter, and one informal test put a model’s horizon for making a cup of coffee at about two minutes. Third, and stated bluntly by METR, a horizon of X hours does not mean you can delegate X-hour tasks. Some jobs — anything where a mistake is expensive and hard to spot — need a 98 per cent success rate to be worth automating at all, and the 50 per cent number tells you nothing about that.
So the dial has a sensible setting for every task, and it is rarely the top. The question is not “can the model do this on its own” but “at what success rate, checked by what, and who approves the step that cannot be undone”.
The anatomy of an agentic workflow
Strip the marketing away and most agentic systems are built from the same eight or nine parts. Not all of them are present in every system, and several can be played by the same model wearing different hats, but it helps to know their names, because the names are what vendors, frameworks and job adverts use.
The model is the reasoning engine, and there may be more than one: a small, cheap model for routing and a large one for planning is a common and sensible split. Everything else exists to feed the model good information and to catch it when it is wrong.
The orchestrator runs the loop. It holds the state of the task, decides what to call next, passes results back in and enforces the limits. In a workflow the orchestrator is your code and the model is a function it calls; in an agent the model takes over much of the orchestrator’s job and your code shrinks to a loop with a step counter. Anthropic names a pattern after it — orchestrator-workers — in which one model breaks a task into sub-tasks it could not have predicted in advance, hands them to worker models and stitches the results together.
The planner turns a goal into a sequence of steps and, in better systems, revises the plan when a step fails. Showing the plan to the user rather than hiding it is one of Anthropic’s three stated principles for agent design; a visible plan is the cheapest form of transparency there is.
Tools are the functions the model is allowed to call: a search engine, a code interpreter, a database query, a file system, a browser, the API of a payment or messaging system. Each comes with a description the model reads to decide when and how to use it, and that description is an interface in the full sense. When Anthropic built its coding agent for SWE-bench — a test made of real bugs from public code repositories — the team reported spending more time on the tools than on the prompt, and one fix was simply to require full file paths because the model kept getting lost in relative ones.
Memory comes in two kinds. Short-term memory is the context window — the finite stretch of text the model can see at once, into which the orchestrator keeps appending results. Long-term memory is anything the system can write to and search later: files, a database, a vector store that finds passages by meaning rather than by exact words. Retrieval, which means fetching the relevant passage and putting it in front of the model before it answers, is the most common augmentation of all, and often the only one a task needs.
The evaluator checks the work. It may be another model call — read this draft, list what is wrong with it — which is the reflection pattern in its simplest form. It is much stronger when it is deterministic: run the tests, validate the output against a schema, check the total against the purchase order. Anthropic’s evaluator-optimiser pattern pairs a generating call with an evaluating one in a loop until the evaluator is satisfied, and recommends it for exactly the tasks where a human reviewer could articulate what is wrong, literary translation and multi-round research among them.
Guardrails and stopping conditions are the parts beginners forget and veterans build first: a maximum number of iterations, a spending cap, an allow-list of tools, a sandbox in which nothing the agent does can touch the real system. Anthropic’s advice is to get “ground truth” from the environment at every step — the actual result of the tool call, not the model’s belief about it — and to test extensively in a sandbox before anything runs for real.
The human sits at the checkpoints. A system can pause for approval before any action that cannot be undone — sending the email, issuing the refund, deleting the file — or ask for help when it is stuck. The industry shorthand distinguishes a human in the loop, who approves each step, from a human on the loop, who watches and intervenes. Which you want depends entirely on the cost of a wrong step.
And finally evaluation — evals, in the jargon — which is not the evaluator inside the loop but the discipline outside it: logging every run, measuring end-to-end success and the success of each component, and finding out which step is the weak one. Ng devotes a full module of his course to it, Anthropic calls measurement the key to the whole enterprise, and it is the part that separates teams who ship agents from teams who demo them.
| Component | What it does | Who usually does it |
|---|---|---|
| The model | Reasons, writes, decides | A language model — often more than one, sized to the job |
| Orchestrator | Runs the loop, holds the state, enforces the limits | Your code in a workflow; mostly the model in an agent |
| Planner | Turns a goal into steps; revises them when a step fails | The model |
| Tools | Act on the world: search, run code, query, send | Functions you write and describe to the model |
| Memory | Short-term context; long-term store and retrieval | Code and a database; the model chooses what to keep |
| Evaluator | Checks the work before it moves on | Tests and rules where possible; a model as a reviewer where not |
| Guardrails and stop conditions | Caps on steps, spend and permissions; a sandbox | Code, set before the first run |
| Human checkpoint | Approves irreversible steps; answers when the system is stuck | A person |
| Evals | Measures success per run and per component | You, continuously |
Six shapes the loop takes
Anthropic’s guidance catalogues the workflow patterns its customers actually use in production. They are composable; real systems combine two or three, and the last row is the one most people mean when they say “agent”.
| Pattern | What it is | Fits when |
|---|---|---|
| Prompt chaining | A fixed sequence: each call works on the last call’s output, with checks between | The steps are known and the task decomposes cleanly |
| Routing | A classifier sends each input to a specialised prompt or a cheaper model | Inputs fall into distinct categories handled differently |
| Parallelisation | Independent sub-tasks run at once, or the same task runs several times and votes | Speed, or confidence from several independent opinions |
| Orchestrator-workers | One model breaks the task into sub-tasks it could not predict, delegates, synthesises | Code changes across many files; research across many sources |
| Evaluator-optimiser | A generator and a critic loop until the critic is satisfied | Clear criteria, and a task that improves with feedback |
| Autonomous agent | The model plans, acts with tools and judges its own progress in an open loop | The path cannot be scripted, results can be verified, failures are affordable |
Four kinds of software, one job
Here is the question that matters most in practice, and it is better asked as a ladder of its own. For any task you want to automate, four kinds of software compete for it, and the right one is the lowest rung that works.
Ordinary code. If the rule can be written down — calculate the VAT, sort these by date, reject anything over the credit limit — write the rule. Code is fast, costs nothing per run, gives the same answer every time, can be tested exhaustively and can be audited line by line. No model matches any of those properties, and a surprising share of what gets pitched as “agentic” is a rule somebody did not want to write. Gartner’s analyst put it plainly: many of the use cases being sold as agentic do not need an agent.
A single model call. If the task is one transformation of language — summarise this, classify that, extract these fields, translate, draft a reply — a single call, perhaps with retrieved documents and a few worked examples in the prompt, is usually enough. Anthropic’s guidance says so explicitly: for many applications, optimising one call with retrieval and examples is all that is required, and the right move is to find the simplest solution and add complexity only when measurement demands it.
A workflow. If the task has several steps but you know what they are — extract, validate, look up, decide, write — build the sequence in code and use the model only for the steps code cannot do. This is the predictable middle: the model’s judgement where you need it, and a fixed, testable path everywhere else. Most production “agents” are, on inspection, workflows, and that is a compliment.
An agent. If the path really cannot be predicted in advance — the number of steps depends on what turns up, the next action depends on the last result, no fixed sequence covers the cases — then let the model choose. But Anthropic attaches conditions, and they are the whole decision: you must be able to check the result, you must be able to afford the failures, and you must trust the model’s judgement over many turns. The two domains where its customers found agents paid off, customer support and coding, share one property: success is verifiable. The ticket is resolved or it is not; the tests pass or they do not.
Run the invoice through all four. Totals, tax and due dates are code. Reading a scanned invoice into fields is one model call. Matching it to a purchase order, flagging exceptions and routing approvals is a workflow, with a model only at the fuzzy steps. Chasing down why one supplier’s invoices never match — reading the contract, querying three systems, drafting the email to procurement — is the one piece that might justify an agent, and even that piece ends with a human pressing send.
A cost argument belongs here too. A model is charged by the token, and an agent re-reads its entire history on every step, so the cost of a run grows faster than the number of steps. A ten-step agent is not ten times the price of one call; it is considerably more, and slower, because the steps happen one after another. Anthropic’s phrasing is that agentic systems trade latency and cost for task performance, and the trade only makes sense when the performance was not available any other way.
The arithmetic that keeps agents honest
There is a piece of arithmetic every agent builder learns the hard way, and it is worth learning the easy way instead. Suppose each step in a chain succeeds 95 per cent of the time, which is a good day for a model. Over five steps the whole chain succeeds 77 per cent of the time. Over ten, 60 per cent. Over twenty, 36 per cent. At 99 per cent a step — which almost nothing achieves — twenty steps still lose nearly one task in five. Reliability engineers have called this Lusser’s law since the 1940s; in 2025 Utkarsh Kanwat, an engineer who builds these systems for a living, wrote the essay that made it the field’s favourite cold shower.
The maths is correct and incomplete, and the gap is exactly where good agents live. The multiplication assumes each step is a coin toss nobody looks at. The whole purpose of the evaluator is to look: a step that is checked and retried is no longer an independent failure, and a loop with a real verifier can turn 95 per cent components into a reliable system. So the arithmetic is not an argument against agents. It is an argument against agents without verification, and for keeping the number of unchecked steps as small as the task allows.
What verification looks like when it fails is documented in a study from UC Berkeley, published in 2025 under the title “Why do multi-agent LLM systems fail?”. The team went through more than 200 full traces from seven open-source agent frameworks — the first 150-odd by hand, the rest with a model taught to apply the same labels — and classified every failure. Fourteen distinct modes emerged, in three families.
Four failures in ten traced to specification — how the task and the roles were written: an agent ignoring a constraint it was given, repeating steps it had already done, losing track of the conversation, or not knowing it had finished. Nearly four in ten were agents talking past each other: proceeding on a wrong assumption instead of asking, withholding information another agent needed, reasoning one way and acting another. One in five was verification: stopping early, not checking at all, or checking the wrong thing. Their sharpest example is a system that produced a chess program which passed every review stage and then accepted illegal moves, because the reviewer checked that the code compiled and had comments, not that it played chess. Another retrieved ten songs from a playlist one at a time over ten rounds of conversation when a single call would have done, which is the kind of inefficiency that quietly multiplies a bill by ten.
Two of their findings are the ones to carry away. Systems with a dedicated verifier failed less often than systems without — but adding one extra check against the high-level goal, rather than the low-level code, improved one framework’s success by 15.6 points, which tells you how shallow the existing checks were. And their overall conclusion was that most of what went wrong was organisational design, not model intelligence: better models would not have fixed it. An organisation of competent people with a bad structure fails, and so does an organisation of competent models.
Who is actually using this
Ask how many organisations are running agents in earnest and you get a number that depends almost entirely on who asked.
The annual survey from McKinsey, run in the middle of 2025 across 1,993 respondents in 105 countries, found 62 per cent of organisations at least experimenting with agents and 23 per cent scaling them in at least one business function — and in no single function did more than one organisation in ten report scaling. Gartner’s January 2025 poll of 3,412 webinar attendees found 19 per cent making significant investments, 42 per cent investing cautiously and 31 per cent waiting to see. CrewAI, a company that sells an agent platform, surveyed 500 senior executives at large enterprises in early 2026 and reported that 65 per cent were already using agents, 81 per cent had scaled or were actively expanding them, and the average organisation had automated 31 per cent of its workflows.
The numbers are not reconcilable and were never meant to be. They ask different questions of different people — a webinar audience self-selects for interest, executives report ambition, a vendor’s survey frames adoption generously — and “using agents” can mean a scheduled summary email or a system that approves payments. Read together they say something more useful than any one of them: curiosity is nearly universal, production use in a single function is common, and production use across a business is still rare. Gartner’s forecast that more than 40 per cent of projects will be scrapped by the end of 2027 sits comfortably beside its own forecast that a third of enterprise applications will contain agentic features by 2028. Both can be true. A lot of what gets built will be the wrong rung of the ladder.
What this means if you are the one deciding
- Write the rule first. If you can state the logic, code it. The model is for the parts you cannot state.
- Start with one call, and measure it. Add the loop only when the measurement says the single call is not good enough — and keep measuring once you add the loop, because it is now the only way to know which step is failing.
- Choose the rung on purpose. Decide how much of the control flow you are handing to the model, write it down, and make everything else deterministic.
- Build the evaluator before the planner. A system that can check its own work can afford to plan badly. A system that plans beautifully and never checks will fail quietly.
- Count the steps. Every unchecked step you remove is a multiplicative gain in reliability; every one you add has to pay for itself.
- Put a human before anything irreversible, cap the iterations and the spend, and run it in a sandbox until you have watched it fail.
- Pick tasks where success is checkable. Tests pass, tickets close, totals match. The domains where agents have worked all share this; the ones where they have not, mostly do not.
What would change this picture
- Per-step reliability reaching 99 per cent or better on ordinary business tasks, which would flatten the compounding curve enough that long unchecked chains become sane.
- The 80 per cent time horizon closing the gap on the 50 per cent one. Today a model finishes four-times-in-five tasks that are far shorter than its coin-flip tasks; if those converge, autonomy gets cheaper to trust.
- An independent survey — not a vendor’s — finding a majority of organisations running agents across several functions, or Gartner’s 2027 cancellation forecast failing to materialise.
- Verification becoming a solved component: standard, trustworthy checks for tool outputs and plans that work outside coding, where tests have done the job so far.
- A collapse in the price per token, which would change the cost side of the trade but not the reliability side.
Methods and limits
The HumanEval comparison is a synthesis Ng assembled from several teams’ published results, each on its own method, on a test old enough to have leaked into training data; it shows a direction, not a precise effect. METR’s horizons are point estimates with wide, stated confidence ranges, measured on software and research tasks that may not resemble yours, and its authors are the first to say so; the figure here uses the January 2026 revision, and METR has since revised some estimates again as it adjusted its model. The Berkeley failure study examined open-source frameworks running 2024-era models, and its category shares shifted between versions of the paper as more traces were annotated; the shares here are from the April 2025 revision. The three adoption surveys measure different things of different populations and are shown together to make that visible, not to be averaged. Gartner’s figures are forecasts and a poll of a self-selected audience. Nothing here is my own measurement. Where I give an opinion — on which rung to choose, on building the evaluator first — it is the opinion of someone who has sat in the rooms where these projects get approved, and it is offered as such.
Sources
- Andrew Ng, Agentic AI (DeepLearning.AI course), 2025 — module 1, “What is agentic AI?” and “Degrees of autonomy”.
- Andrew Ng, “Agentic Design Patterns Part 1”, The Batch, March 2024 — the HumanEval figures and the four patterns.
- Anthropic, “Building effective agents”, December 2024 — workflows versus agents, the six patterns, when to use which.
- Anthropic, “Donating the Model Context Protocol and establishing the Agentic AI Foundation”, December 2025.
- Hugging Face, Agents Course, “What is an agent?” — the agency levels table from the smolagents guide.
- METR, “Measuring AI Ability to Complete Long Software Tasks”, March 2025 (arXiv 2503.14499).
- METR, “Time Horizon 1.1”, January 2026 — the point estimates and ranges in Figure 4.
- Thomas Kwa (METR), “Clarifying limitations of time horizon”, January 2026.
- Cemri et al., “Why Do Multi-Agent LLM Systems Fail?”, UC Berkeley, 2025 (arXiv 2503.13657; NeurIPS 2025).
- Gartner prediction and January 2025 poll, as reported by IT Brief, June 2025.
- McKinsey, “The State of AI in 2025: Agents, innovation, and transformation”, November 2025.
- CrewAI, “2026 State of Agentic AI”, February 2026 — vendor survey.
- Utkarsh Kanwat, “Why I’m Betting Against AI Agents in 2025 (Despite Building Them)”, July 2025.
- Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, October 2022; Shinn et al., “Reflexion”, March 2023.