Loop engineering: when a system steers the AI agent
For almost two years the lever in working with AI agents was the prompt. That is coming to an end. What loop engineering is, when it pays for itself and when it only burns budget.
For almost two years, the leverage in working with AI agents sat in the prompt: you write an instruction, read the output, correct it and write again. Your hand never leaves the tool. That is now changing. Loop engineering means designing a small system that runs this cycle for you, so that you supervise the system rather than individual prompts.
In short
- Loop engineering is building a small system that finds a task, hands it to an AI agent, checks the result and decides the next move. Addy Osmani breaks such a loop into six parts. You design it once, and from then on the loop prompts the agent.
- A loop pays off only when four conditions hold at once: the task repeats at least once a week, verification is automatic, the token budget can absorb waste, and the agent has the tools a senior engineer would have.
- The heart of every good loop is a hard gate: a test, a build or a linter that can reject work without you. Without it the agent signs off its own work. Anthropic described separating the author from the verifier as the evaluator-optimizer pattern back in December 2024.
- To be honest, most companies do not need a loop yet. If yours does, start small with four parts: one automation, one skill, one state file and one gate.
What loop engineering is, and why the shift is real
Loop engineering is the design of a system that, in your place, feeds tasks to an AI agent, checks the output and chooses the next step. Until now the lever was the prompt, meaning whatever you typed. Now the lever is the loop, meaning the system that types for you. The point of leverage has moved one level up.
Addy Osmani, an engineer known for his web performance work at Google, breaks such a loop into six parts: find a task, hand it to the agent, check the result, record what happened, decide the next move, repeat. Anthropic reports that its engineers now merge about 8 times more code per day than they did in 2024, although the company itself notes that the figure almost certainly overstates the real productivity gain.
That caveat is welcome, because this particular number is disputed. The mechanism is not. The lever has moved from writing prompts to designing the loop that writes them. For anyone running a business the pattern is familiar: it is the same logic as any process automation, only with an AI agent in the middle instead of a simple "if this, then that" rule.
The four-condition test: when a loop pays off and when it burns budget
Before you build any loop, check four conditions. A loop pays for itself only when all of them hold at the same time. If a single one is missing, the cost outweighs the benefit. It is the cheapest test you can run, because it does not use a single token.
- The task repeats, at least once a week, so the one-off set-up cost has something to be amortised against.
- Verification is automatic: a test suite, a type check, a linter or a build can reject the work without you in the room.
- The token budget can absorb waste: loops re-read context, retry and explore, and the tokens burn whether or not a given run delivers anything.
- The agent has a senior engineer's tools: logs, an environment in which to reproduce the bug, and the ability to run code and see what actually breaks.
It reads like a checklist for the IT department, because it is one. The conclusion, though, is universal and identical to any other automation decision: hand the machine what repeats and what can be checked without a person. Everything else remains work for people.
Who gains, and who should hold off
Loops favour those who can afford to feed them. Teams with repetitive, machine-checkable work and a budget for it come out ahead. Teams that build a loop where there is nothing to check, or where the bottleneck is a person rather than the writing, lose out.
Who gains: teams with repetitive, verifiable tasks such as triaging CI failures, updating dependencies, applying bulk linter fixes, or turning a draft ticket into a finished pull request on well-tested code. Strong test suites and an asynchronous working culture help them.
Who loses: solo builders on cheap plans (the token bill arrives before the gain), code without automatic verification (the loop simply agrees with itself), and teams whose real bottleneck is code review rather than writing (the loop only lengthens the review queue). For one-off tasks, or work that calls for judgement, a single well-aimed prompt still wins. It is the same calculation you make when choosing between an off-the-shelf tool and a bespoke one.
The five building blocks of a loop
Every loop breaks down into five elements. You do not need all of them at once, but it is worth knowing what each one is for before someone sells you a "swarm of agents" that nobody can verify. In plain terms:
- Automations are the loop's heartbeat. They fire on a schedule or on an event, and everything else hangs off them.
- Work isolation (in code, a separate working directory on its own branch) stops two agents from treading on each other's files. Mechanical collisions disappear, but your capacity to review remains the ceiling on how many agents can run in parallel.
- Skills are project knowledge written down once: conventions, build steps, "we don't do it this way because it broke things once". Without them the loop rebuilds its context from scratch on every cycle.
- Connectors (through MCP, the Model Context Protocol) let the agent touch real tools: the ticketing system, the database, an API, the team chat. That is the difference between "here is a fix" and a loop that opens the ticket itself and tells you when the tests are green.
- Subagents separate whoever writes from whoever checks. It is the most important structural move in the whole arrangement.
That last block is the key one. As Osmani puts it, the model that wrote the code is far too kind when it grades its own work. A second agent, with different instructions, catches what the first one talked itself into. None of this is new vocabulary: Anthropic described the arrangement as the evaluator-optimizer pattern in its engineering post from December 2024, a year and a half before loop engineering became fashionable. The practical conclusion: a verifier you trust is the only reason you can walk away from the screen.
The minimum viable loop, and the one number worth measuring
If you passed the four-condition test, build the smallest version you can, not a swarm. Four parts are enough: one automation, one skill, one state file and one gate. The order matters, and it is usually what decides whether the loop works.
First, get one manual run to the point where it is repeatable. Then turn it into a skill. Then wrap it in a loop. Only at the very end put it on a schedule. The state file (a plain text file in the repository, or a task board) exists so that tomorrow the loop resumes its work instead of starting from zero. The rule is simple: the agent forgets, the repository does not.
Then there is the one number that really counts: cost per accepted change, not tokens consumed and not the number of runs. If fewer than half of the changes the loop produces are accepted, you are doing by hand exactly the work the loop was meant to save you. At that point the loop is not earning its keep, it is adding cost.
The work did not get easier. The point of leverage moved. Build the loop, but stay an engineer.
Where loops fail silently
The most dangerous loop failure is a silent one: it keeps costing money while delivering nothing useful. The engineer Geoffrey Huntley called it the Ralph Wiggum loop. The agent is meant to report completion only once the task is finished, but it reports it too early, so the loop closes halfway through the job and carries on burning tokens. These are the typical traps and their cures:
- No real verifier: two optimists nodding at each other. The cure is an objective gate that can reject work (the test passes or it does not, the build compiles or it does not), not a reviewer armed only with an opinion.
- Goal drift in long sessions: every summary drops a little context, and "don't do X" vanishes somewhere around turn 47. The cure is a fixed specification loaded on every run.
- Self-serving bias: the author is too lenient with its own work. The cure is a separate subagent acting as the verifier.
- Agent laziness: the agent declares a partial result "good enough". The cure is a stop condition checked by a fresh, independent model.
The security tax and comprehension debt
A loop running unsupervised is an unsupervised attack surface. The more efficiently it ships code you did not write, the wider the gap between what sits in the repository and what you actually understand. These are two bills that arrive later, and both can hurt.
The security tax. A loop puts changes up for review faster than a person can read them. Without a gate that includes vulnerability scanning, a dependency audit and secret detection, unsafe code merges itself. Skills can also be an attack vector: a loop that installs community skills on its own inherits every prompt injection hidden inside them. In one audit, 520 of the 17,022 skills examined were leaking credentials, so read the source before you install anything. Watch for permission creep as well: a loop tested as "read-only" is granted "just this one" write permission, and nobody ever goes back to review it.
Comprehension debt. Here the cures are not technical. Read the diffs. Spot-check whether the gate really catches the failure you care about, because gates degrade over time. Keep the loop away from architectural decisions. Osmani describes the biggest risk as a kind of cognitive surrender: the temptation to stop holding your own view and simply accept whatever the loop returns. It is the cheapest risk to avoid and the most expensive one once it has happened.
What this means for your business
For most companies the honest conclusion is simple: not yet. A loop makes sense only when the task repeats, verification is automatic, the budget can absorb waste and the agent has the tools a senior engineer would have. Before anyone sells you an "autonomous agent", walk through those four conditions with them, out loud.
It is the same logic you apply to any automation decision. Does the thing you want to hand to a machine really repeat, and can the result be checked without a person? That question holds whether you are comparing workflow tools or choosing a partner to implement AI for you.
The same shift applies to visibility in Google and in AI answers, the territory of Generative Engine Optimization. Repetitive, checkable work, such as rank monitoring, keeping content fresh and validating structured data, can be handed to a loop. Deciding what to publish, and why, is still a human call. If you would rather hand that repetitive part to us, start with the free SEO and GEO audit or check our pricing.
Start small: one automation, one skill, one state file, one gate. Get one manual run to the point of repeatability, turn it into a skill, wrap it in a loop, and only then schedule it. The work did not get easier. The point of leverage moved.
Common questions
What is loop engineering in one sentence?
It is the design of a small system that, instead of you, finds a task, hands it to an AI agent, checks the result and chooses the next step. You used to write the prompt yourself; now you design the loop once and the loop prompts the agent. The lever moves from writing instructions to designing the system that writes them.
When does building an agent loop pay off?
When four conditions hold at once: the task repeats at least once a week, verification is automatic (a test, a build, a linter), the token budget can absorb waste, and the agent has a senior engineer's tools (logs, an environment, the ability to run code). Miss one and the loop costs more than it returns.
What is the Ralph Wiggum loop?
It is a silent failure described by the engineer Geoffrey Huntley: the agent is meant to report completion only once the task is finished, but it reports it too early, so the loop stops halfway through the job and keeps burning tokens. The cure is a hard gate that can objectively reject unfinished work.
Why does a loop need a separate agent to do the checking?
Because the model that wrote the code is too lenient with its own work. A second agent with different instructions catches what the first one talked itself into. Anthropic described this as the evaluator-optimizer pattern back in December 2024. A verifier you trust is the only reason you can walk away from the screen.
Does my company need agent loops right now?
Usually not yet. If the task does not repeat, cannot be checked by a machine, or the bottleneck is review rather than writing, one well-aimed prompt wins. A loop pays for itself only with repetitive, automatically verifiable work and a budget for waste.
Read next
Find out whether AI recommends your company.
Start with the free SEO and GEO audit, delivered in 5 working days. We check how the models describe your brand and hand back a prioritised list of changes.