How to Know When an AI Agent Has Hit a Wall and Needs a Human, Instead of Retrying Forever
TL;DR
An agent has hit a wall when the page stops changing in response to its actions, when it lands on a login, code, CAPTCHA or permission screen, or when its last actions repeat. Give every run a retry budget of two. At the wall, hand the exact screen to a person once, and keep what they did so the next run does not need them.
You know an agent has hit a wall when its actions stop changing the world. Not when the model says it is stuck; models rarely say that, they say "let me try a different approach" and try the same one. Read the page, not the transcript. If the screen after an action is the screen before it, or it is a screen that exists to stop automation, the run is at a wall and every further retry is a cost with no upside.
The five signals
All five are observable from outside the model. That matters: a detector that depends on the model noticing its own loop will miss most loops, because the model is the thing that is looping.
| Signal | How to measure it | What it usually is |
|---|---|---|
| No state change | Hash the DOM or the screenshot after each action. Two identical hashes in a row after actions that should change something | A click that did nothing: an overlay, a disabled button, a form that rejected the input silently |
| Same URL after navigation | The action was a submit or a link and the URL did not move, or moved and bounced back | A login redirect, a validation error, a bot check that swallowed the POST |
| A wall page | A password field, a one-time-code input, a CAPTCHA iframe, a 401 or 403, the words "not authorised" or "verify you are human" | Something the agent structurally cannot do: a credential, a second factor, a permission, a human check |
| Repeated action | The same tool call with the same arguments twice in the last three steps | The model has run out of ideas and is re-rolling |
| Time to progress | Wall-clock since the last state change, not since the run started | A slow page is fine for thirty seconds; three minutes with nothing new is a wall |
Why retrying forever is the worst outcome, not a neutral one
Three costs stack. Money: every retry is a full model call with the page in context, and loops run until a timeout that is usually measured in hours. Trust: sites score sessions, and five identical failed logins teach the site that this session is a bot and often lock the real account. Time: the run holds a browser, a queue slot and sometimes a person's attention while it produces nothing. An agent that stops at attempt two and asks is cheaper in every one of those currencies than an agent that tries thirty times and asks anyway.
The budget
Two. One retry covers the genuinely flaky cases, a selector that rendered late, a network blip. The second identical failure is information: the wall is real. After that the right move is not a third attempt with a cleverer prompt, it is a different actor.
from doubleoh import DoubleOh
oo = DoubleOh() # reads DOUBLEOH_API_KEY
RETRY_BUDGET = 2
def run(page, task):
hint = oo.skills_for(task, url=page.url) # has a person already cleared this wall?
if hint:
return follow(hint, page)
for attempt in range(RETRY_BUDGET):
before = page.state_hash()
act(page, task)
if page.state_hash() != before and not looks_like_wall(page):
return finish(page, task)
# Two identical failures: a wall. Hand the live screen to a person, once.
fix = oo.request_fix(url=page.url, task=task)
oo.wait_for_fix(fix.id) # long-polls; returns when a person is done
return finish(page, task)What to do at the wall
Stop the loop and change the actor. The person needs three things: the exact screen the agent saw, the task in one sentence, and a way to act on that same session rather than starting over in their own browser. That last part is what turns a handoff into a fix that can be kept: if the person does the step on the agent's live screen, what they did can be recorded, minus anything they typed, and compiled into a procedure. The next time any agent hits the same wall on the same site, the first call in the code above returns that procedure and nobody is paged.
What a wall is not
| Looks like a wall | Actually is | Right response |
|---|---|---|
| Element not found once | Render timing | Wait and retry once |
| Page slow, then changes | A slow site | Wait for a state change, not a fixed sleep |
| HTTP 429 | Rate limit | Back off with jitter; do not hand to a person |
| Same screen after two real attempts | A wall | A person, once, on the live screen |
The distinction is whether time or a different action can change the outcome. If yes, wait or retry. If no, no amount of either helps, and the only variable left is who is acting.
Questions people ask
How do I know when my agent has hit a wall and needs a human instead of just retrying forever?
When two attempts leave the page unchanged, when the page is a login, code, CAPTCHA or permission screen, or when the agent repeats the same action with the same arguments. Measure it from the page, not from the model's transcript. At that point hand the exact screen to a person once, and keep what they did.
How many times should an AI agent retry before asking a human?
Two. One retry absorbs flakiness. A second identical failure means the wall is real, and a third attempt costs money and raises the site's suspicion score without changing the outcome.
What counts as a wall in an AI agent run?
Anything time and a different action cannot change: a credential the agent does not have, a second factor, a permission it was not granted, a human check, or a rule of the business nobody wrote down. Rate limits and slow pages are not walls.
Can I detect a loop from the agent's own reasoning?
Not reliably. The model that is looping is the model you would ask, and it will describe a new approach while repeating the old one. Hash the page state and compare tool calls instead.
Does DoubleOh detect walls for me?
The detection lives in your loop, because only your code sees the page. What DoubleOh adds is what happens next: the SDK asks whether a person has already cleared this wall on this site, and if not, requests a fix that puts a person on the agent's live screen once. The fix compiles into a skill your fleet reuses, with an explicit boundary of what still needs a human.
DoubleOh is the fix desk for AI agents. When one gets stuck, a person fixes it once in a live browser, and the fix becomes a skill the whole fleet follows from then on.
Start free