← All articles

"Session ended", "did not respond in time", "invalid handoff request": What Agent Session Errors Mean

TL;DR

"Did not respond in time" is a liveness timeout: the runner stopped hearing from the agent or its browser. "Session ended" means the browser or the run was torn down underneath the task. "Invalid handoff request" means a human handoff was attempted against a run that no longer exists or was already claimed. All three share one root: the run and the thing waiting on it disagree about who is still alive.

These three messages come from three different layers, but they are one bug seen from three sides. A runner has a clock. A browser session has a clock. A human being asked to step in has a clock. When one of them expires while another is still working, the waiting side declares the working side dead, tears it down or refuses its next request, and you get one of these strings. The fix is never a longer timeout. It is to stop using one clock for two different kinds of waiting.

The three messages

MessageWhich layer says itWhat actually happenedFirst thing to check
did not respond in timeThe runner or orchestratorA liveness check fired. The agent was mid tool call, mid page load, or waiting for a person, and produced no heartbeat inside the windowWas anything still happening? A long tool call and a dead process look identical to a sweeper
session endedThe browser provider or sandboxThe browser was closed or reclaimed: idle limit, max session length, the run above it was cancelled, or the container restartedThe provider's session length limit, and whether a cancel from the runner cascaded down
invalid handoff requestThe human-in-the-loop layerA handoff was requested for a run that is gone, already handed off, or already resolved. Often the same task asked twice after a retryWhether the retry created a second handoff for a task that already had one

One root cause: two clocks

A run that is executing and a run that is waiting on a person are different states with different budgets, and most stacks give them the same one. Executing should time out in seconds to minutes. Waiting on a person is legitimately five minutes to a day. When a sweeper built for the first sees the second, it sees a stale run, marks it dead, and either kills it, which becomes "session ended", or dispatches a second copy, which becomes "invalid handoff request" when both copies try to hand off.

Fix one: waiting is its own state

When the agent decides a person is needed, it should record that, with its own deadline, and stop consuming a runner slot. Persist the run, exit, and resume when the person is done. A process parked on an await through a human wait is the shape that breaks under every sweeper ever written.

from doubleoh import DoubleOh

oo = DoubleOh()

def on_wall(run, page, task):
    fix = oo.request_fix(url=page.url, task=task)
    run.mark_waiting(fix_id=fix.id, deadline=fix.expected_seconds)   # a different budget
    run.persist(); return  # exit the worker; nothing is held open

def on_resume(run):
    status = oo.intervention(run.fix_id)      # or wait_for_fix, which long-polls for 55 s
    if status.resolved:
        continue_task(run)

Fix two: handoffs are idempotent

Give every handoff a stable key derived from the task and the wall, not from the attempt. A retry, a second worker, or a sweeper's duplicate then gets the same handoff back instead of creating a second one, and "invalid handoff request" stops appearing because there is nothing invalid to request. The same rule applies to the side effect the handoff protects: the human should be asked once, not once per copy of the run.

What DoubleOh does about this

A fix request returns a link that stays valid until the person is done, an expected duration, and a wait URL that long-polls for up to 55 seconds so the caller never needs a tight loop. Asking again for the same wall while a fix is open returns the same fix rather than a new one, so one stuck task does not become ten messages. And because the person works on the agent's live screen, the session the agent needs is the session the person is in; nothing has to be rebuilt after the handoff.

Questions people ask

"Session ended as the agent did not respond in time": what does it mean?

A liveness timeout on the runner fired while the agent was busy or waiting, and the session was torn down as a result. Check whether the run was waiting on a long tool call or on a person. Those need a longer, separate budget from ordinary execution, and ideally the run should persist and exit rather than hold a process open.

"Invalid browser session handoff request": how do I fix it?

Something asked for a human handoff on a run that no longer exists or already has one. The usual cause is a retry or a duplicate worker creating a second handoff for the same task. Make the handoff idempotent: a stable key per task and wall, so a repeat returns the existing handoff.

How long should an agent wait for a human?

As long as the business allows, which is minutes to a day, and that budget must be different from the one for executing steps. Do not stretch the execution timeout to cover it; record the waiting state with its own deadline and resume later.

Should the agent process stay alive while a human works?

No. Persist the run, release the worker, and resume on a signal or a poll. A process parked on an await through a human step is what sweepers kill and what produces "session ended".

DoubleOh is the fix desk for AI agents. When one gets stuck, a person fixes it once in a live browser, and the fix becomes a skill the whole fleet follows from then on.

Start free