Long-horizon work spans enough steps that early decisions and retained information can affect much later actions. The agent must manage dependencies, unfinished work and changes in its environment rather than only producing a response to one short request.
Errors can accumulate through stale memories, lost sources or plans that no longer match the task. Evaluation should examine behavior across extended scenarios, while preserving provenance and checking whether stored state actually improves subsequent decisions.
ELI5
This is an AI agent that works through a job with many connected steps. It needs to keep track of what happened and what remains, not simply remember a large amount of text.
For example, a research agent may collect sources, compare claims and produce a report over several sessions. If a source was rejected as unreliable, later sessions should retain that decision and its reason rather than treating the rejected claim as fact.
Does a large context window make an agent reliable over long tasks?
Not by itself. Planning, source retention, state updates and outcome checks also matter.
How should long-horizon performance be tested?
Use extended, representative scenarios and inspect both final outcomes and whether important state remains accurate throughout the run.



