
AgencyBench is a benchmark for evaluating long-horizon autonomous-agent workflows in real-world-style settings. The supplied arXiv record describes 32 scenarios with iterative user feedback, Docker sandboxes, detailed rubrics, and an average of approximately 90 tool calls per scenario. The official GAIR-NLP repository provides V1/V2 benchmark materials, including scenarios, queries, deliverables, rubrics, evaluati...
AgencyBench is presented in the supplied arXiv record as a benchmark for long-horizon autonomous-agent workflows in “1M-token real-world contexts.” Its unit of evaluation is a scenario requiring sustained execution rather than a compact, single-turn answer. According to the supplied paper excerpt, AgencyBench contains 32 real-world scenarios, includes iterative user feedback and Docker sandboxes, uses detailed rubrics, and requires an average of approximately 90 tool calls per scenario.
This framing distinguishes the benchmark from evaluations centered on isolated reasoning questions or a small number of function calls. An evaluated agent is expected to progress through a multi-step workflow, invoke tools repeatedly, produce task-specific deliverables, and potentially revise its work after feedback. The repository description further indicates that AgencyBench provides detailed task queries, deliverables, rubrics, evaluation scripts, and Docker-based evaluation support.
The available material supports viewing AgencyBench primarily as an environment-and-deliverable benchmark, not as a new agent architecture. It does not describe a required planning loop, memory system, retrieval component, reflection procedure, verifier, subagent topology, ReAct-style prompt format, or orchestration framework. Those implementation choices appear to be left to systems evaluated on the suite.
The benchmark title foregrounds “1M-token” contexts, but the supplied sources do not specify how this quantity is measured. In particular, they do not establish whether one million tokens refers to an initial prompt, accumulated interaction history, attached files, artifacts in the sandbox, retrieved material, environmental state, or another context-accounting method. Nor do they specify whether every scenario reaches that scale or whether participating models must support a particular context-window length.
AgencyBench evaluates agents through scenario-based workflows with several supported operational components:
The design places pressure on an agent’s ability to preserve task state across many dependent actions. It also creates opportunities for failures that may not appear in short-horizon benchmarks: losing track of user constraints, failing to inspect intermediate artifacts, looping on ineffective tools, making invalid environment changes, or producing a final deliverable that does not satisfy the rubric. These are plausible long-horizon failure modes, but the supplied sources do not provide a measured error taxonomy or frequencies for them.
The supplied evidence establishes an evaluation setting with 32 scenarios, iterative user feedback, Docker sandboxes, detailed deliverables and rubrics, and trajectories averaging about 90 tool calls. The official GAIR-NLP/AgencyBench repository is described as containing AgencyBench V1/V2 materials, task scenarios, detailed queries, expected deliverables, evaluation rubrics, and evaluation scripts.
This setup emphasizes operational features that compact agent benchmarks often abstract away: maintaining task progress over many actions, managing dependencies between tool outputs, producing artifacts rather than merely conversational responses, and revising work after feedback. Scenario-specific deliverables may also make it possible to inspect partial completion and deviations from requested workflows, depending on how the associated rubrics and scripts are implemented.
However, the supplied sources do not specify the scoring procedure. It is not established whether evaluation uses binary success, partial credit, multi-axis rubric scores, automated checks, human judgment, artifact validation, trajectory efficiency, tool-call counts, wall-clock completion time, token consumption, or a composite metric. The role of detailed rubrics is supported, but their dimensions, weights, and automatic-versus-manual grading split are not.
Other consequential protocol details are likewise unspecified: model prompting templates, sampling temperatures, repeated-trial policy, scenario splits, held-out evaluation sets, agent permissions, persistent-file behavior, external-service access, and contamination controls. As a result, the available record supports the benchmark’s broad evaluation framing but not strong conclusions about comparability across agent implementations.
The supplied excerpts provide benchmark characteristics, not numerical comparative results. They do not include a leaderboard, evaluated-model list, baseline-agent implementation, aggregate score table, task-level success rates, confidence intervals, cost estimates, or a published failure analysis.
Accordingly, the source packet does not establish which frontier model, open-weight model, context-management method, planning approach, or agent framework performs best on AgencyBench. It also does not specify whether performance is measured principally through rubric completion, valid deliverables, automated tests, human adjudication, efficiency, or another outcome measure.
The central quantitative operational detail available is the reported average of approximately 90 tool calls per scenario. This makes long-horizon reliability a likely determinant of success: an agent can be locally competent at individual actions while failing overall because an earlier mistake propagates through the workflow. That interpretation follows from the benchmark’s multi-step design, but it should not be treated as a reported experimental finding because the supplied sources do not present causal analyses or ablations.
The strongest reproducibility signal in the supplied material is the stated availability of official benchmark artifacts. The GAIR-NLP/AgencyBench repository is described as providing:
Together, these resources can enable implementers to inspect task requirements, recreate at least some execution conditions, evaluate artifacts using available scripts, and compare outputs against published rubric definitions. The official AgencyBench project site is also listed alongside the paper and repository as a benchmark resource.
The supplied sources do not specify a software license, pinned dependencies, release tags, container registry, hardware requirements, model-provider settings, seed policy, rate-limit handling, or exact instructions for reproducing reported experiments. They also do not establish whether every scenario is executable using only public local resources, or whether some tasks require credentials, paid APIs, network access, or additional setup. Prospective users should inspect the repository and project site directly before assuming a turnkey reproduction path.
Several caveats are clear from the source record.
First, 32 scenarios may offer useful task diversity, but scenario count alone does not establish broad coverage. The supplied sources do not characterize domains, geographic or organizational settings, user populations, software ecosystems, task complexity, or representativeness.
Second, Docker support can improve environmental control, but it does not automatically ensure reproducible outcomes. Agent performance can vary with model version, prompt construction, sampling configuration, controller logic, tool implementations, external dependencies, and nondeterministic execution. Controls for these factors are not specified in the supplied sources.
Third, iterative feedback may improve realism while adding ambiguity. Because the available excerpts do not define who or what generates feedback, different implementations of the feedback policy could materially change task difficulty and agent behavior.
Fourth, an average of 90 tool calls demonstrates trajectory length, not necessarily that every action requires difficult planning. Without task traces, action distributions, and ablation studies, the available material cannot separate genuine planning burden from workflow or interface overhead.
Finally, the paper title’s “frontiers” framing should not be read as evidence of comparative frontier-model performance from this source packet. No baseline results or leaderboard evidence are included in the supplied excerpts.
AgencyBench is relevant to agent builders because it evaluates properties closer to operational workflow execution than conversational plausibility: persistent task state, repeated tool interaction, artifact production, adaptation after feedback, and execution within bounded environments. A system that appears capable in short tool-calling demonstrations may fail when it must coordinate dozens of dependent actions and satisfy a detailed deliverable rubric.
The published-material description makes the suite potentially useful for testing practical agent behaviors, including whether an agent maintains user intent across feedback turns, recovers after tool errors, checks intermediate artifacts, validates outputs before submission, and avoids unproductive action loops. The supplied sources do not confirm that each behavior is independently scored, but they are natural concerns for systems operating under the benchmark’s scenario-and-deliverable format.
AgencyBench should nevertheless be treated as one evaluation signal rather than a complete measure of deployable autonomy. The supplied materials do not specify assessments of security, unsafe action handling, privacy, adversarial instructions, latency, cost, reliability under external-service failure, or robustness across model updates. Those dimensions require separate evaluation even for systems that perform well in Docker-supported long-horizon workflows.
Published Jan 16, 2026, 12:00 AM
By Kainotomic Team
Published Jan 16, 2026, 12:00 AM