New Zero2Repo Benchmark Finds Top AI Coding Agent Builds 10 of 11 Repositories From a Spec, Weakest Manages Seven

A 26-author benchmark released this week swaps patch-the-repo coding tests for a harder task β€” turning a requirements document into a working project from an empty folder.

EduFabTech Β· 2 October 2026 Β· 4 min read Β· 2 views
An 11-cell task grid shows GPT-6 Astra (Codex) building 10 of 11 repositories from a spec, the best score in the Zero2Repo benchmark.
EduFabTech · Own work

A benchmark released this week asks a question that most coding-agent evaluations have quietly avoided: can these systems build a piece of software from nothing, rather than patch one that already exists? Zero2Repo, posted to arXiv on September 29, 2026 by a 26-author team led by researchers at Gradient Data and McGill University, hands an agent a product requirements document, an interface contract, and an empty workspace, and checks whether what comes out actually works.

The headline result is narrow but pointed. Across the benchmark's 11 released tasks, the strongest agent tested β€” OpenAI's GPT-6 Astra running inside the Codex harness β€” completed 10. The weaker configuration in the study, xAI's Grok 4.7 High running inside Cursor CLI, completed 7. These are not toy exercises: the tasks are rebuilt from real, version-pinned open-source projects, meaning frontier models had a reasonable chance of having seen the original code during training, yet a fully correct repository was still not something every agent could reliably produce.

A bar chart compares pass rates across the three tested configurations: GPT-6 Astra/Codex at 90.9%, Claude Opus 5.5/Claude Code at 81.8%, and Grok 4.7 High/Cursor CLI at 63.6%.
A bar chart compares pass rates across the three tested configurations: GPT-6 Astra/Codex at 90.9%, Claude Opus 5.5/Claude Code at 81.8%, and Grok 4.7 High/Cursor CLI at 63.6%.EduFabTech · Own work

How the test was built

Most existing coding benchmarks fall into one of two categories. Tools like SWE-bench hand an agent an existing repository with a bug and ask for a patch; tools like HumanEval ask for a single isolated function. Neither resembles the workflow the paper's authors say many users now lean on: describe what you want, and let the agent build the whole project. Zero2Repo targets that gap specifically, covering parsers, protocols, command-line tools, security utilities, and software that depends on external resources, written in Python, TypeScript, Go, and C++.

Each task is produced by a pipeline that converts a real open-source project into a behavioral specification, a reproducible container environment, and a hidden acceptance-test suite. Before release, every task must pass a reference implementation and reject deliberately broken ones, so the tests are checked against both correct and incorrect code, not written once and assumed to work. Agents run inside isolated containers with no access to the original repository or to the hidden tests, and scoring is strictly binary: a task counts as solved only if every hidden test passes, with no partial credit and no language model acting as judge.

What the agents actually got wrong

The paper's results table, covering one run per task per configuration, looks like this:

ModelHarnessTasks solvedPass rateMean time
GPT-6 AstraCodex10 / 1190.9%14.4 min
Claude Opus 5.5Claude Code9 / 1181.8%19.7 min
Grok 4.7 HighCursor CLI7 / 1163.6%51.5 min

The more telling number sits beneath the pass/fail count. On every task an agent failed, its submission still passed the large majority of hidden tests: Grok 4.7's four failures passed between 89.9% and 98.4% of the test suite, Claude Opus 5.5's two failures passed 97.0% and 99.4%, and GPT-6 Astra's single failure passed 95.8%. The authors trace most of these gaps to a single omitted edge case or a low-frequency rule buried in the specification, rather than an entire missing subsystem. In other words, the agents were not failing to understand the software they were asked to build β€” they were leaving small, specific holes in it, the kind that a binary pass/fail test suite catches and a casual read of the code would not.

A side-by-side comparison shows how a failed task can still pass 95.8% of hidden tests, contrasted with the 100% binary threshold actually required to score a solve.
A side-by-side comparison shows how a failed task can still pass 95.8% of hidden tests, contrasted with the 100% binary threshold actually required to score a solve.EduFabTech · Own work

Why a count of 11 still matters

The sample is small by design for now. The project's maintainers, who publish the benchmark's code and task data separately from the paper, describe the current 11 tasks as a first release and say they plan to expand to 30–50 tasks with broader language coverage. That makes the specific percentages above provisional rather than definitive β€” a handful of additional tasks could shift any single model's score by close to ten points.

What the result does establish, with the kind of granularity a benchmark is supposed to provide, is a specific failure mode. Tools that already report strong scores on patch-style benchmarks are not automatically equipped for end-to-end construction, and the gap between "mostly working" and "fully working" in that setting is concentrated in exactly the edge cases that hidden, deterministic tests are built to catch. For teams deciding how much unsupervised scope to hand a coding agent β€” a single function, a bug fix, or an entire new service β€” that distinction between near-complete and complete is the one the benchmark was built to measure.

A quick question for readers


References
  1. Pei Yang, Tianyu Shi, Yuhang Yao, et al.. Zero2Repo: Can Coding Agents Build Repositories from Scratch?. arXiv preprint, 2026. link
  2. OpenEdgeHQ. Zero2Repo: from a spec to a working repo (benchmark code and task data). GitHub, 2026. link
  3. OpenAI. GPT-6 Astra: the next generation in intelligence for work. OpenAI, 2026. link
  4. Anthropic. Claude Opus 5.5. Anthropic, 2026. link
  5. xAI. Introducing Grok 4.7. xAI, 2026. link