New Zero2Repo Benchmark Finds Top AI Coding Agent Builds 10 of 11 Repositories From a Spec, Weakest Manages Seven
A 26-author benchmark released this week swaps patch-the-repo coding tests for a harder task β turning a requirements document into a working project from an empty folder.
A benchmark released this week asks a question that most coding-agent evaluations have quietly avoided: can these systems build a piece of software from nothing, rather than patch one that already exists? Zero2Repo, posted to arXiv on September 29, 2026 by a 26-author team led by researchers at Gradient Data and McGill University, hands an agent a product requirements document, an interface contract, and an empty workspace, and checks whether what comes out actually works.
The headline result is narrow but pointed. Across the benchmark's 11 released tasks, the strongest agent tested β OpenAI's GPT-6 Astra running inside the Codex harness β completed 10. The weaker configuration in the study, xAI's Grok 4.7 High running inside Cursor CLI, completed 7. These are not toy exercises: the tasks are rebuilt from real, version-pinned open-source projects, meaning frontier models had a reasonable chance of having seen the original code during training, yet a fully correct repository was still not something every agent could reliably produce.

How the test was built
Most existing coding benchmarks fall into one of two categories. Tools like SWE-bench hand an agent an existing repository with a bug and ask for a patch; tools like HumanEval ask for a single isolated function. Neither resembles the workflow the paper's authors say many users now lean on: describe what you want, and let the agent build the whole project. Zero2Repo targets that gap specifically, covering parsers, protocols, command-line tools, security utilities, and software that depends on external resources, written in Python, TypeScript, Go, and C++.
Each task is produced by a pipeline that converts a real open-source project into a behavioral specification, a reproducible container environment, and a hidden acceptance-test suite. Before release, every task must pass a reference implementation and reject deliberately broken ones, so the tests are checked against both correct and incorrect code, not written once and assumed to work. Agents run inside isolated containers with no access to the original repository or to the hidden tests, and scoring is strictly binary: a task counts as solved only if every hidden test passes, with no partial credit and no language model acting as judge.
What the agents actually got wrong
The paper's results table, covering one run per task per configuration, looks like this:
| Model | Harness | Tasks solved | Pass rate | Mean time |
|---|---|---|---|---|
| GPT-6 Astra | Codex | 10 / 11 | 90.9% | 14.4 min |
| Claude Opus 5.5 | Claude Code | 9 / 11 | 81.8% | 19.7 min |
| Grok 4.7 High | Cursor CLI | 7 / 11 | 63.6% | 51.5 min |
The more telling number sits beneath the pass/fail count. On every task an agent failed, its submission still passed the large majority of hidden tests: Grok 4.7's four failures passed between 89.9% and 98.4% of the test suite, Claude Opus 5.5's two failures passed 97.0% and 99.4%, and GPT-6 Astra's single failure passed 95.8%. The authors trace most of these gaps to a single omitted edge case or a low-frequency rule buried in the specification, rather than an entire missing subsystem. In other words, the agents were not failing to understand the software they were asked to build β they were leaving small, specific holes in it, the kind that a binary pass/fail test suite catches and a casual read of the code would not.

Why a count of 11 still matters
The sample is small by design for now. The project's maintainers, who publish the benchmark's code and task data separately from the paper, describe the current 11 tasks as a first release and say they plan to expand to 30β50 tasks with broader language coverage. That makes the specific percentages above provisional rather than definitive β a handful of additional tasks could shift any single model's score by close to ten points.
What the result does establish, with the kind of granularity a benchmark is supposed to provide, is a specific failure mode. Tools that already report strong scores on patch-style benchmarks are not automatically equipped for end-to-end construction, and the gap between "mostly working" and "fully working" in that setting is concentrated in exactly the edge cases that hidden, deterministic tests are built to catch. For teams deciding how much unsupervised scope to hand a coding agent β a single function, a bug fix, or an entire new service β that distinction between near-complete and complete is the one the benchmark was built to measure.
A quick question for readers
Which skill will matter most for getting a job in 2030?
Everyone has a prediction. Add yours and see what students are betting on.
Pick an answer, create a free account in a minute, and your vote counts. Already a member? Sign in
- Pei Yang, Tianyu Shi, Yuhang Yao, et al.. Zero2Repo: Can Coding Agents Build Repositories from Scratch?. arXiv preprint, 2026. link
- OpenEdgeHQ. Zero2Repo: from a spec to a working repo (benchmark code and task data). GitHub, 2026. link
- OpenAI. GPT-6 Astra: the next generation in intelligence for work. OpenAI, 2026. link
- Anthropic. Claude Opus 5.5. Anthropic, 2026. link
- xAI. Introducing Grok 4.7. xAI, 2026. link