Google's FlowAgent AI Fixed 28,554 of 295,508 Flagged Pre-Submit Test Failures

A paper posted October 5, 2026 describes an AI agent that Google wired into its code-review tools to repair failing tests before developers even switch context.

EduFabTech · 8 October 2026 · 4 min read · 24 views
A failing pre-submit test moves through FlowAgent's propose-and-validate loop and comes out passing, alongside the headline counts of changes flagged, fixes previewed, and fixes applied.
EduFabTech · Own work

Google engineers posted a paper on arXiv on October 5, 2026 describing FlowAgent, an AI agent the company has built into its internal code-review and development tools to fix failing tests before a developer's change reaches review. The paper, by Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro and Lorenzo Dini, has been accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026), set for October 12–16, 2026 in Munich, Germany, according to the comments field on the paper's arXiv listing.

The target is a specific moment in a developer's day: the pre-submit phase, when a change triggers continuous-integration tests and some of them fail. Fixing that failure usually means a developer stops what they are doing, reads the test output, finds the bug and patches it. The paper's authors call this "time-consuming and disruptive," and argue that most prior automated program repair research has focused on post-submit, offline settings rather than the real-time, low-latency setting a developer actually works in, as stated in the paper's abstract.

A funnel chart shows the drop-off from 295,508 flagged changes to 65,069 previewed fixes to 28,554 applied fixes, with the 67.18% manual-evaluation accuracy rate noted below.
A funnel chart shows the drop-off from 295,508 flagged changes to 65,069 previewed fixes to 28,554 applied fixes, with the 67.18% manual-evaluation accuracy rate noted below.EduFabTech · Own work

How FlowAgent works

FlowAgent is integrated into two of Google's internal developer tools, Critique and Cider, the authors write in the paper. It runs a ReAct-style "generate-and-validate" loop: the agent proposes a candidate fix, then checks it by re-running the failing test, iterating until it either produces a change it is confident in or gives up. That second option matters as much as the first — the system applies what the authors describe as "rigorous pre-execution and post-execution abstention filters," designed to hold back a suggestion rather than show developers a fix that is likely wrong, under what the paper calls "strict latency constraints."

That abstention design separates FlowAgent from a tool that simply tries to repair every failure it sees. Posting a wrong suggestion in a live workflow costs a developer's attention directly, so the system is built to suggest less often rather than suggest badly.

What the numbers show

The paper reports two separate measurements. First, a manual evaluation on 195 real-world test failures found FlowAgent produced a correct fix 67.18% of the time, according to the abstract. Second, after Google rolled the agent out company-wide, FlowAgent suggested fixes on 295,508 changes; developers previewed 65,069 of those suggestions and applied 28,554, the authors report.

Read together, those production figures mean roughly 44% of previewed suggestions were applied, and about 10% of all changes that received a suggestion ended with the developer taking it. The gap between fixes suggested and fixes applied is the paper's central operational finding: even a system tuned to abstain rather than guess still produces far more suggestions than developers choose to use, which is consistent with the authors' point that "interesting challenges and opportunities still remain" in bringing autonomous repair agents into daily engineering work.

A three-step diagram of FlowAgent's ReAct-style cycle: a failing test triggers the agent, which proposes and re-runs a fix in a loop, then either gets applied by the developer or is abstained on.
A three-step diagram of FlowAgent's ReAct-style cycle: a failing test triggers the agent, which proposes and re-runs a fix in a loop, then either gets applied by the developer or is abstained on.EduFabTech · Own work

Where this sits in the research line

FlowAgent is not Google's first publication on AI-driven program repair. A separate Google team — Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano and Satish Chandra — previously described an earlier system, Passerine, in "Evaluating Agent-based Program Repair at Google" (2025). That paper evaluated agent-based repair offline, against a curated set of bugs from Google's internal issue tracker, without the pre-submit latency constraint that defines FlowAgent. The new paper frames its contribution specifically around fitting repair into the "flow" of a developer's pre-submit workflow, rather than treating repair as a background or offline task.

Why it matters for engineers and researchers

For software engineering researchers, FlowAgent is a rare published account of an agentic coding tool measured against real production use rather than a static benchmark — the kind of deployment data that benchmark-only evaluations of coding agents cannot provide. For practitioners building or evaluating similar tools, the paper's abstention-filter design and its gap between "previewed" and "applied" rates are concrete reference points: a system can be accurate in controlled evaluation and still see most of its live suggestions declined by the people it is built to help.

The authors describe developer interviews as part of their evaluation, reporting that engineers found the agent "useful in suggesting correct fixes" and that integrating autonomous repair agents into industrial workflows was "received well," while also flagging that "interesting challenges and opportunities still remain," per the paper's abstract. Neither the paper nor Google's public documentation gives a date for when FlowAgent was first switched on internally, so the reported figures should be read as describing the deployment period covered by the study, not a fixed calendar window.

A quick question for readers

Source: arXiv (Google)

Sources (3)
  1. Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini. Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale. arXiv (accepted at ASE 2026), 2026. arxiv.org ↗ · checked 8 Oct 2026
  2. Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano, Satish Chandra. Evaluating Agent-based Program Repair at Google. arXiv, 2025. arxiv.org ↗ · checked 8 Oct 2026
  3. IEEE/ACM. ASE 2026 — 41st IEEE/ACM International Conference on Automated Software Engineering. conf.researchr.org (IEEE/ACM), 2026. conf.researchr.org ↗ · checked 8 Oct 2026