Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Google Releases Auto-Diagnose: LLM-Based System for Diagnosing Integration Test Failures at Scale
00

Google Releases Auto-Diagnose: LLM-Based System for Diagnosing Integration Test Failures at Scale

Apr 19, 2026

Google AI has released Auto-Diagnose, an LLM-based system that automatically diagnoses integration test failures by analyzing logs and posting root cause findings directly into code reviews. Using Gemini 2.5 Flash with prompt engineering rather than fine-tuning, the system achieved 90.14% accuracy in identifying root causes across 71 real-world failures spanning 39 teams. Since launching in production in May 2025, Auto-Diagnose has processed 224,782 test executions across 91,130 code changes from 22,962 developers, ranking 14th in helpfulness among 370 tools in Google's Critique system with 84.3% of feedback requesting continued use. The system addresses a major productivity pain point identified in Google surveys, where 38.4% of integration test failures take over an hour to diagnose and the issue ranked among the top five developer complaints in a company-wide survey of 6,059 engineers.

Auto-Diagnose System Architecture and Technical Implementation

  • ▪Auto-Diagnose post-processes the model's response into markdown with Conclusion, Investigation Steps, and Most Relevant Log Lines sections
  • ▪Gemini 2.5 Flash was not fine-tuned on Google's integration test data for Auto-Diagnose
  • ▪Auto-Diagnose calls Gemini 2.5 Flash with temperature set to 0.1 and top p set to 0.8

Performance Metrics and Production Deployment Results

  • ▪Auto-Diagnose has been in production since May 2025
  • ▪Auto-Diagnose received 517 total feedback reports from 437 distinct developers
  • ▪436 of Auto-Diagnose's feedback reports were Please fix from 370 reviewers, representing 84.3% of total feedback
  • ▪Auto-Diagnose has a Not helpful rate of 5.8% based on feedback received
  • ▪Auto-Diagnose has a helpfulness ratio of 62.96% among developer-side feedback
  • ▪Auto-Diagnose has run on 52,635 distinct failing tests across 224,782 executions on 91,130 code changes authored by 22,962 distinct developers
  • ▪Auto-Diagnose ranks number 14 in helpfulness across 370 tools that post findings to Critique, placing it in the top 3.78%
  • ▪Auto-Diagnose averages 110,617 input tokens and 5,962 output tokens per execution
  • ▪Auto-Diagnose correctly identified the root cause 90.14% of the time in a manual evaluation of 71 real-world failures spanning 39 distinct teams
  • ▪Auto-Diagnose posts findings with a p50 latency of 56 seconds and p90 latency of 346 seconds

The Integration Test Debugging Problem at Google

  • ▪A Google survey of 116 developers found that 38.4% of integration test failures take more than an hour to diagnose
  • ▪A separate Google survey of 239 respondents found that 78% of integration tests at Google are functional
  • ▪Diagnosing integration test failures showed up as one of the top five complaints in EngSat, a Google-wide survey of 6,059 developers

Perspective of Google developers using Auto-Diagnose

  • ▪370 distinct Google developers submitted Please fix feedback for Auto-Diagnose, indicating they want to continue receiving its diagnoses

Perspective of Google AI researchers behind Auto-Diagnose

  • ▪Google AI researchers validated Auto-Diagnose's 90.14% accuracy rate through manual evaluation across 39 distinct teams to ensure cross-team reliability

Perspective of Google engineering leadership

  • ▪Google engineering leadership measures Auto-Diagnose success through its top 3.78% helpfulness ranking among 370 tools in the Critique code review system
  • ▪Google engineering leadership deployed Auto-Diagnose to production in May 2025 after validating its performance across 22,962 distinct developers

2 sources

Marktechpost
Google AI Releases Auto-Diagnose: An Large Language Model LLM-Based System to Diagnose Integration Test Failures at Scale
View source article
Marktechpost
Google AI Releases Auto-Diagnose: An Large Language Model LLM-Based System to Diagnose Integration Test Failures at Scale - MarkTechPost
View source article

Featured stories

View more in Large language models (LLMs)

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

OpenAI unveils ChatGPT overhaul with shared workspaces, plugin system, and $500 monthly tier at DevDay

Sep 29, 2026 · 13 sources

Story comments

Loading comments…

Related entities

Developer tools

Related Projects

Google AI

Topics

Large language models (LLMs)AI coding assistantsAI for developers

Featured stories

View more in Large language models (LLMs)

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

OpenAI unveils ChatGPT overhaul with shared workspaces, plugin system, and $500 monthly tier at DevDay

Sep 29, 2026 · 13 sources