As clinical and biological large language models proliferate, researchers are grappling with benchmarking their reasoning capabilities. A study evaluating general-purpose models like GPT-5.2 Pro, Gemini 3.0 Pro-Preview, and Claude Opus 4.6 reveals they can reconstruct 24% to 65% of literature-curated signaling network reactions. However, these AI-generated networks struggle with downstream regulatory pathways and show low functional accuracy, predicting perturbation responses at rates of only 5.8% to 33.3% compared to over 78% for manually curated models.
Aug 10, 2026 · 1 source
Aug 8, 2026 · 7 sources
Aug 7, 2026 · 3 sources
Aug 7, 2026 · 2 sources
Story comments
Loading comments…