As clinical and biological large language models proliferate, researchers are grappling with benchmarking their reasoning capabilities. A study evaluating general-purpose models like GPT-5.2 Pro, Gemini 3.0 Pro-Preview, and Claude Opus 4.6 reveals they can reconstruct 24% to 65% of literature-curated signaling network reactions. However, these AI-generated networks struggle with downstream regulatory pathways and show low functional accuracy, predicting perturbation responses at rates of only 5.8% to 33.3% compared to over 78% for manually curated models.
Sep 29, 2026 · 14 sources
Sep 29, 2026 · 14 sources
Sep 29, 2026 · 13 sources
Sep 29, 2026 · 7 sources
Story comments
Loading comments…