Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Snorkel AI Highlights First Wave of Open Benchmarks Grants Projects
00

Snorkel AI Highlights First Wave of Open Benchmarks Grants Projects

Jul 24, 2026

Snorkel AI highlighted the first wave of projects supported through its $3 million Open Benchmarks Grants program, launched in February 2026. The company said the initiative addresses a gap between rapid AI advancement and rigorous performance measurement. Supported projects developed with research institutions and communities include Frontier-Bench, Agents' Last Exam, OSWorld 2.0, Continual Learning Bench, SlopCode Bench, and Terminal-Bench 2.1. Separately from the grants program, Snorkel AI led the development of Senior SWE-Bench with research teams at Princeton University and the University of Wisconsin–Madison.

Open Benchmarks Grants recipients

  • ▪The Open Benchmarks Grants program provides selected research teams with funding, expert data development support, research and engineering collaboration, and platform resources.
  • ▪Snorkel AI launched the Open Benchmarks Grants program in February 2026 as a $3 million commitment to support open-source datasets, benchmarks, and evaluation research.

AI evaluation challenges

  • ▪Fred Sala, a member of the Open Benchmarks Grants steering committee, said the projects address evaluation challenges including complex environments, huge autonomy horizons, and rich, sophisticated outputs.
  • ▪Snorkel AI described AI systems as advancing faster than the field's ability to rigorously measure their performance on realistic, consequential work.

Funded benchmark projects

  • ▪OSWorld 2.0, developed with XLANG Lab, evaluates computer-use agents on 108 long-horizon workflows across 31 self-hosted web environments and professional desktop applications.
  • ▪Frontier-Bench, developed with Laude Institute and the Harbor community, evaluates agents in terminal environments under continuous adversarial review.
  • ▪Agents' Last Exam, developed with UC Berkeley RDI and the RDI Foundation, evaluates AI agents on long-horizon, economically valuable professional workflows.
  • ▪Terminal-Bench Science is in development with support from Open Benchmarks Grants to extend the Terminal-Bench framework to computational research workflows.
  • ▪SlopCode Bench, developed with the University of Wisconsin–Madison, measures how code quality degrades as coding agents repeatedly modify and extend their own solutions.
  • ▪Continual Learning Bench, developed with UC Berkeley SkyLab and the University of Wisconsin–Madison, measures whether AI agents improve across sequential, stateful tasks.
  • ▪Terminal-Bench 2.1, developed with Stanford University, Laude Institute, and the Harbor community, evaluates AI agents on challenging work in terminal environments.

Senior SWE-Bench development

  • ▪Snorkel AI led the development of Senior SWE-Bench with research teams at Princeton University and the University of Wisconsin–Madison.
  • ▪Senior SWE-Bench evaluates coding agents on senior-level engineering work, including implementing features, investigating bugs, and following existing codebase conventions.

Program partners

  • ▪The Open Benchmarks Grants program was established with support from Hugging Face, Prime Intellect, Together AI, Factory, Harbor, and PyTorch.
  • ▪Fred Sala, an assistant professor at the University of Wisconsin–Madison, serves as a member of the Open Benchmarks Grants steering committee.

Application process

  • ▪Applications for the Open Benchmarks Grants program remain open and are reviewed on a rolling basis.
  • ▪The Open Benchmarks Grants program has received hundreds of applications from researchers, labs, and engineers since its launch in February 2026.

1 source

Prnewswire
Snorkel AI Highlights First Wave of Open Benchmarks Grants Projects
View source article

Featured stories

View more in AI research & benchmarks

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

FTC opens investigation into OpenAI and Anthropic over consumer protection

Sep 30, 2026 · 7 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

Accelevation raises $540 million in US IPO

Sep 29, 2026 · 2 sources

Story comments

Loading comments…

Topics

AI research & benchmarksAI startupsModel evaluation methodologyOpen-source AI

Featured stories

View more in AI research & benchmarks

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

FTC opens investigation into OpenAI and Anthropic over consumer protection

Sep 30, 2026 · 7 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

Accelevation raises $540 million in US IPO

Sep 29, 2026 · 2 sources