Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
OpenAI's Codex wins AI security race at Devcon, DeepSeek finishes second
00

OpenAI's Codex wins AI security race at Devcon, DeepSeek finishes second

Sep 3, 2026

In an AI security race organized by Austin Griffith of the Ethereum Foundation on September 3, 2026, OpenAI's Codex (running GPT-5.5) was the only model to clear all 12 Solidity challenges, taking the top three spots. The medium reasoning setting proved fastest at 40 minutes and 7 seconds. DeepSeek V4 Pro finished fourth, capturing 11 flags for just $1.45 in compute, significantly outperforming Anthropic's Claude Opus 4.8, which cost $7.61 for 10 flags.

Devcon Solidity CTF AI benchmark

  • ▪The Solidity challenges used in the AI security race on September 3, 2026, were originally designed as a capture-the-flag course for human developers at Ethereum's Devcon conferences in Bangkok and Buenos Aires.
  • ▪Austin Griffith of the Ethereum Foundation ran an AI security race on September 3, 2026, featuring ten AI coding agents tackling 12 Solidity challenges.

OpenAI Codex performance results

  • ▪OpenAI's Codex took the top three finishing places in the September 3, 2026 AI security race.
  • ▪OpenAI's Codex, running GPT-5.5, was the only model to clear all 12 Solidity challenges in the September 3, 2026 security race.

GPT reasoning setting comparison

  • ▪In the September 3, 2026 AI security race, OpenAI's Codex running on an extra-high reasoning setting finished last among the three Codex entries in 50 minutes and 26 seconds, consuming nearly a third more tokens than the medium setting.
  • ▪In the September 3, 2026 AI security race, OpenAI's Codex running on a medium reasoning setting cleared all 12 flags fastest in 40 minutes and 7 seconds.

DeepSeek V4 Pro cost efficiency

  • ▪Anthropic's Claude Opus 4.8 finished behind DeepSeek V4 Pro in the September 3, 2026 AI security race, capturing 10 flags at a compute cost of $7.61.
  • ▪DeepSeek V4 Pro, an open-weight Chinese model, finished in fourth place in the September 3, 2026 AI security race by capturing 11 of 12 flags at a compute cost of $1.45.

Open-weight Chinese model performance

  • ▪Among other open-weight Chinese models in the September 3, 2026 AI security race, GLM 5.3 captured six flags.
  • ▪In the September 3, 2026 AI security race, open-weight Chinese models Kimi K3 and Qwen captured two flags each.

AI coding evaluation methodology

  • ▪Austin Griffith of the Ethereum Foundation and developer collective BuidlGuidl published the system prompts, human interventions, and standings of the September 3, 2026 AI security race, describing the exercise as a transparent single-run evaluation rather than a universal model ranking.
  • ▪During the September 3, 2026 AI security race, each AI coding agent was provided with an isolated instance and its own wallet, with a flag capture counted only when the mint landed onchain.

1 source

Unchainedcrypto
Only Codex Finished Austin Griffith's AI Security Race, DeepSeek Close Behind - Unchained
View source article

Featured stories

View more in Smart contracts

OpenAI agents hijacked German website in previously undisclosed breakout incident

Sep 3, 2026 · 5 sources

Debian votes to allow AI-generated code in Linux distribution

Aug 30, 2026 · 2 sources

Google AI introduces EnvHarness to transform static agent environments into adaptive training systems

Aug 30, 2026 · 1 source

Anthropic remains Pentagon supply chain risk despite Lutnick comments

Sep 3, 2026 · 2 sources

Story comments

Loading comments…

Related entities

Devcon

People Involved

Austin Griffith

Related Projects

CodexOpenAIDeepSeek

Topics

Smart contractsAI agentsAI research & benchmarksAI securityAI coding assistantsOpen source AI ecosystems & communities

Featured stories

View more in Smart contracts

OpenAI agents hijacked German website in previously undisclosed breakout incident

Sep 3, 2026 · 5 sources

Debian votes to allow AI-generated code in Linux distribution

Aug 30, 2026 · 2 sources

Google AI introduces EnvHarness to transform static agent environments into adaptive training systems

Aug 30, 2026 · 1 source

Anthropic remains Pentagon supply chain risk despite Lutnick comments

Sep 3, 2026 · 2 sources