White Circle Raises $11M Seed Round from OpenAI, Anthropic, DeepMind and Hugging Face Leaders
Paris-based White Circle raised $11 million in seed funding from Romain Huet (head of developer experience at OpenAI), Durk Kingma (OpenAI cofounder now at Anthropic), Guillaume Lample (cofounder and chief scientist at Mistral), and Thomas Wolf (cofounder and chief science officer at Hugging Face). Denis Shilov discovered a universal jailbreak prompt in late 2024. White Circle published KillBench in May 2026.
Universal jailbreak discovery
▪Anthropic invited Denis Shilov to test their models privately after the universal jailbreak post went viral.
▪Denis Shilov posted about the universal jailbreak on X and the post went viral by the next morning.
▪Denis Shilov discovered a universal jailbreak prompt in late 2024 that could bypass safety filters of every leading AI model by instructing models to behave like an API endpoint rather than a chatbot with safety rules.
White Circle seed funding
▪Ophelia Cai is a partner at Tiny VC.
▪White Circle raised $11 million in seed funding from backers including Romain Huet (head of developer experience at OpenAI), Durk Kingma (OpenAI cofounder now at Anthropic), Guillaume Lample (cofounder and chief scientist at Mistral), and Thomas Wolf (cofounder and chief science officer at Hugging Face).
▪Denis Shilov founded White Circle.
▪White Circle is based in Paris.
▪White Circle has a team of 20 people distributed across London, France, Amsterdam, and elsewhere in Europe, with almost all of them being engineers.
▪White Circle will use the seed funding to expand its team, accelerate product development, and grow its customer base across the U.S., U.K., and Europe.
Real-time AI control platform
▪White Circle builds software that sits between a company's users and its AI models, checking inputs and outputs in real time against company-specific policies.
▪White Circle's platform can catch when a model starts hallucinating, leaking sensitive data, promising refunds it cannot issue, or taking destructive actions inside a software environment.
▪White Circle's platform can flag or block requests when users try to generate malware, scams, or other prohibited content.
▪White Circle's platform has processed more than one billion API requests.
▪White Circle's platform is used by Lovable, the vibe-coding startup, as well as several fintech and legal companies.
Model provider incentives
▪The alignment tax refers to the idea that training models to be safer can sometimes make them less performant on tasks such as coding.
▪AI companies charge for input and output tokens even when a model refuses a harmful request, which reduces the financial incentive to block abuse before it reaches the model.
KillBench research
▪White Circle published KillBench in May 2026, a study that ran more than one million experiments across 15 AI models from OpenAI, Google, Anthropic, and xAI to test how systems behaved when forced to make decisions about human lives.
▪KillBench results showed models making different choices depending on attributes such as nationality, religion, body type, or phone brand, suggesting hidden biases can surface in high-stakes settings even when models appear neutral in ordinary use.
▪KillBench experiments asked models to choose between two fictional people in scenarios where one had to die, with details such as nationality, religion, body type, or phone brand changed between prompts.
▪KillBench found that the bias effect became worse when models were asked to give their answers in a format that software can easily read, such as choosing from a fixed set of options or filling out a form.
Story comments
Loading comments…