WorkBuddy Bench is an openly released, 260-task evaluation suite for agentic work across repository engineering, front-end development, office workflows, and security. Its tasks are written as realistic, deliberately underspecified requests and evaluated through track-specific verifiers in isolated workspaces. The benchmark’s central empirical finding is that capability depends on the complete model–harness configuration: changing the agent harness can materially alter both scores and model rankings. The release is reproducible and auditable, but remains exposed to post-release contamination, model-judge bias in Web and Office, Python-heavy Code coverage, and version-specific serving and harness effects.
Story comments
Loading comments…