Research
Environments are the new data.
We publish the environments and benchmarks we use to evaluate agents on real work — papers, datasets, and live sites. SkillsBench, ClawsBench, PostTrain, and FrontierPhysics. The Discord is 1,500 members.
01· BenchmarkSkillsBench
The first benchmark for whether procedural skills — instructions, scripts, references an agent loads on demand — make agents better at real work. 86 tasks, 11 domains.
02· EnvironmentClawsBench
Five mock workplaces — Gmail, Calendar, Drive, Docs, Slack — wire-compatible with the upstream `gws` and Slack APIs. Production agents and skills run unchanged against a safety-evaluable replica.
03· ArenaPostTrain
Contribute an environment. We post-train and score what generalizes.
04· ScienceFrontierPhysics
Evaluating agents for end-to-end frontier physics research.