Work in progress · accepting task contributions

Evaluate agents for end-to-end frontier physics research.

FrontierPhysics: Benchmark how AI agents do frontier physics research.

A merged task earns 6 points, a referral 2, a review 2.At 12 points you are a co-author.

How FrontierPhysics Works

FrontierPhysics is a benchmark evaluating how AI agents do frontier physics research iteratively. We evaluate realistic research challenges with iteration loops from literature deep review to research plan implementation. Tasks come from real research problems that take at least weeks of effort for a physics PhD to do deep research and implement, and SOTA LLM agents struggle with. The tasks are evaluated with verifiable graders and per-task rubric-based reviewer agents to make sure agents are doing research in ways aligned with real frontier researchers.

Research &PlanImplement &ExperimentEvaluate &Feedback

Example task

multiplexing-ion-chain-qnet

hardtrapped-ionscalculationsimulationoptimization
tasks/multiplexing-ion-chain-qnet/task.md
    • g2_data54 files
    • assets32 files

Loading…

Run eval

claude-agent-acp · effort max · no-skill

Timeline

We will submit to ICLR first and then submit to Nature after further polish.

  • v0.1

    Get tasks merged by 31 August to join the author list of ICLR (and all future paper versions).

  • v1.0

    Get tasks merged by 31 December to join the author list of the draft submitted to Nature.

Merged, not opened — reviewing and revising a task takes days of back-and-forth, so a PR opened close to a deadline is unlikely to land in time.