Fixed task set
The same Selected_1000 task IDs and dataset revision are used for every group. Each score is successes ÷ 1,000; failed or unknown tasks stay in the denominator.
LEADERBOARD · SEPTEMBER 2026
Models and harnesses on 1,000 econometric replication tasks.
01 / RESULTS
DSH is an independent run with technical recovery, not a paired lift. Details ↗
Loading archived results…
Results could not be loaded. Please retry or open the results archive.
No matching models.
Research snapshot · local scoring; official scorer parity unverified. Run budgets differ.
Locally scored, archived results. Official scorer parity is not yet verified. Different budgets and run protocols mean the lift is descriptive, not an equal-cost comparison.
Single-pass Agent
One code-generation call, followed by execution. No iterative agent loop.
DeepAgents
DeepAgents plans, uses tools and revises code. Up to 6 model calls, plus final execution.
DeepSeek Harness · separate run
Official DSH SDK, requested v4-pro API alias. A fixed technical recovery selection; underlying model weights are unverified.
The same Selected_1000 task IDs and dataset revision are used for every group. Each score is successes ÷ 1,000; failed or unknown tasks stay in the denominator.
Full replication is the local paper metric. The four HF-style metrics reproduce the published descriptions locally; they are not yet verified against the official scorer.
Baseline: 1 model call. DeepAgents and DSH: up to 6 model calls and 4 trial tool calls per attempt. DSH adds 87 recovery attempts; full environment and reasoning parity is unverified. DSH has no paired lift; these are not equal-cost causal estimates.
“Metric unknown” means the selected metric cannot be assessed, including failed execution or unavailable outputs. It is different from a task whose final status is unknown. Neither is dropped or counted as a success.
| Model / API alias | Protocol | Evidence |
|---|
HF-style counts are checked against archived task flags. Full-replication counts for Sol and Opus come from accepted summaries; all other displayed runs, including DSH, have per-task local flags. This refresh does not rerun or rescore any experiment.
Per-run archive files remain frozen, including their original publication metadata. This page is a new display derived from those archives. GPT-5.5 and unarchived experiments are not included in this edition.