LEADERBOARD · SEPTEMBER 2026

InferenceNet Challenge Leaderboard

Models and harnesses on 1,000 econometric replication tasks.

model aliases experiment groups tasks per group task records

01 / RESULTS

Agent & Harness leaderboard

Research results · provisional
Download CSV ↓

Loading archived results…

Every score uses all 1,000 tasks. Unknown and invalid outcomes earn no successes.

Research snapshot · local scoring; official scorer parity unverified. Run budgets differ.

Evaluation protocol, definitions & evidence

Locally scored, archived results. Official scorer parity is not yet verified. Different budgets and run protocols mean the lift is descriptive, not an equal-cost comparison.

Single-pass Agent

One code-generation call, followed by execution. No iterative agent loop.

DeepAgents

DeepAgents plans, uses tools and revises code. Up to 6 model calls, plus final execution.

DeepSeek Harness · separate run

Official DSH SDK, requested v4-pro API alias. A fixed technical recovery selection; underlying model weights are unverified.

Fixed task set

The same Selected_1000 task IDs and dataset revision are used for every group. Each score is successes ÷ 1,000; failed or unknown tasks stay in the denominator.

Two scoring profiles

Full replication is the local paper metric. The four HF-style metrics reproduce the published descriptions locally; they are not yet verified against the official scorer.

Read lift with the budgets

Baseline: 1 model call. DeepAgents and DSH: up to 6 model calls and 4 trial tool calls per attempt. DSH adds 87 recovery attempts; full environment and reasoning parity is unverified. DSH has no paired lift; these are not equal-cost causal estimates.

Evidence status & metric definitions

“Metric unknown” means the selected metric cannot be assessed, including failed execution or unavailable outputs. It is different from a task whose final status is unknown. Neither is dropped or counted as a success.

    Status from each accepted archive
    Model / API aliasProtocolEvidence

    HF-style counts are checked against archived task flags. Full-replication counts for Sol and Opus come from accepted summaries; all other displayed runs, including DSH, have per-task local flags. This refresh does not rerun or rescore any experiment.

    Per-run archive files remain frozen, including their original publication metadata. This page is a new display derived from those archives. GPT-5.5 and unarchived experiments are not included in this edition.

    Historical leaderboard ↗