DATA · SELECTED_1000

The benchmark, task by task.

Each task names an outcome, a treatment, controls, an estimation method and a data file. The goal is to recover one reported effect: its coefficient, standard error and p-value.

Dataset downloads require a Hugging Face account with access. Open the dataset to request access before using the download link.

1,000 fixed evaluation tasks
15 journals in this subset
286 distinct article titles
5 estimation classes

Selected_1000 · Counts from the pinned task list, revision 59f9512.

01 / DISTRIBUTION

Different research questions. Different demands.

Estimation methods

Five classes · 1,000 tasks

Loading distributions…

OLS includes panel OLS; IV includes IV-2SLS. Other combines Probit, Logit, negative binomial, LMM, PSM and Cox models.

Journal coverage

The five largest sources; all 15 below

Loading distributions…
All 15 journals

Percentages use all 1,000 tasks. These distributions are recomputed from the pinned CSV, not from the moving dataset main branch.

02 / TASK FORMAT

A clear input and output contract.

What the agent receives

Target and explanatory variables, controls, estimator, analysis requirements, and mounted source data.

What the agent produces

An executable program and a result JSON containing coefficient, standard_error and p_value.

What the evaluator keeps separate

Reference answers and replication programs are kept outside the model input during evaluation. Dataset downloads can contain references.

03 / EXAMPLE

One task, in context.

TASK 0011 · OLS

Economic uncertainty and the “two Spains”

Recover the relationship between socio-economic conflict and policy uncertainty, with three controls and month-clustered standard errors.

See the agent walkthrough →
y
EPU0month_simsn_w
x
Wscmonth_simsn_w
Data
data_np.dta
Sample
1905–1945 · 977 estimation observations

04 / SOURCE

A fixed, inspectable snapshot.

Dataset
CamoAiLab/InferenceNet ↗
Pinned revision
59f9512a38e594528807744214a60ee00367434e
Task list
Selected_1000/1000_new.csv
Subset labels
1,000 Stata-labelled tasks; 14 raw method labels folded into five classes. Article count is the number of distinct title strings.