O*NET for AI R&D

Categories of work that researchers/engineers at frontier AI companies do · Backfilling Epoch’s initial work

Rating scale 0–5
Methodology

Current ratings (2026-07) are Epoch AI’s published author judgments from “Toward an O*NET for AI R&D” (Denain, Kwon & Ho) — self-described “best-guess starting points,” describing typical frontier-lab practice rather than maximum elicitation.

Historical ratings (2022-07 – 2026-01, semiannual) are an independent AI-assisted reconstruction. One research agent per task traced its full trajectory under an evidence rule: every rating must cite dated evidence — no interpolation backwards from the 2026 anchor. Contemporaneous sources were preferred (lab engineering posts, system cards, measured studies such as METR and SWE-bench); retrospective testimony was admissible but flagged and down-weighted. Ratings are non-decreasing by default, and tasks that did not yet exist as jobs are rated 0 by construction.

Verification: per-category consistency review plus blind spot-checks of a 15% sample (checkers committed to independent ratings before reading the file). Eight flags were raised and individually adjudicated; the dominant error classes were future-dated evidence supporting earlier snapshots and early ratings inflated toward high 2026 anchors. Each rating’s rationale and evidence are visible by expanding its task row.

Uplift tab: a Monte Carlo model mapping rating levels to labor-time multipliers, weighting tasks by estimated time-use, aggregating Amdahl-style, and applying a Cobb-Douglas labor/compute chain (parameters after Greenblatt 2025; Davidson et al. 2026). Bands are p10–p90. No direct frontier-lab time-use data exists; the weights are triangulated from adjacent surveys and are the model’s weakest input.

Caveats: pre-2024 internal lab practice is thinly documented (low confidence is common in early snapshots); tasks new since 2022 inflate apparent growth; labs publicize automation wins, not failures.