# Verified source register

All rows established by live fetch on 2026-07-30 by two independent verification agents. Every
URL below was fetched successfully unless the row says otherwise. Builders cite the source ID
(S1, S2, ...) and never author a URL. Values are deliberately omitted: this register establishes
that a source EXISTS, is FETCHABLE, and is MACHINE-READABLE, not what it currently says.

Auditability classes: IND = independent third party runs the eval. LAB = laboratory
self-reported. AGG = aggregator of others' results. GOV = government or intergovernmental.

Verdicts: MR = live and machine-readable. MAN = live but manual or scrape-only.
DORM = dormant. RET = retired. NF = not found.

## Tier 1: spine (machine-readable, licensed, current)

| ID | Source | URL (fetched 2026-07-30) | Class | Feed | Licence | Cadence | Verdict |
|---|---|---|---|---|---|---|---|
| S1 | Epoch AI Benchmarking Hub / AI Capabilities database | https://epoch.ai/benchmarks | IND + AGG | `benchmark_data.zip` at https://epoch.ai/benchmarks/use-this-data plus `pip install epochai` (Airtable-backed). Dump observed updated 2026-07-30 | CC-BY | continuous to daily | MR |
| S2 | Epoch AI Notable AI Models (training compute) | https://epoch.ai/data/notable-ai-models | IND | https://epoch.ai/data/notable_ai_models.csv and https://epoch.ai/data/all_ai_models.csv (both 200) | CC-BY | weekly automated sweep; major models within ~2 weeks | MR |
| S3 | Epoch AI Machine Learning Hardware | https://epoch.ai/data/machine-learning-hardware (page stamped updated 2026-07-21) | IND | https://epoch.ai/data/ml_hardware.csv (200) | CC-BY | weeks | MR |
| S4 | Epoch Capabilities Index (ECI) | https://epoch.ai/eci | IND | https://epoch.ai/data/eci_benchmarks.csv (200, raw scores only). Fitting code https://github.com/epoch-research/eci-public (200, MIT) | CC-BY data, MIT code | continuous | MR |
| S5 | METR task time horizons | https://metr.org/time-horizons/ (200) | IND | https://metr.org/assets/benchmark_results_1_1.yaml and https://metr.org/assets/task_results_1_1.yaml (both re-verified 200 on 2026-07-30; note the asset path is root-relative, NOT nested under /time-horizons/); analysis code https://github.com/METR/eval-analysis-public | not stated (confirmed: the eval-analysis-public repo reports license null via the GitHub API). CORRECTION 2026-07-30: the first verification pass reported a "DO NOT TRAIN ON THIS DATA" canary on this page. A second independent fetch of the page (678,581 bytes) plus both YAML assets and the repo README found ZERO occurrences of that string. The canary claim is DISCONFIRMED for this page and is withdrawn; it may exist elsewhere in METR material, but this register no longer asserts it | irregular, self-described "limited capacity"; updates observed Mar to May 2026 | MR, low cadence |
| S6 | SWE-bench submission artefacts | https://github.com/SWE-bench/experiments (200) | AGG | results in-tree: `evaluation/<split>/<date>_<model>/all_preds.jsonl`, `metadata.yaml`, `logs/`, `trajs/`. Reasoning traces required since Jul 2024 | not stated | continuous, submission-driven | MR |
| S7 | Anthropic Economic Index | https://huggingface.co/datasets/Anthropic/EconomicIndex (200) | LAB | HF dataset repo, dated release folders. Releases observed: 2025-02-10, 2025-03-27, 2025-09-15, 2026-01-15, 2026-03-24, 2026-06-26, plus `labor_market_impacts` | data CC-BY, code MIT (stated verbatim on the page) | roughly quarterly, tightening | MR |
| S8 | Epoch AI GPU Clusters | https://epoch.ai/data/gpu_clusters.csv (listed on the fetched /data index, 500+ clusters) | IND | CSV | CC-BY | UNKNOWN | MR |
| S9 | Epoch AI open-versus-closed capability gap insight | https://epoch.ai/data-insights/open-closed-eci-gap (200) | IND | no download on the insight page; underlying scores via S4. The S4 CSV carries NO open-weight flag, so the split must be joined by hand | CC-BY | insight text lags ~2 months | MAN for the gap figure, MR for the inputs |

## Tier 2: usable, manual or constrained

| ID | Source | URL (fetched 2026-07-30) | Class | Notes | Verdict |
|---|---|---|---|---|---|
| S10 | ARC Prize (ARC-AGI-1, -2, -3) | https://arcprize.org/leaderboard (200); tasks https://github.com/arcprize/ARC-AGI-2 (200, Apache-2.0); runner https://github.com/arcprize/arc-agi-benchmarking; ARC-AGI-3 human baseline board https://arcprize.org/arc-agi/3/leaderboard (re-verified 200 on 2026-07-30) | IND | ARC-AGI-3 exists and is described on the site as unbeaten and agentic. Two withheld sets: semi-private for commercial testing, fully private for the competition. ARC Prize runs verified scores itself; community submissions are explicitly unverified. Tasks are JSON in-tree; the leaderboard SCORES are HTML only | MAN (tasks MR) |
| S11 | Humanity's Last Exam | https://agi.safe.ai/ (200) | IND | CAIS plus Scale AI. Public 2,500-question set at https://huggingface.co/datasets/cais/hle. Private held-out set used to detect overfitting. Public set finalised 2025-04-03, Nature paper 2026-01-28, HLE-Rolling dynamic fork launched 2025-10-08. Licence not stated | MAN |
| S12 | SWE-bench Pro (Scale AI) | https://labs.scale.com/leaderboard/swe_bench_pro_private (200) | IND, single vendor | 1,865 tasks over 41 repos: 731 public, 276 commercial-private drawn from 18 proprietary startup codebases, 858 held-out never published. Scale runs the evals. No data download. Page copyright line still reads 2025 | MAN |
| S13 | Artificial Analysis | https://artificialanalysis.ai/ (200); API docs https://artificialanalysis.ai/data-api/docs (200) | IND | REST API base https://artificialanalysis.ai/api/v2, `x-api-key` header (the bare base path returns 404 by design; it is not an endpoint, checked 2026-07-30). Free tier 100 req/day (headline indices plus pricing), Pro 500 req/day, Commercial adds 7/30/90-day series. Proprietary licence, attribution required, redistribution needs a contract. Runs its own evals across ~180 models and publishes confidence intervals. Independently re-implements GDPval as GDPval-AA at https://artificialanalysis.ai/evaluations/gdpval-aa (200) | MR, licence-constrained |
| S14 | Terminal-Bench | https://www.tbench.ai/leaderboard (200); repo https://github.com/laude-institute/terminal-bench (200, Apache-2.0) | AGG | Versions live: 1.0, 2.0, 2.1 current, plus Terminal-Bench Science 1.0 in development. Results are NOT in-tree; hosted on tbench.ai. Submissions must pin a Harbor dataset version | MAN |
| S15 | GPQA | https://github.com/idavidrein/gpqa (200, MIT) | dataset only | No first-party leaderboard exists. Dataset at https://huggingface.co/datasets/idavidrein/gpqa, password-gated on GitHub to resist scraping. Diamond split size not verified today. Use S1 as the score feed | MAN |
| S16 | Stanford HELM | https://github.com/stanford-crfm/helm (200) | IND, strongest reproducibility | Verified status quote: "HELM entered maintenance mode on June 1, 2026." Boards: HELM Capabilities, HELM Safety, VHELM. CRFM runs the models itself with prompt-level transparency. https://crfm.stanford.edu/helm/ is a JS shell and yielded no fetchable content | MAN, trending DORM |
| S17 | Stanford HAI AI Index | https://hai.stanford.edu/ai-index (200); 2026 edition https://hai.stanford.edu/ai-index/2026-ai-index-report | AGG | Annual only, editions 2017 to 2026. Public dataset download referenced but the exact CSV URL was not established | MAN |
| S18 | LMArena / Arena | https://lmarena.ai/leaderboard now 301-redirects to https://arena.ai/leaderboard (redirect observed today) | IND, human preference | Operator now "Arena Intelligence". Boards include Agent, Text, WebDev, Vision, Document, Search. Measures preference, not capability. Vote data and the Elo pipeline are operator-controlled; pre-release anonymous testing advantages labs that iterate against it | MAN |
| S19 | OSWorld / OSWorld-Verified | https://os-world.github.io/ 301-redirects to http://osworld-v1.xlang.ai/ (redirect observed today; note the v1 host and plain HTTP) | IND | XLANG Lab with Salesforce Research, CMU, Waterloo. OSWorld-Verified announced 2025-07-28. Code https://github.com/xlang-ai/OSWorld. Last dated site update 2025-07-28, roughly a year stale. Site content CC BY-SA 4.0 | MAN, possibly superseded |
| S20 | GDPval (original) | https://openai.com/index/gdpval/ (returns HTTP 403 with cf-mitigated: challenge under a desktop browser user agent, re-checked 2026-07-30: the host serves a Cloudflare bot challenge, so the page cannot be fetched programmatically and was confirmed by search instead); dataset https://huggingface.co/datasets/openai/gdpval; grading UI https://evals.openai.com/gdpval/leaderboard (fetched, JS shell, no readable content) | LAB | Lab self-reported on its own economic-value framing. The independently runnable version is S13's GDPval-AA, whose numbers are not OpenAI's numbers | MAN |
| S21 | BEHAVIOR Challenge 2026 | https://behavior.stanford.edu/challenge/index.html (200); leaderboard https://huggingface.co/spaces/behavior-1k/2026-challenge-leaderboard (curl 200) | IND academic | Second edition. Launch 2026-07-02, submissions close 2026-10-16, winners 2026-11-04. Code https://github.com/StanfordVL/BEHAVIOR-1K. Simulation, not unstructured reality. No results CSV or JSON; the board is a Gradio Space | MAN |
| S22 | Epoch AI economic-value-benchmark critique | https://epoch.ai/publications/what-do-economic-value-benchmarks-tell-us (200, published 2026-02-13) | IND | Establishes that economic-value benchmarks measure isolated tasks rather than end-to-end work situated in real context. Public task counts stated: RLI 10 of 240, APEX-Agents 480 of 480, GDPval 220 of 1320 | MAN, qualitative |
| S23 | Manifold Markets | https://api.manifold.markets/v0/search-markets?term=AGI&limit=2 returned HTTP 200, application/json | AGG, play money | Documented REST API, no key required for reads, full order-book history. Play-money incentives and thin liquidity make individual AGI markets poorly calibrated | MR |
| S24 | Epoch AI AI Companies / Chip Owners / Chip Sales | https://epoch.ai/data/ai_companies.zip, https://epoch.ai/data/ai_chip_owners.zip, https://epoch.ai/data/ai_chip_sales.zip (all listed on the fetched /data index) | IND | Revenue, funding, staff, compute, and compute concentration by actor | MR |

## Tier 3: blocked, gated, dormant, or dead. Named so the spec never cites them as live.

| ID | Source | Status established 2026-07-30 |
|---|---|---|
| S25 | Metaculus | API GATED. The AGI question page returned 403 to fetch. Three API probes (`/api/posts/5121/`, `/api2/questions/5121/`, `/api/posts/?search=AGI`) all returned HTTP 403 with the body "Permission Error: The API is only available to authenticated users." An API exists but now requires a token. No question ID or community median is cited here because none could be confirmed |
| S26 | Papers with Code | RETIRED. https://paperswithcode.com/ returns 302 to https://huggingface.co/papers/trending, which has no leaderboard or SOTA function. Any historical SOTA curve must come from S1 instead |
| S27 | Hugging Face Open LLM Leaderboard | RETIRED / archived. The Space still renders but HF publishes an "Archived Open LLM Leaderboard (2024-2025)" collection and an archive doc page, and the project's own discussion thread announces its end. Historic results at https://huggingface.co/datasets/open-llm-leaderboard-old/results. Open-weights only; never covered closed frontier models |
| S28 | LiveBench | LIKELY DORMANT. https://github.com/LiveBench/LiveBench (200) promises monthly refreshed questions, but the README states the current release is 2025-04-25, so the monthly promise appears unmet. Results on HF at livebench/model_answer and livebench/model_judgment. https://livebench.ai/ is a JS shell |
| S29 | LiveCodeBench | UNFETCHABLE. https://livecodebench.github.io/leaderboard.html renders JS only and showed "Loading..." with no data extractable on 2026-07-30 |
| S30 | RoboArena | FEED NOT FOUND. https://robo-arena.github.io/ returns 200 but serves a 466-byte SPA shell; `/leaderboard` and `/api/leaderboard` both 404. Paper https://arxiv.org/abs/2506.18123. Crowdsourced double-blind pairwise comparisons on DROID hardware, so the Elo has no fixed zero across time |
| S31 | OpenAI usage data portal | BLOCKED. https://openai.com/signals/data/ returned 403 to both WebFetch and curl. The underlying paper is NBER Working Paper 34255, "How People Use ChatGPT", Chatterji et al., https://www.nber.org/papers/w34255 (re-verified 200 on 2026-07-30) |
| S32 | US Census BTOS | FRAGILE. https://www.census.gov/hfp/btos/data returns 200 but is a JS SPA with no extractable content. CRITICAL: the Census server returns HTTP 200 for fabricated filenames under `/hfp/btos/downloads/`, including one the verifier invented, so it soft-404s. No BTOS download URL may be trusted on status code alone; the file must be opened and its contents checked |
| S33 | OECD.AI | PARTIALLY DORMANT. https://oecd.ai/en/data (200) carries the banner "This section is currently being updated and may not reflect the most recent data." No API, CSV, licence, or cadence established |
| S34 | Eurostat `isoc_eb_ai` | SEARCH-ONLY, not fetched. Enterprise AI adoption series. Eurostat operates a documented dissemination API (inferred, not verified) |
| S35 | Epoch AI inference-price and US-versus-China insights | both re-verified 200 on 2026-07-30: https://epoch.ai/data-insights/llm-inference-price-trends and https://epoch.ai/data-insights/us-vs-china-eci. Both are believed CC-BY per site policy. No longer pending: both pages fetch. The underlying data remains reachable via S1 and S4, which is the preferred route because the insight text lags the data |
| S36 | Epoch AI algorithmic-efficiency work | NO MAINTAINED SERIES. https://epoch.ai/topics/software-progress (200) carries publications only: "Algorithmic progress in language models" (2024-03-12) and "Revisiting algorithmic progress" (2022-12-12). No feed. Any efficiency trend must be re-derived from papers that are now one to four years old |
| S37 | FrontierMath | https://epoch.ai/frontiermath (200). Epoch-run, Tiers 1 to 4 unpublished plus an Open Problems set, held out. Problem data is not public, so contamination protection comes at the cost of external auditability. Scores reachable via S1 |

## Established non-existence (searched 2026-07-30, not found)

| Gap | What was searched | Consequence for the spec |
|---|---|---|
| G1 | No live public leaderboard for model CALIBRATION or run-to-run variance. Partial coverage only: HELM includes calibration in its metric matrix (S16, now in maintenance mode) and S13 publishes confidence intervals | Any reliability indicator must be self-computed or declared unavailable. It cannot be sourced |
| G2 | No live eval or leaderboard tracking generation-over-generation REGRESSION | Must be self-computed by diffing successive S1 dumps. This is cheap and is the recommended approach |
| G3 | No published QUANTIFICATION of the benchmark-to-real-world gap. S22 argues one exists; METR's own note at https://metr.org/notes/2026-01-22-time-horizon-limitations/ states they do not know its size (re-verified 200 on 2026-07-30) | The gap can be framed and bounded but not stated as a number. Any numeric gap claim in the spec would be invented |
| G4 | No robotics benchmark with both unstructured-environment tasks and a machine-readable time series | Embodied competence is a supporting indicator with annual granularity at best (S21) |
| G5 | No dedicated China-specific capability tracker with a machine-readable feed. S4's explorer carries a country filter | Geopolitical comparison is derivable from S4 but is largely collinear with the open-weight gap, since most leading Chinese models are open-weight. It is not independent evidence |

## Confounders established during verification, to be carried into the failure-mode register

1. S4's ECI is a latent-trait fit that is RE-ANCHORED as benchmarks saturate, so historical
   values move retroactively and any cached figure silently drifts. Named anchors observed:
   Claude 3.5 Sonnet = 130, GPT-5 = 150.
2. S2's training-compute figures are Epoch ESTIMATES derived from parameter and token counts,
   not laboratory disclosures, with wide error bars that are revisable without notice.
3. S3 records RELEASE price, not realised or rental price, so any FLOP-per-dollar series built
   on it ignores the cost frontier labs actually pay.
4. S7 measures Claude's own user mix, so movements reflect Anthropic's product and customer
   changes as much as any change in the economy.
5. Source mortality is empirically high. Of the sources checked, four are retired, dormant, or
   in maintenance mode (S16, S26, S27, S28) and two have moved host or frozen a version line
   (S18 lmarena to arena, S19 pinned to a v1 host). A tracker that hardcodes URLs will break.
6. Version churn breaks longitudinal comparability: S14 has moved 1.0 to 2.0 to 2.1, and S6 has
   five splits. Every observation must pin its benchmark version.
7. S32 soft-404s with HTTP 200 on fabricated filenames, which means status-code validation alone
   is insufficient anywhere in the pipeline.
8. This register applies its scepticism unevenly, and the gap is named here rather than left for a
   reader to find. S5 is awarded the strongest auditability class (IND) and carries all three horizon
   indicators, which is the sole rendering of the "extended horizons" clause of the measurement target.
   It receives no examination of funding, of its pre-deployment access relationships with the
   laboratories whose systems it times, or of the contested representativeness of its task suite and
   the log-linear form of its fit, while S13 and S37 are both examined on exactly those grounds. The
   register states no wrongdoing and has no evidence of any; it records that one supplier holding up
   one entire axis was not held to the standard the others were. Treat S5 claims as MEDIUM auditability
   pending that examination, not IND.
9. Six of the ten contextual indicators and five of the nine launch indicators resolve to a single
   operator (Epoch AI, S1 to S4, S8, S9, S24). FM-25 rates this CRITICAL and mitigates it by naming a
   fallback per indicator, but for A1, A2, D4, G1i, G3i and H2 no non-Epoch feed of equivalent type
   exists anywhere in this register. That mitigation is therefore an aspiration, and this register is
   the evidence against it.
