AGI Progress Tracker: conceptual framework and minimum viable specification


Version 0.1.0 (draft for review). Compiled 2026-07-30. Maintainer: Victor del Rosal. Every source named in this document was fetched and checked on 2026-07-30. The register of those checks, including what was found dead, gated, or dormant, is a companion file (SOURCES.md) and is part of the specification, not an appendix to it.

#

This document specifies an instrument, not a forecast. It contains no AGI date, no probability of AGI, and no claim about whether AGI is near. It is built so that it would register stagnation during a boom and acceleration during a bust, which is the only property that makes such an instrument worth maintaining.

There is no agreed definition of AGI, and this tracker does not supply one. It measures progress toward a stipulated operational target and says so on every screen. The target: growth in the reliability-weighted, novelty-discounted volume of cognitive tasks a system completes without human assistance over extended horizons, per unit of resource consumed. Four constitutive axes (breadth, depth, reliability, horizon), one contamination discount (novelty), one denominator (resource).

What counts, and what is only context. Raw capability, reliability, generality and autonomy count as progress. Resources required is the denominator, contextual but load-bearing, because it is what separates capability that was earned from capability that was purchased. Economic usefulness, deployment scale, and safety or controllability are tracked separately and are never summed into a capability figure. A safer model is not a more general one, and a widely deployed model is not a more capable one.

The headline is two numbers and a flag, never one number. POSITION is the count of pre-registered capability thresholds met, in native units, never a normalised level. It is floored by the weakest core dimension so that one hot dimension cannot carry the headline, and the floor is computed only over dimensions that actually have a live, non-stale threshold and a published dispersion estimate. The number of dimensions entering the floor is printed inside the headline string itself, because a floor that fell because a source went quiet must never be readable as a capability claim. RATE is a doubling-time estimate with a published interval. The FLAG is a primary state (CLEAN, REGRESSION, SATURATION, INSTRUMENT-FAILURE, STALE) plus a set of co-occurring conditions that are never suppressed, so an instrument failure cannot hide the staleness that caused it. The whole thing is emitted as one atomic citation string, and the API and the social preview image serve only that string, because a headline that can be quoted without its qualifiers will be. The Kurzweil-style log chart is retained as the rate view only. A clock face is assessed and rejected in section 1: minutes-to-midnight is expert judgment wearing the visual grammar of measurement, and it implies a metric distance that does not exist.

What this instrument is not, stated in one place before anything else. A hostile expert review of the first draft returned the judgment that the document was candid clause by clause and never assembled the clauses into the verdict they imply. That verdict is assembled here, at the front, because a reader who stops after one page should leave with it. At launch: the tracker makes ONE original measurement (a deliberately small rerun-variance probe, section 9), and every other reading it publishes is a restatement of two other organisations' work, since five of the nine launch indicators resolve to Epoch AI and three resolve to METR. Of the four axes its own target sentence calls constitutive, TWO ARE UNINSTRUMENTED at launch: breadth has no indicator of its own and is read across other dimensions because no task-family census exists, and reliability carries NO threshold at all at launch, resting on one deliberately small probe and one regression counter, so the word "reliability-weighted" in the target sentence is not yet discharged by anything. The resource denominator named in the definition is published beside the headline and not inside it. The position figure rests on four to six thresholds chosen by one person; pre-registration makes that choice auditable and does not make it any less determinative. The honest summary is therefore that this instrument's early rigour is a property of its method and its provenance discipline, not yet of its evidentiary base, and its most defensible contributions on day one are the source register, the failure-mode register, and the regression counter. Anyone citing the headline before stage 3 is citing a scorecard laid over Epoch AI and METR.

Twenty-one indicators across nine dimensions, of which nine form the shippable MVP. Two things are unusual about the roster and both are deliberate. First, several indicators can only be computed by the maintainer because no live source publishes them: verification established that no public leaderboard tracks model calibration, run-to-run variance, or generation-over-generation regression anywhere. Rather than record all three as gaps, the tracker builds the two that are cheap: the regression count is computed by diffing retained data snapshots, and run-to-run variance is measured by a deliberately small published probe. Calibration remains a declared gap with a costed construction. Nothing is filled with a proxy. Second, the roster includes indicators whose job is to detect the instrument failing rather than the systems improving: benchmark saturation half-life and the source-freshness probe exist so that the tracker can tell "progress stopped" from "our ruler melted". A tracker that cannot separate those two is worthless in precisely the moment it matters.

The four layers are separated by construction, not by convention. Data, methodology, interpretation, and forecasting live in different directories, and the dashboard carries a toggle that removes interpretation and forecasting entirely. The test of whether the separation is real is that the page with interpretation switched off is still a complete, readable document. No editorial or predictive language appears in the data or methodology layers.

The findings from source verification that most shaped the design, all established by live fetch on 2026-07-30:

  1. One source family is strong enough to be the spine. Epoch AI publishes machine-readable,

    CC-BY licensed capability, model, hardware and cluster data with a Python client, refreshed at a cadence measured in days (S1 to S4, S8, S9, S24).

  2. Source mortality is high and was not on the original risk list. Of the sources checked, four

    are retired, archived, dormant, or in maintenance mode, and two have moved host or frozen a version line. A dead source silently freezes an indicator, and a frozen indicator reads as a plateau. This became failure mode FM-27.

  3. A latent-factor headline already exists in the wild and demonstrates the method's central

    danger. The Epoch Capabilities Index (S4) re-anchors as benchmarks saturate, so historical values move retroactively. Any tracker consuming it must freeze and archive each vintage and never overwrite history.

  4. Status codes are not validity checks. One government source returns HTTP 200 for filenames

    that do not exist, including one the verifier invented (S32). Content and hash validation is therefore mandatory in the pipeline, not optional.

  5. This register applies its scepticism unevenly and says so: METR (S5) carries the whole horizon

    axis and was awarded the strongest auditability class without the funding and access scrutiny applied to other suppliers. It is downgraded to MEDIUM pending that examination.

  6. Three things the field simply does not publish: a quantification of the benchmark-to-reality

    gap, a calibration or reliability time series, and a robotics benchmark with both unstructured tasks and a machine-readable feed. These are gaps G3, G1 and G4 respectively, and are stated as gaps rather than estimated. Two further gaps are recorded in the register: G2, generation-over-generation regression, which is self-computed and shipped rather than absent, and G5, a China-specific capability feed, which is derivable but largely collinear with the open-weight lag.

How to read this document. Sections 1 to 3 define what is being measured and why. Section 4 is the reference specification, sixteen fields for each of twenty-one indicators, and is meant to be consulted rather than read. Section 5 records what was deliberately excluded, which is where a sceptical reader should look first. Sections 6 and 7 compare four candidate headline methods and specify the sensitivity analysis that must run on every release. Section 8 is the failure-mode register, twenty-seven modes with a detector and a verdict for each. Sections 9 to 11 are the buildable part: the MVP, the schema, and the governance. Section 12 comes last on purpose, because the visual form is downstream of the measurement, and designing it first is how trackers end up with a beautiful number nobody can defend.

Provenance conventions used throughout. Sources are cited by register ID (S1 to S37) and established non-existence by gap ID (G1 to G5); URLs live only in SOURCES.md so that link rot has one place to be fixed. Every quantitative claim carries a verification class: independently verified, self-reported, estimated, expert judgment, inferred, or missing. Confidence bands are HIGH, MEDIUM or LOW and decay across horizons, a convention inherited from the AGI-handbook project. Every observation carries a last-verified date and is flagged STALE after ninety days, a convention inherited from the future-of-work knowledge base.

Status of this draft. The framework, the indicator specifications, the methods, the register and the schema are complete and internally consistent. No indicator has yet been populated with data: no value in this document is a measurement, and the sample record in section 10 is explicitly flagged as a schema illustration. Building the MVP is the next step, not a completed one.

#

1.1 The tracker does not define AGI

There is no agreed definition of artificial general intelligence. Competing definitions disagree on whether the target is a capability profile, an economic threshold, a learning mechanism, or a comparison against a human population, and they disagree in ways that change which observations count as evidence. This tracker supplies no definition and adjudicates between none. It measures progress toward a stipulated operational target, states that target in full, and labels every reading as a reading against that stipulation rather than against AGI. Readers who reject the stipulation should read the component axes directly, which is why the axes are published unbundled (section 3).

1.2 The stipulated target

Progress toward AGI is measured as growth in the reliability-weighted, novelty-discounted volume of cognitive tasks a system completes without human assistance over extended horizons, per unit of resource consumed. That decomposes into four constitutive axes (BREADTH, DEPTH, RELIABILITY, HORIZON), one contamination discount (NOVELTY), and one denominator (RESOURCE). Nothing else enters.

1.3 Operational rendering of each axis

BREADTH. The count of distinct task families in which a system clears a competence threshold pinned before the observation, from a fixed versioned taxonomy. Moved when a system clears the threshold in a family no prior system cleared. Falsified if the gain is confined to families already cleared (that is depth), if the taxonomy widened between observations, or if the new family is a relabelling. No verified source publishes a task-family census (verification class: missing), so breadth is read across dimensions A, B, E and F rather than from one indicator.

DEPTH. Residual headroom on evaluations no system has saturated, and the rate at which it closes. Moved when headroom narrows at a pinned benchmark version. Falsified by a ceiling effect on an already-saturated evaluation, by a version change, or if the gain appears only on the public split of a benchmark that also maintains a held-out split (S10, S11, S12, S37; verification class: independently verified that held-out splits exist).

RELIABILITY. Stability of output across repeated attempts at one task, plus agreement between stated confidence and realised accuracy. Moved by a narrowing first-attempt to best-of-k spread, a fall in rerun variance under a fixed seed policy, or a fall in calibration error. Falsified if the improvement came from sampling or scaffold changes between observations, or if the best run improved while median and worst did not. No live source publishes calibration or rerun variance (G1; verification class: missing), so this axis is self-computed or declared absent.

HORIZON. Duration of unassisted work completed at a stated success rate, in units of human expert time on the same task. Moved when the duration at a fixed success rate rises at a pinned task-suite version (S5; verification class: independently verified that the series exists, at a self-described low cadence). Falsified if the success rate was lowered to buy duration, if the suite changed, or if human assistance entered the run.

NOVELTY (discount, not an axis). Ratio of performance on tasks the system could not have seen to performance on tasks it might have. Moved when that ratio changes at a fixed pair of splits. It is not itself a capability gain; it is the factor deflating a claimed one. Falsified as a real change if the held-out split was refreshed, resized, or partially released between observations.

RESOURCE (denominator). Compute and money to reach a fixed capability threshold, and inference cost per solved task at fixed reliability. Moved when the same threshold is reached for less. Falsified if the threshold moved, if the price series records list rather than realised cost (S3 records release price; verification class: independently verified as a documented limitation), or if an external compute estimate is compared against a laboratory disclosure (S2; class: estimated).

1.4 The eight-way distinction, resolved

Raw capability COUNTS. It is the object the stipulation names. Capability without reliability is a demonstration rather than a competence, which is why it enters weighted rather than raw; excluding it would leave the instrument measuring the conditions of capability and never capability itself.

Reliability COUNTS. A system that solves a task one time in ten holds a lottery ticket on the task, not the task. Since the target is a volume of tasks completed without human assistance, and unreliable completion requires a human to check and retry, reliability is constitutive of the target rather than a quality attribute layered over it.

Generality COUNTS. Generality is the breadth axis under another name, and it is what separates the target from narrow superhuman performance. Without it, saturating one domain would register as progress equivalent to acquiring ten.

Autonomy COUNTS. The stipulation says "without human assistance over extended horizons", so autonomy is the condition under which the other axes are read rather than an added virtue. A capability that appears only under supervision is still measured; it is measured as a lower reading.

Resources required is the DENOMINATOR, contextual but load-bearing. Resources are not a cognitive property, so they do not belong in the numerator; a system is not more general for being cheap. They are load-bearing because a numerator reached for progressively less input is the signature of mechanism change, while a rising numerator reached by rising input is the signature of scaling within a fixed mechanism. A denominator keeps both readable and lets neither impersonate the other.

Economic usefulness is TRACKED SEPARATELY. It depends on labour prices, integration cost, regulation, and firm-level adoption capacity, none of which are properties of the system. It is retained because it is what most readers care about, and excluded from the composite because a change in it is not evidence about cognition. S22 establishes that economic-value benchmarks measure isolated tasks rather than end-to-end situated work (class: independently verified, qualitative).

Deployment scale is TRACKED SEPARATELY. Deployment is a distribution and marketing variable, moving on pricing, bundling, procurement cycles, and default placements. A frontier system with no distribution and a mediocre system inside a billion devices would rank in the wrong order on any capability reading that admitted scale. Self-reported usage data compounds this: S7 measures one laboratory's own user mix (class: self-reported) and the equivalent portal for another laboratory is blocked to automated fetching (S31; class: missing).

Safety and controllability is TRACKED SEPARATELY and NEVER SUMMED. A safer model is not a more general one. The properties are conceptually orthogonal and empirically capable of moving in opposite directions, so summing them creates two symmetrical failures: a safety improvement could offset and mask a capability plateau, and a capability jump could offset and mask a controllability regression. Either failure destroys the instrument's only useful property, which is returning a verdict its operator did not want. Safety is therefore published as a divergence indicator (I1) with no arithmetic path into the headline.

1.5 The remaining named terms

  • Transfer to unfamiliar tasks: folded into BREADTH, measured through dimension B.
  • Sample efficiency: folded into RESOURCE on the input side; no maintained public series for

    algorithmic efficiency exists (S36; class: missing), so it is carried as a declared gap.

  • Adaptability: folded into BREADTH and partly unmeasurable. In-context adaptation is observable

    through held-out performance; adaptation across sessions has no verified source.

  • Tool use: folded into HORIZON, since long-horizon unassisted completion is where tool use is

    exercised rather than declared (S14, S19; class: independently verified, manual-only feeds).

  • Learning from feedback: currently unmeasurable. No verified source isolates within-deployment

    learning from version-to-version change.

  • Robustness outside benchmark conditions: tracked separately in dimension E, framed and bounded

    but never stated as a number (G3).

  • Ability to replace or complement human cognitive work: tracked separately in dimension E and

    never summed, for the reasons given for economic usefulness.

1.6 The two reference models

Kurzweil-style long-run log progression charts. Useful because rate, not level, is the question a progress tracker exists to answer; because a log axis makes doubling behaviour visible as slope instead of leaving the reader to infer it; and because a long-run chart survives the saturation of any single underlying metric by carrying successive metrics on one axis. Dangerous because the log frame smuggles in an assumption of continuity the data cannot supply; because the metrics on such charts are selected after the fact for having been exponential, which is survivorship in the metric choice rather than a finding; and because the form invites extrapolation past the regime in which the generating mechanism held. Design consequence: retained as the RATE view only, over pinned metric versions, with no fitted line extended beyond the last observation.

The Doomsday Clock. Useful because it yields one legible position a lay reader can hold; because scheduled revision makes the instrument a public ritual with a cadence; and because a named committee owns the judgment, so the reading is attributable. Dangerous because it presents expert judgment in the visual grammar of measurement, importing precision the process does not have; because minutes-to-midnight implies a metric distance to an event when no such distance is defined; and because it cannot move for a reason a reader can audit, since no reader can recompute the position from published inputs. Design consequence: the clock is rejected as the primary headline, and its one transferable property, a fixed revision cadence with an attributable owner, is kept in the governance layer.

#

These are the operating rules that justify the roster in section 3, stated so as to be capable of rejecting a candidate. The confounders list at the foot of SOURCES.md and the established gaps G1 to G5 are the evidence that each addresses a realised problem rather than a hypothetical one.

  1. Demonstrated capability over inputs and claims. An indicator records what a system did under

    an evaluation someone else ran, not what was spent on it or announced about it. Inputs belong in the denominator and announcements belong nowhere. Laboratory self-reported figures are admitted only when tagged self-reported and only where no independent runner exists (S7, S20).

  2. The indicator must be able to move backwards. A metric monotonic by construction cannot

    register stagnation or regression, so it cannot fail, which makes it decoration. Every dimension carries at least one indicator whose value can fall. G2 establishes that no live source tracks generation-over-generation regression, so that capacity is built rather than sourced.

  3. Contamination resistance is a first-class requirement, not a caveat. A candidate states how

    its held-out split is protected before it enters the roster. Benchmarks with a private or semi-private set (S10, S11, S12, S37) are preferred over fully public ones; where only a public set exists, the novelty discount is applied explicitly rather than footnoted.

  4. Prefer sources whose scores a third party can reproduce. Reproducibility outranks coverage: a

    narrower benchmark with published harness, prompts and traces (S6, S16) beats a broader one whose pipeline is operator-controlled (S18). Where the auditable option is weaker, both are carried and the divergence between them becomes the diagnostic.

  5. Prefer duration-denominated and cost-denominated metrics over accuracy percentages. Accuracy

    percentages have a ceiling and stop moving exactly when the frontier is most interesting; duration and cost have none. Hence dimension C is duration-based (S5) and dimension G cost-based (S13), and an accuracy-only candidate is rejected unless it also carries a saturation half-life.

  6. Every observation pins a benchmark version. Confounder 6 records one tracked evaluation moving

    through three version lines and another maintaining five splits (class: independently verified), so an unpinned observation is not comparable to its predecessor. A candidate with no stable version identifier is rejected, and a version change opens a new series rather than continuing the old one.

  7. Measure the median and the worst case, never the best case. Best-of-k and best-reported-run

    figures measure the reporting process, not the system. Indicators are defined on the median run with the worst observed run published beside it, and a candidate that exists only as a leaderboard maximum is recomputed or rejected.

  8. An indicator with no live source is a declared gap, never an estimate. Where verification

    established non-existence (G1, G3, G4, G5), the field publishes as missing with the date checked. Interpolation, proxying by a related metric, and expert-judgment fill are prohibited inside the data layer, which is why the benchmark-to-reality gap is bounded but never given a number.

  9. Retroactively mutable series are re-fetched, never cached as truth. Confounder 1 records a

    latent-trait index re-anchored as benchmarks saturate, so historical values move after publication. Indicators built on such a series store fetch date, upstream revision, and prior value, and show revisions as revisions instead of overwriting silently.

  10. At least one indicator must only get worse if progress is real. Frontier eval headroom (A1)

    and benchmark saturation half-life (A2) both fall as capability rises. Their inclusion stops the instrument reading as a cheer, and a period in which they stop falling is itself a finding.

  11. At least one indicator must reveal the instrument itself failing. Rerun variance (D2) and the

    generation-over-generation regression count (D4) rise when either the systems or the measurement harness degrade, and open-weight lag (H1) moves when the observable population of models changes rather than the frontier. These are the instrument's smoke detectors.

  12. No indicator may depend on a single host, and no fetch is validated by status code.

    Confounder 5 establishes high source mortality: of the sources checked, four are retired, dormant or in maintenance mode (S16, S26, S27, S28) and two changed host or froze a version line (S18, S19); class for both counts: independently verified by live fetch. Confounder 7 records a government host returning HTTP 200 for a fabricated filename (class: independently verified). Every indicator names a primary and a fallback source or is flagged single-sourced and fragile, and every fetch validates content.

  13. Human preference is not capability. Preference and vote-based rankings (S18) measure what

    raters liked, are operator-controlled, and have no fixed zero across time. They are admitted as context only and are ineligible for any composite.

  14. Every reading carries a confidence band and a staleness flag. Bands are HIGH, MEDIUM or LOW

    and decay across horizons, per sBs/AGI-handbook/bible/SPINE.md; epistemic tags are ESTABLISHED, PROBABLE, CONTESTED or SPECULATIVE with a 90-day staleness flag, per sBs/fow/kb/INDEX.md. An indicator whose cadence cannot support a 90-day freshness check is admitted only at LOW confidence with the cadence stated. The jagged-frontier argument in sBs/AGI-handbook/chapters/03-capability-map.md is why no single indicator stands for a dimension.

#

Core dimensions A to D

Conventions. Confidence grades use the inherited house band convention (BRIEF, sBs/AGI-handbook/bible/SPINE.md): HIGH, MEDIUM or LOW attached to a stated horizon and decaying as the horizon lengthens; grades are 12-month unless stated. Verification classes are the six-level set fixed in CRITERIA C8. Thresholds, basket sizes and k values are specification parameters chosen here, not empirical claims. Every observation pins benchmark version, harness or scaffold identifier, and capture date, per SOURCES.md confounder 6.


A1. Frontier Eval Headroom

Operational definition For each benchmark b in a rotating basket B, headroom_b(t) = (ceil_b - best_b(t)) / (ceil_b - base_b), where best_b(t) is the highest verified score on the version-pinned benchmark at date t, ceil_b its declared maximum (replaced by a lower published human-expert ceiling where one exists, substitution recorded), base_b its published naive baseline. The indicator is the unweighted median across B, so no single collapsing benchmark drives it. Basket size pinned at eight. Rotation, both directions documented: ADMIT when a benchmark is carried in S1 with verified scores from three or more distinct developers, headroom above 0.60, and a pinnable version line; RETIRE when headroom stays below 0.10 for two consecutive quarters, its operator stops publishing, or its version line breaks comparability. Retirement draws the top entry from a published admission queue ranked on headroom then auditability class (IND before AGG before LAB). ARC-AGI-3 exists at S10 and is described there as unbeaten and agentic (S10, self-reported), making it a standing admission candidate; no score for it is stated anywhere in this spec. Rotation is chain-linked: both baskets are computed for one overlapping observation and the ratio of medians becomes a splice factor. The unspliced series is retained.

What it measures Unclaimed scoring room on the hardest evaluations independent parties currently run.

Why it signals progress toward general capability Falling headroom on a basket that keeps admitting harder instruments is the least ambiguous public evidence of depth gain.

What it does not measure Off-distribution generalisation, reliability, horizon, cost, transfer to paid work.

Unit Dimensionless ratio on [0, 1], two decimals.

Primary source S1.

Fallback or corroborating source S4 raw score CSV, plus S13 for independent re-runs on its covered subset. S4 raw scores only, never its fitted index, per SOURCES.md confounder 1.

Historical data availability Deep in S1 per benchmark and version, but honest backfill of the median starts only when eight qualifying benchmarks first coexisted, a property of the assembled basket.

Update cadence Weekly, against a feed updating continuous to daily (S1, independently verified by live fetch 2026-07-30).

Expected reporting lag Days to weeks per model, set by S1 ingestion.

Can it move backwards YES: admission of a harder benchmark, a re-graded or retracted submission lowering a standing best, or a tightened harness. The splice and unspliced series separate basket effects from score effects.

Known biases and confounders Best-of-any aggregation reads absent submissions as absent capability. Mixed auditability inside S1. Human-expert ceilings are estimates. Scaffolds unnormalised.

Gaming and contamination risk HIGH, by frontier developers: public-split training, selective submission, scaffold-heavy harnesses. Priced by dimension B, not assumed away.

Confidence grade HIGH at 12 months, lower at 36 given high basket mortality (an INFERRED extension of SOURCES.md confounder 5, which establishes source mortality rather than benchmark retirement, independently verified). Raised by a posted admission queue and a second independent runner.

Composite eligibility ELIGIBLE.

Class and lead-lag Core, leading.


A2. Benchmark Saturation Half-Life

Operational definition For each benchmark ever admitted to the A1 basket, let t0 be its first verified frontier submission in S1 and h0 its headroom then. Half-life is t_half - t0, where t_half is the first date headroom reaches or falls below h0 / 2. Benchmarks that have not halved are right-censored, excluded from the point estimate, and their count published beside it. The indicator is the median half-life over a trailing 36-month window, reported with contributing and censored counts, never as a bare figure.

What it measures How fast the measuring instruments stop discriminating. A property of the eval suite, not of the systems, and the reader-facing text says so.

Why it signals progress toward general capability It does not, directly. It is one of the named detectors separating a plateau from a melted ruler: A1 stalling while half-life is short indicts the ruler, A1 stalling while half-life lengthens indicts the systems.

What it does not measure Capability level, reliability, horizon, cost. Explicitly instrument-health.

Unit Months.

Primary source S1.

Fallback or corroborating source S4 raw score CSV as a second extraction path, not independent measurement, since both trace largely to one operator.

Historical data availability Back to S1's earliest version-pinned scores. Benchmarks retired before S1 took its current form are absent, biasing history toward survivors.

Update cadence Quarterly. Higher frequency adds noise to a duration.

Expected reporting lag One quarter plus S1 ingestion on the triggering observation.

Can it move backwards YES, as a lengthening median, when slow-saturating benchmarks enter the window or when submissions slow so halvings stop being recorded. Submissions per window are published so that activity artefact is visible.

Known biases and confounders Right-censoring, small per-window samples, survivorship. Benchmarks preserving headroom by design via withheld splits (S10, S11, S12, S37) saturate slower for design reasons.

Gaming and contamination risk MODERATE, by benchmark operators, who control version admission and can flatter a suite by scheduling harder variants. Low from developers.

Confidence grade MEDIUM; censoring and small samples dominate. Raised by a longer window and a preregistered basket history, so admissions cannot be chosen after seeing halvings.

Composite eligibility EXCLUDED. Summing an instrument measure into a capability figure mixes meta-measurement with measurement.

Class and lead-lag Core, leading, as a detector: it moves before A1 loses interpretability.


B1. Held-Out Novelty Gap

Operational definition For each pair p of one public and one withheld split of the same benchmark, and each model m evaluated on both under one harness inside a 90-day window, gap_{m,p} = score_public - score_withheld in percentage points on that benchmark's scale. The indicator is the median across qualifying model-pair observations, published with observation count and pair identities. Qualifying pairs: ARC-AGI-2 public tasks against the semi-private and the fully private sets, pinned by tier used since S10 operates two with different access rules; the public 2,500-question HLE set against its private set (S11, self-reported); SWE-bench Pro's public 731 tasks against its 858 never-published held-out tasks (S12, self-reported). FrontierMath (S37) joins only where its operator publishes a matched public comparator under the same grading; absent that, S37 feeds B2 only.

What it measures The drop when a system meets same-distribution items it could not have seen.

Why it signals progress toward general capability A shrinking gap at constant or rising public score is the strongest public evidence that measured competence is not memorised.

What it does not measure Generalisation outside the benchmark's own distribution. A withheld split protects against contamination, not against a narrow task family.

Unit Percentage points on the source scale, medianed across pairs.

Primary source S10.

Fallback or corroborating source S11 and S12; S12 is single-vendor and self-run, so it corroborates rather than confirms. Stated in reader-facing text: at S10, S11, S12 and S37 contamination protection is bought with lost external auditability, since nobody outside the operator can inspect the withheld items. At S37 the funding arrangement behind the held-out set is a known controversy the verifier could not confirm from the page (S37, missing), which lowers S37's weight without excluding it.

Historical data availability Short and irregular: the S11 public set was finalised 2025-04-03, its rolling fork launched 2025-10-08 (S11, self-reported). A single-digit number of comparable quarters per pair, none behind each withheld set's creation.

Update cadence Quarterly, event-driven: the series advances only when an operator posts matched results.

Expected reporting lag One to two quarters. Community submissions at S10 are explicitly unverified (S10, independently verified) and are excluded.

Can it move backwards YES: public-split contamination, an operator refreshing the public split but not the withheld one, or a changed withheld difficulty mix. Widening is a contamination detector.

Known biases and confounders No operator publishes an item-level difficulty match between splits, so part of any gap is composition, not contamination. Semi-private access at S10 is commercial, selecting who appears. Cross-run harness variance is undocumented.

Gaming and contamination risk MODERATE on the withheld side, HIGH on the public side, by frontier developers, via training-corpus inclusion. Semi-private sets degrade under repeated commercial access.

Confidence grade MEDIUM at 12 months, LOW at 36 given access dependencies. Raised by a published item-level difficulty match and a second operator adopting paired reporting.

Composite eligibility CONDITIONAL: eligible only with three or more qualifying model-pair observations spanning two or more operators; below that, diagnostic at zero weight.

Class and lead-lag Core, lagging: a withheld result is posted after the public result it is compared against.


B2. Contamination-Controlled Score Ratio

Operational definition On B1's qualifying pairs, ratio_{m,p} = score_withheld / score_public, computed only where score_public exceeds 10 points, since the ratio is unstable near zero. The indicator is the median across qualifying observations, normally on [0, 1] and permitted above 1 where a withheld set is easier. Published jointly with B1, never instead of it: the ratio is scale-free and comparable across differing score ranges but degenerates toward 1 as public scores approach the ceiling, which is when B1 stays informative; B1 degenerates when scales differ, which is when the ratio stays informative. Both carry the same observation counts.

What it measures The proportion of measured public-split performance surviving contact with unseen items.

Why it signals progress toward general capability A ratio near 1 across several operators and task families is the cleanest available signal that scores reflect competence rather than exposure.

What it does not measure Absolute capability: a ratio near 1 fits high and low performance equally, so it is never read without A1 levels.

Unit Dimensionless ratio, two decimals.

Primary source S12.

Fallback or corroborating source S10, S11 and S37. S37 contributes only where matched grading over an unpublished tier and a comparable published comparator both exist, and its unconfirmed funding arrangement (S37, missing) is disclosed wherever it contributes.

Historical data availability As B1 and slightly worse, since both scores must sit on one scale inside one window. Non-comparable before each withheld set existed.

Update cadence Quarterly, event-driven.

Expected reporting lag One to two quarters.

Can it move backwards YES: rising public-split contamination, a hardened withheld set, or a developer optimising against a semi-private set it repeatedly buys access to.

Known biases and confounders Unmatched split difficulty maps straight into the ratio. Small denominators inflate volatility, only partly controlled by the 10-point floor. One operator's grading choices move the primary series.

Gaming and contamination risk MODERATE overall, HIGH on the public denominator, by frontier developers. The ratio rises either by generalising better or by scoring worse publicly, so both components are published.

Confidence grade MEDIUM. Raised by matched-difficulty publication and a third operator adopting paired reporting.

Composite eligibility CONDITIONAL: B1's three-observation, two-operator threshold, and only when reported jointly with B1, so the pair cannot be quoted as one reassuring figure.

Class and lead-lag Core, lagging.


C1. 50% Task Time Horizon

Operational definition The task length, in human expert completion time, at which estimated success probability equals 0.50 under the logistic fit published by S5 over its task suite. Fitted by S5, not by the tracker: the tracker ingests benchmark_results_1_1.yaml and task_results_1_1.yaml, pins the file version, and records model identifier, scaffold and fit variant. The frontier series takes the monthly maximum across models, naming the contributing model. Where S5 publishes several scaffolds for one model, the tracker records highest and lowest and publishes the spread rather than silently taking the best.

What it measures The duration of task a system completes unaided at even odds.

Why it signals progress toward general capability Horizon is a constitutive axis; longer unaided task chains are something a shorter-horizon system cannot do, independent of raw accuracy.

What it does not measure Reliability anyone would deploy at: a 0.50 threshold describes a coin flip. Nor breadth beyond S5's suite, nor cost.

Unit Minutes of human expert completion time, log axis.

Primary source S5.

Fallback or corroborating source None verified. S14 and S6 measure agentic task success but publish no time-horizon fit. The cost is a single point of failure: if S5 stops updating, the horizon dimension collapses to whatever C3 can infer from stale data.

Historical data availability A retrospective series across model generations, but irregular, over a sparse model subset with heavy scaffolding variance; updates observed only March to May 2026 and the operator self-describes limited capacity (S5, independently verified by live fetch 2026-07-30). Comparable only within a pinned fit version.

Update cadence Irregular, treated as unscheduled, with a staleness flag at 90 days without a new file.

Expected reporting lag Months, bounded by nothing the tracker controls.

Can it move backwards YES: a refit on a revised suite, a scaffold change, or an added long task that all models fail lowers the curve for everyone.

Known biases and confounders Human baseline times are estimates. Scaffolding variance dominates several comparisons. The suite is software and research work, so the horizon generalises to those families. The do-not-train canary (S5, independently verified) is a control and also an admission contamination is plausible.

Gaming and contamination risk MODERATE, by developers, via scaffold engineering aimed at S5's suite and training exposure the canary discourages but cannot prevent.

Confidence grade MEDIUM at 12 months, LOW at 36 given single-source dependency and observed cadence. Raised by a second operator publishing a horizon fit, or by S5 restoring regular updates at a stable fit version.

Composite eligibility CONDITIONAL: eligible only while S5 sits inside the 90-day staleness window and only when reported beside C2, so the 0.50 threshold never stands alone as the horizon figure.

Class and lead-lag Core, leading.


C2. 80% Task Time Horizon

Operational definition Identical to C1 with the threshold at 0.80. S5 exposes both as chart variants on one underlying fit and the 0.80 variant is the less emphasised of the two (S5, independently verified by live fetch 2026-07-30). Same ingestion, version pinning, scaffold-spread reporting and frontier-maximum construction.

What it measures The duration of task a system completes unaided at a reliability level closer to one a user would accept.

Why it signals progress toward general capability The operational target is reliability-weighted, so the higher threshold is the more faithful reading of the horizon axis. C2 is the more honest of the two figures and C1 is the more quoted; the spec says so in reader-facing text rather than resolving it silently. Where they disagree in direction, C2 governs the dimension and the divergence is published as a finding.

What it does not measure Reliability above 0.80, breadth outside S5's suite, cost per completed task.

Unit Minutes of human expert completion time, log axis.

Primary source S5.

Fallback or corroborating source None verified. Same single-point-of-failure cost as C1 and slightly worse: as the less emphasised variant it is the likelier of the two to go unrefreshed in a reduced-capacity release.

Historical data availability As C1, plus greater long-tail noise, since fewer successful long-task observations constrain the 0.80 fit. Comparable only within a pinned fit version.

Update cadence Irregular, tied to S5 releases, same 90-day staleness flag.

Expected reporting lag Months.

Can it move backwards YES, by C1's mechanisms and more sensitively: further out on the logistic curve, it moves more per added long-task failure.

Known biases and confounders All of C1's, plus fit sensitivity at the threshold. The sparse model subset means the frontier maximum is often set by one scaffold-model pairing.

Gaming and contamination risk MODERATE, by developers, by C1's mechanisms.

Confidence grade MEDIUM at 12 months, LOW at 36. Raised by S5 publishing per-threshold uncertainty intervals, and by any second operator replicating the fit.

Composite eligibility ELIGIBLE while S5 sits inside the staleness window. C2 is the horizon dimension's primary input; C1 is carried as context.

Class and lead-lag Core, leading.


C3. Horizon Doubling Time

Operational definition Ordinary least squares regression of log2(C2) on time over a trailing 24-month window of frontier-maximum observations, six observations minimum. Doubling time is 1 / slope in months, published with the regression interval and observation count, never as a point figure. The same regression is run on C1 and reported beside it, so a reader sees whether the two thresholds imply different rates. Windows are recomputed from scratch each release, because S5 revisions move historical points.

What it measures The rate at which the measured autonomy horizon is changing.

Why it signals progress toward general capability Position and rate are different questions and the recommended headline is a two-number profile; C3 is the rate input for horizon, and a lengthening doubling time is the clearest available plateau detector on autonomy.

What it does not measure The horizon level, breadth, or anything outside the window. It describes observed history; the data layer states no projection.

Unit Months per doubling, with an interval.

Primary source S5.

Fallback or corroborating source None verified, being a transform of C2. The observation set is published each release so a third party can refit, substituting reproducibility for source redundancy.

Historical data availability Bounded by C2's history and the six-observation minimum. At the observed S5 cadence the window sits at or near that minimum, disclosed with every value.

Update cadence Recomputed on every S5 release, reported quarterly.

Expected reporting lag Months, inherited from S5.

Can it move backwards YES: recent frontier observations falling below trend, or mechanically when a scaffold-driven outlier leaves the window. Each release names the observations that entered and left.

Known biases and confounders Small-sample regression on irregularly spaced, revisable points. The log-linear form is imposed and untested here, so a change in functional form appears as drifting residuals rather than a changed doubling time; residuals are published.

Gaming and contamination risk LOW directly, since no external party reports this figure. Inherits C2's MODERATE risk through its input.

Confidence grade LOW at 12 months; six observations on a revisable series supports no more. MEDIUM on a regular S5 cadence giving a dozen or more stable observations plus published residual diagnostics.

Composite eligibility EXCLUDED from any level composite, being a rate. It is the horizon input to the rate half of the two-number headline profile.

Class and lead-lag Core, lagging: a rate is estimable only after its levels exist.


D1. pass@1 to pass@k Spread

Operational definition For each model and version-pinned agentic benchmark with at least k independent attempts per task, spread = pass@k - pass@1, k pinned at 8: the smallest power of two at which the estimator is stable enough to publish and still cheap. pass@1 is the mean per-task success rate over attempts, not the best attempt; pass@k is the unbiased estimator over the same attempt set, computed per task then averaged. The indicator is the median across qualifying model-benchmark cells, cell count published. Single-attempt submissions are excluded and counted in a published exclusion tally, because treating one submission as pass@1 and a best-of-many leaderboard entry as pass@k would manufacture a spread from two different runs.

What it measures How much of a headline capability depends on being allowed to retry.

Why it signals progress toward general capability A narrowing spread at constant pass@k means the system reaches its own ceiling on the first attempt, which is what unassisted completion requires.

What it does not measure Calibration, stability under identical conditions (D2), whether retries are affordable.

Unit Percentage points.

Primary source S6, whose artefacts include all_preds.jsonl, per-attempt logs and required reasoning trajectories (S6, independently verified by live fetch 2026-07-30), which is what makes per-attempt recomputation possible.

Fallback or corroborating source S14, whose submissions must pin a Harbor dataset version (S14, independently verified), giving a second agentic suite with occasional multi-attempt runs. S14 results are not in-tree, so extraction is scrape-only.

Historical data availability From July 2024 for the trajectory-carrying portion of S6, since traces became required then (S6, independently verified). Earlier submissions lack the artefacts.

Update cadence Monthly, against a continuous submission-driven tree.

Expected reporting lag Weeks to months after a release, depending on somebody submitting a multi-attempt run.

Can it move backwards YES: a best-of-many submission configuration, a new split raising task variance, or a scaffold introducing nondeterminism. Widening at constant pass@k is a reliability regression and a named regression detector.

Known biases and confounders Submission selection: developers submit flattering runs. Attempt independence is assumed and rarely verified, since a scaffold with cross-attempt memory violates it. Small k is noisy. Scaffolds unnormalised across submissions.

Gaming and contamination risk MODERATE, by developers, through submitting high-k runs while never publishing single-attempt rates. The exclusion tally makes that visible as missing data.

Confidence grade MEDIUM. Raised by operators mandating a fixed attempt budget and publishing per-attempt results, which removes the selection and independence problems together.

Composite eligibility CONDITIONAL: eligible with five or more qualifying cells across two or more developers in the period; otherwise diagnostic only.

Class and lead-lag Core, lagging.


D2. Rerun Variance

Operational definition DECLARED GAP with a self-computed construction. No public source publishes run-to-run variance (G1, established by search 2026-07-30). Construction: a fixed probe of 100 tasks drawn from version-pinned benchmarks in the A1 basket, each model run 10 times per task at fixed temperature, scaffold, prompt and harness version; take the standard deviation of the aggregate score across the 10 runs; the indicator is the median of that standard deviation across models in the period. Every probe parameter is published, because the number is meaningless without them. Construction cost, stated so the gap is not hidden: paid inference at 1,000 task-runs per model per period plus a maintained harness, the largest recurring cost in the indicator set and the one likeliest to exceed a one-person budget. Unfunded, D2 publishes as DECLARED GAP G1 with no value.

What it measures How much a system's measured score moves when nothing about the task changes.

Why it signals progress toward general capability Reliability is a constitutive axis; a system whose score swings between identical runs has not established the competence its mean reports.

What it does not measure Variance under prompt or scaffold changes, calibration, capability level.

Unit Score standard deviation in percentage points, medianed across models.

Primary source None verified: gap G1. Nearest partial coverage is S13, which publishes confidence intervals on its own re-runs (S13, independently verified by live fetch 2026-07-30), and S16, whose metric matrix includes reproducibility-relevant metrics but which entered maintenance mode on 1 June 2026 (S16, independently verified). Neither publishes rerun variance as a series.

Fallback or corroborating source None verified. The cost: the reliability axis cannot be reported from public data at all, so the tracker either funds the probe or shows the axis as an explicit hole. The spec chooses the hole, because a proxy here is the exact failure the brief guards against.

Historical data availability None. Retired model endpoints cannot be re-probed, so history starts the day the probe starts.

Update cadence Quarterly if funded; while declared a gap, the tracker shows the gap with its establishment date rather than an empty cell.

Expected reporting lag Days after a probe run, the tracker owning the pipeline. The only indicator with no external lag, the compensating advantage of self-computation.

Can it move backwards YES: variance rises when a provider changes serving infrastructure, quantisation or routing behind a stable model name, without announcement, so D2 doubles as a detector for silent endpoint drift.

Known biases and confounders Provider-side and model-side nondeterminism are not separable. Temperature choice sets the answer. Transient errors contaminate the run set unless retried, and retry policy itself moves the number. The probe set ages as its benchmarks saturate.

Gaming and contamination risk LOW from developers, who neither report this nor select the runs. MODERATE from providers, weakly, through undisclosed serving changes. Self-computation removes the developer-selection risk afflicting D1.

Confidence grade LOW at 12 months as an unfunded declared gap. MEDIUM on the first funded probe with published parameters; HIGH only on independent replication.

Composite eligibility EXCLUDED while declared a gap; CONDITIONAL on funding, eligible after two consecutive funded quarters with published probe parameters.

Class and lead-lag Core, lagging.


D3. Calibration Error

Operational definition DECLARED GAP with a self-computed construction. No public leaderboard publishes model calibration (G1, established by search 2026-07-30). Construction: expected calibration error over a probe of multiple-choice and short-answer items with elicited numeric confidence, binned into 10 equal-width confidence bins, computed as the sample-weighted mean absolute difference between mean stated confidence and observed accuracy per bin. The probe reuses the D2 task set where item format permits, sharing one harness and one cost line. Published parameters: bin count, elicitation prompt, whether verbalised or token-probability confidence is used, and item set version. Those two confidence sources are different quantities and the tracker labels which it uses rather than mixing them. Cost is marginal on a funded D2; standalone cost is one probe run per model per period plus the elicitation harness. Unfunded, D3 publishes as DECLARED GAP G1, with no value and no proxy.

What it measures Whether stated confidence matches observed accuracy.

Why it signals progress toward general capability Unassisted operation over long horizons requires a system to know when it does not know, since that is what lets it stop, ask or escalate instead of proceeding wrongly.

What it does not measure Accuracy itself, honesty under adversarial pressure, calibration on open-ended generation, where no scoring rule is agreed.

Unit Expected calibration error as a proportion on [0, 1], three decimals.

Primary source None verified: gap G1. S16 includes calibration in its metric matrix but has been in maintenance mode since 1 June 2026 (S16, independently verified by live fetch 2026-07-30) and its web interface yielded no fetchable content, so it cannot serve as a live feed. S13's confidence intervals are intervals over its own measurements, not model calibration, and using them as such is a category error.

Fallback or corroborating source None verified. The cost: the reliability dimension rests on two self-computed indicators, both unfunded at specification time, stated in MVP scoping rather than concealed.

Historical data availability None, for D2's reason. S16's archive could supply a one-off historical anchor for the models it covered, not a comparable series, and any use of it is labelled as such.

Update cadence Quarterly if funded; otherwise shown as a gap with its establishment date.

Expected reporting lag Days after a probe run.

Can it move backwards YES: calibration degrades when post-training optimises for confident presentation, and a model can gain accuracy while losing calibration. Catching that divergence is why this indicator exists, and it is a named regression detector.

Known biases and confounders Elicitation format dominates the result, so cross-model comparison needs an identical prompt, which favours models tuned to it. Multiple-choice items overstate calibration relative to open-ended work. Bin count moves the number. Refusals are excluded under a documented rule with the exclusion rate published.

Gaming and contamination risk LOW from developers, who neither report nor select these runs. MODERATE from prompt-format sensitivity, which is not gaming but does the same comparability damage.

Confidence grade LOW at 12 months as an unfunded gap. MEDIUM on a funded probe with published elicitation parameters; HIGH only if an independent evaluation body restores a public calibration series.

Composite eligibility EXCLUDED while declared a gap; CONDITIONAL on funding, on D2's two-consecutive-quarters rule.

Class and lead-lag Core, leading: calibration decay is observable before it surfaces as failures in horizon or task-completion indicators.


D4. Generation-over-Generation Regression Count

Operational definition DECLARED GAP G2 with a self-computed construction that is cheap and is the recommended approach (G2, established by search 2026-07-30). Retain every S1 dump as an immutable dated snapshot. For each consecutive snapshot pair, join on (developer, benchmark identifier, benchmark version, harness identifier), order each developer's models by release date from S2, and count every case where a newer model from the same developer scores lower than its immediate predecessor on the same version-pinned benchmark by more than a materiality threshold of 2 percentage points, set to exclude grading noise. The indicator is that count per period, published over a denominator of comparable model-benchmark pairs so it is never read as a bare number. Retroactive-drift handling, the reason for this precision: the diff runs exclusively on RAW benchmark scores from S1, never on fitted index values from S4, because S4's ECI is a latent-trait fit re-anchored as benchmarks saturate, so its historical values move retroactively and any cached figure silently drifts (SOURCES.md confounder 1, independently verified). Diffing S4 index values would generate regressions that are artefacts of refitting. Because S1 may itself revise a historical raw score, each snapshot pair is first diffed for revisions; revised cells leave the regression count and are reported separately as a revision count, so a corrected score is never reported as a capability regression.

What it measures How often a developer's newer model is worse than its own predecessor on a fixed instrument.

Why it signals progress toward general capability Public commentary assumes monotonic progress and the inherited jagged-frontier material says otherwise. A non-zero count is direct evidence that gain is not uniform, and a rising count against a stable denominator is a named plateau detector.

What it does not measure The magnitude or importance of any regression, whether it was a deliberate trade against cost or safety, capability level.

Unit Count of regressing model-benchmark pairs per period, over a published denominator of comparable pairs.

Primary source S1, diffed across dated snapshots, with S2 supplying release ordering.

Fallback or corroborating source S4's raw score CSV, a second extraction path used to detect extraction errors only. S13 corroborates on its covered subset with its own independent runs, the only genuinely independent check available here.

Historical data availability Only from the first snapshot the tracker itself retains. S1's dump gives current values, not a version history, so backfill is impossible after the fact; snapshot retention from day one is a hard pipeline requirement and the cheapest irreversible decision in the build.

Update cadence Monthly diffs, matched to the snapshot schedule.

Expected reporting lag One month plus S1's ingestion lag on the newer model.

Can it move backwards YES, as a falling count, the ordinary case when a generation is uniformly better. It rises on narrower models shipped under a successor name, a harness change penalising a newer model, or a capability-for-cost trade. The harness mechanism is a false positive, which is why the join key includes the harness identifier.

Known biases and confounders Lineage is not always well defined, since developers rename and rebrand, and same-developer ordering relies on S2 release dates that are estimated for some entries (S2, estimated). Denominator instability: more submissions means more chances to regress. Version-pinning failure at source would silently compare unlike things.

Gaming and contamination risk LOW, by developers, who do not report this and whose already published data it uses. The realistic risk is selective non-submission: a developer expecting a regression does not submit, suppressing the count. Denominator and per-developer submission counts are published so suppression is visible.

Confidence grade MEDIUM at 12 months, rising as the snapshot archive lengthens. Raised by S1 publishing its own revision history, which removes dependence on private snapshot retention, and by a second aggregator adopting version-pinned score archiving.

Composite eligibility EXCLUDED from any capability level composite, being a count of exceptions rather than a measure of level. It is a mandatory published detector beside the headline and feeds the weakest-link floor in the headline methodology.

Class and lead-lag Core, lagging.

Contextual and supporting dimensions E to I

Ten indicators. Field 3 states whether the indicator signals general capability; for G, H and I it does not, and the field says what it contributes instead.

Identifier footnote, stated once: the resource indicators are written G1i, G2i, G3i, G4i because SOURCES.md uses G1 to G5 for established non-existence gaps, and an unqualified "G1" would be ambiguous between the compute-threshold indicator and the calibration gap.


E1. Real-Task Win Rate vs Human Expert

  1. Operational definition wins / (wins + losses + ties) over a pinned suite of

occupationally grounded deliverables, model output against human-expert output, one grader protocol. Stored with suite ID, suite version, grader-protocol version and index methodology version. Both a ties-as-half-wins and a ties-excluded value are kept; the rule is never changed retroactively.

  1. What it measures Head-to-head parity on discrete professional deliverables against a named

human comparison group.

  1. Why it signals progress toward general capability Weakly. It is a breadth proxy over

occupational task space, not a capability measure: the tasks are curated, isolated and stripped of the organisational context that makes real work hard (S22, independently verified as a published critique). It contributes evidence that scored capability reaches deliverables a domain expert recognises, which no academic benchmark supplies.

  1. What it does not measure End-to-end job performance, multi-day or multi-stakeholder work,

economic value realised, willingness to delegate unsupervised, anything outside the suite.

  1. Unit Proportion, 0 to 1, dimensionless.
  2. Primary source S13 (GDPval-AA re-implementation).
  3. Fallback or corroborating source S20, for framing only: it is lab self-reported on the

lab's own economic-value construct, and S13's numbers are not S20's numbers. No second independent re-implementation is verified. Cost: single-vendor dependency, and no way to separate re-implementation error from genuine disagreement on task difficulty.

  1. Historical data availability Shallow. Both pages were live on 2026-07-30 (independently

verified as reachable); neither establishes a multi-year back-series, and S13's 7/30/90-day series sit behind its Commercial tier (independently verified from its API docs). Series begins at first ingestion.

  1. Update cadence Release-driven, not calendar-driven. Poll monthly.
  2. Expected reporting lag Days to weeks after a release for S13; unbounded for S20.
  3. Can it move backwards YES: new grader protocol; harder tasks at a suite version bump; a

model regressing on long-form deliverables while gaining elsewhere (jagged frontier); recomputation after a tie-rule change.

  1. Known biases and confounders Grader identity and instructions dominate. The human

comparison group is a convenience sample, not best-in-field. Task selection encodes the author's model of knowledge work. S13's index methodology is versioned, so only within-version segments compare: every observation pins the methodology version and the render shows a visible break at each boundary rather than interpolating across it.

  1. Gaming and contamination risk HIGH. By labs: suite tasks are partly public (S22 records

public fractions of 10 of 240, 480 of 480 and 220 of 1320 across three economic-value suites, independently verified as stated there), so training on the published portion is available. By graders drawn from a lab's own model family. By suite authors, through task retirement.

  1. Confidence grade LOW. Raised by a second independent re-implementation, a published

inter-grader agreement statistic, a stable held-out partition, or an S13 methodology changelog with recomputed history.

  1. Composite eligibility CONDITIONAL: admissible only inside one S13 methodology-version

segment, under a written licence permitting derived publication (S13 is proprietary, attribution required, redistribution by contract, independently verified from its docs), and at a capped weight so E1 alone cannot move the headline. Otherwise dashboard-only.

  1. Class and lead-lag Supporting. Lagging.

E2. Benchmark-to-Reality Gap

  1. Operational definition Not a scalar. A bounded directional indicator with a declared

measurement gap. Stored value: direction in {widening, stable, narrowing, undetermined}, assigned only when two independent structural signals agree; bound_evidence, source-tagged structural facts that constrain the gap without sizing it; quantified = false, permanently, until the roadmap construction below is executed. A numeric E2 fails schema validation.

  1. What it measures Direction of travel, and the size of the evidentiary hole, between

curated-evaluation scores and the same competence exercised in situ.

  1. Why it signals progress toward general capability It does not, in either direction. It is an

interpretation-layer control on every other indicator, stating how much of a benchmark gain may be read as a real-world gain. Its contribution is to stop benchmark motion becoming a capability claim silently.

  1. What it does not measure Gap magnitude, its shape across domains, or whether the gap is a

property of models or of benchmark construction.

  1. Unit Ordinal direction plus a declared-unmeasured flag. No numeric unit exists by

construction.

  1. Primary source S22.
  2. Fallback or corroborating source S5's limitations note, in which the eval team states it does

not know the gap's size (search-verified, per gap G3), and S12's held-out and commercial-private partitions for structural bounds. No quantified source exists: gap G3 is established, which is why field 1 refuses a number. Cost: E2 cannot be audited numerically, cannot enter a composite, and its direction rests on judgment.

  1. Historical data availability None as a series. S22 is dated 2026-02-13 (independently

verified as its publication date). The flag begins at first ingestion.

  1. Update cadence Quarterly manual review, plus an event trigger when a suite publishes a

held-out-versus-public score split.

  1. Expected reporting lag Months. Qualitative publications are the input.
  2. Can it move backwards YES: the flag can turn from narrowing to widening, or revert to

undetermined when the two signals stop agreeing. Undetermined is the expected default, not a failure.

  1. Known biases and confounders The flag is expert judgment and inherits the assigner's

priors, so it records the assigner, the date and the two signals relied on. Publication bias: benchmark critiques get written when benchmarks look strong.

  1. Gaming and contamination risk MODERATE. Not gameable as a number because there is no

number. Gameable rhetorically by any party, including this tracker's author, through selection of what enters bound_evidence. Mitigated by requiring a source ID per entry.

  1. Confidence grade LOW for direction, HIGH for the claim that no quantification exists.

Raised only by the roadmap item.

  1. Composite eligibility EXCLUDED. An unquantified indicator cannot be summed; it acts on the

composite by constraining interpretation.

  1. Class and lead-lag Supporting. Lagging.

Roadmap item. For N tasks, obtain a curated benchmark instance and an in-situ instance of the same competence, scored under one grader protocol, the in-situ instance embedded in real context (real repository, real stakeholder, real ambiguity, real tool access). E2 then becomes the difference of the two scores with a confidence interval. Cost: original instrument construction with human-expert grading, out of scope for a one-person tracker at MVP. Until then E2 stays a declared gap.


Resolution of a challenge to this slot, recorded rather than buried. A reviewer argued for deleting this indicator outright: at most two non-comparable points, since S21 is a second-edition challenge whose suite changes between editions (independently verified), an annual cadence, a scrape-only feed, and a measured quantity that is not the named one, which is a slot with a name on it rather than an instrument, and the name is what gets cited. The argument on the name is correct and is accepted. Deletion is rejected on two grounds. The first is mechanical and decisive: the frozen evaluation surface enumerates this indicator ID by name and a missing F1 block fails the gate, and that surface is not a builder's to edit. The second is substantive: the floor argument in field 3 does real work, because sustained near-zero embodied competence is evidence against a generality claim whatever the text scores say, and a deleted indicator cannot carry a floor. The remedy taken is therefore the narrower one, aimed at the actual defect: the indicator is renamed to what it measures, the simulation status is stated in fields 2 and 15, and the derivation of any embodied-competence claim from it is prohibited in the block rather than merely discouraged beside it.

F1. Simulated Embodied-Task Success Rate

  1. Operational definition successful_episodes / attempted_episodes on a pinned suite of

simulated household or workplace manipulation tasks at a fixed success criterion, per suite version and per embodiment, carrying environment_class drawn from the single vocabulary declared in the section 10.2 schema, which is {simulated, structured-physical, unstructured-physical, unspecified-by-source}. Only simulated is populated by any verified source, and it is the class the indicator name now states. The schema is the authority for this field: an earlier draft of this block and of the schema carried two different vocabularies, which would have made every F1 record fail validation, and the two were reconciled to one set after a cold verifier found the collision. The two physical classes are separate series that no verified source currently fills; they are never merged with the simulated series and never averaged into it.

  1. What it measures Success on a simulated benchmark of embodied household tasks, and nothing

outside simulation. This is a simulation indicator. It is named for what it measures rather than for the dimension it sits under, because the earlier name asserted unstructured-environment manipulation and no verified source supplies that quantity (gap G4, established).

  1. Why it signals progress toward general capability Weakly, prospectively, and in one direction

only. Embodiment is where the cognitive-task framing of the operational target is most incomplete, so the slot stops generality being declared on text alone. Its contribution is a floor, and the floor runs downward only: sustained near-zero success even in simulation is evidence against a generality claim whatever the text scores say, because simulation is the easier case. A high simulated value is not the converse and carries no evidence about physical competence, since simulation success is necessary and not sufficient for it.

  1. What it does not measure Physical robustness, novel geometry, clutter, lighting, human

presence, sim-to-real transfer, hardware reliability, the long tail of physical failure.

  1. Unit Proportion, 0 to 1, dimensionless.
  2. Primary source S21.
  3. Fallback or corroborating source None verified. S30's feed is not found: an SPA shell, with

/leaderboard and /api/leaderboard both 404 (independently verified 2026-07-30). Its rating is a crowdsourced pairwise Elo with no fixed zero across time, so recovering the feed would still not supply a comparable level. Cost: F1 rests on one academic challenge, and a single maintainer decision retires the indicator.

  1. Historical data availability At most one comparable point per year. S21 is a second-edition

challenge (independently verified), so at most two editions exist and the suite changes between them, breaking cross-edition comparability.

  1. Update cadence Annual, event-driven. The 2026 edition opened 2026-07-02, submissions close

2026-10-16, winners 2026-11-04 (all independently verified from the challenge page).

  1. Expected reporting lag Weeks after announcement plus manual extraction: the board is a

hosted interactive Space with no results CSV or JSON (independently verified), so ingestion is a scrape.

  1. Can it move backwards YES: a harder suite in a new edition; fewer or weaker entrants, since

the value reflects who competed rather than the state of the art; a changed success criterion.

  1. Known biases and confounders Simulation is not reality, a category confound rather than a

correctable bias. Entrant self-selection: a thin field reads as regression. Overfitting to one simulator's physics. Academic-calendar effects on participation.

  1. Gaming and contamination risk MODERATE. By entrants, through simulator-specific exploits

that do not transfer. Retrieval contamination is lower than for text benchmarks because episodes are generated, but the suite is published in an open repository (independently verified), which permits targeted tuning.

  1. Confidence grade LOW. Raised by any benchmark meeting both of gap G4's conditions, a

physical-robot evaluation with a stable suite across at least three periods, or recovery of an S30-class feed on an anchored scale.

  1. Composite eligibility EXCLUDED at MVP and, on the naming ground, further constrained. The

exclusion grounds are unchanged: annual cadence, a suite that changes between editions, and a simulation-versus-reality category mismatch. Added constraint: this is a simulation indicator, and no embodied-competence claim may be derived from it, by this tracker or by any reader of it. A statement of the form "embodied capability reached X" is not a permitted reading of any F1 value at any level; the only permitted readings are "simulated embodied-task success reached X on suite version V" and the field 3 floor reading. The render carries the simulation label inside the indicator title, not only in a caption, and environment_class is displayed on every point.

  1. Class and lead-lag Supporting. Lagging.

Position stated honestly. Embodiment, as the dimension is named, is unmeasured. Gap G4 is established: no robotics benchmark carries both unstructured-environment tasks and a machine-readable time series. What F1 now carries is a simulated proxy under its own name, plus the declaration that the physical quantity has no source. Making a real embodiment indicator requires a third party running a fixed physical suite on comparable hardware across time, publishing per-episode outcomes machine-readably, and versioning the suite. That is beyond a one-person tracker's build capacity. The interpretation layer therefore states, in one sentence and without qualification, that embodiment is unmeasured, and F1 is displayed as a simulation series that supports the floor argument and no positive embodiment claim.


G1i. Compute to Reach a Fixed Capability Threshold

  1. Operational definition Self-computed. Fix threshold T as a score on a pinned benchmark

version; per period, find the model with the lowest estimated training compute scoring at least T. Store (threshold_benchmark, benchmark_version, threshold_value, min_compute_FLOP, model_id, period); the series is the trajectory of min_compute_FLOP. Thresholds are never respecified: a saturated threshold is retired and a higher one opened alongside, both displayed.

  1. What it measures The falling compute cost of a frozen capability level, bundling algorithmic,

data and training-recipe improvement into one observable.

  1. Why it signals progress toward general capability It does not: capability is held fixed by

construction, so the indicator moves on efficiency. Its contribution is the denominator of the operational target. It is the only indicator separating capability bought with resources from capability obtained more cheaply, and the only one that would show an efficiency plateau under rising headline scores.

  1. What it does not measure Frontier capability, inference cost, realised dollar cost,

algorithmic progress isolated from hardware or data, anything above the threshold.

  1. Unit FLOP, log scale; derived view in FLOP-halving-time in months.
  2. Primary source S1 for scores at a pinned benchmark version, joined to S2 for compute.
  3. Fallback or corroborating source S36, as literature context only, never as a series. This is

load-bearing: S36 establishes that no maintained algorithmic-efficiency series exists. The topic page carries two publications, dated 2024-03-12 and 2022-12-12 (independently verified as the two dates listed), and no feed, so any efficiency trend drawn from that literature rests on papers between roughly one and four years old as of 2026-07-30 (inferred from those dates), covering generations no longer at the frontier. G1i is therefore specified as a self-computed construction from S1 plus S2. Cost: the tracker owns the methodology and its errors, with no external series to cross-check.

  1. Historical data availability Deepest of the contextual family. S2 covers notable models

historically with compute estimates in a downloadable CSV (independently verified as fetchable); S1 supplies scores as far back as each benchmark's own history.

  1. Update cadence Monthly recompute. Inputs move faster: S1 continuous to daily, S2 on a weekly

automated sweep (both independently verified as stated cadences).

  1. Expected reporting lag Roughly two weeks for major models to appear in S2 (independently

verified as the source's stated target), plus the recompute interval.

  1. Can it move backwards YES: an S2 estimate revised upward raises the historical minimum

retroactively; a small model's score revised down on rerun removes the threshold-holder; a benchmark version change invalidates the pinned threshold.

  1. Known biases and confounders S2's figures are estimates from parameter and token counts

rather than lab disclosures, with wide error bars revisable without notice (independently verified as a recorded confounder). Coverage is limited to notable models, so a cheap capable model outside the set is invisible. Distillation makes a student's own training compute a misleading measure of the compute needed to reach T. Threshold choice materially changes the trend.

  1. Gaming and contamination risk MODERATE. By labs, through benchmark-targeted training of

small models, which lowers apparent compute-to-threshold with no real efficiency gain. Contamination inflates small-model scores specifically, since small models gain most from memorisation. Mitigation: a threshold benchmark with a held-out partition, cross-checked against the B-family contamination indicators.

  1. Confidence grade MEDIUM. Raised by lab disclosure of training compute, a maintained

third-party efficiency series replacing the dormant S36 line, a published threshold-sensitivity fan across at least three thresholds, or distillation provenance in model metadata.

  1. Composite eligibility CONDITIONAL: admissible only as an explicit denominator in a

capability-per-resource ratio, never additive, and only with the threshold-sensitivity fan published alongside. Any construction letting falling compute cost add to a capability score is prohibited.

  1. Class and lead-lag Contextual. Leading: efficiency gains surface in small-model results

before frontier capability.


G2i. Inference Cost per Solved Task at Fixed Reliability

  1. Operational definition total_inference_cost / tasks_solved, where solved counts only at a

declared reliability level (for example, correct on k of k independent attempts), and cost sums input and output token charges at published list price for the pinned model version, including reasoning tokens where billed. Store (benchmark, benchmark_version, reliability_rule, model_id, price_snapshot_date, cost_per_solved_task). A separate hardware_price_performance series is never blended into it.

  1. What it measures The marginal price of a unit of reliable completed work, for one named

benchmark.

  1. Why it signals progress toward general capability It does not: capability is fixed by the

benchmark and the reliability rule, so the indicator moves on price and token efficiency. Its contribution is to make the resource denominator observable at the point of use, where G1i does not, and to expose the case where headline scores rise only because more inference-time compute is spent.

  1. What it does not measure Realised cost (discounts, committed use, self-hosting, batch tiers),

latency, capability, retries beyond the reliability rule, cost on any other benchmark.

  1. Unit USD per solved task at the stated reliability, with the price snapshot date. Companion

series in FLOP per USD.

  1. Primary source S13, live per-token pricing joined to per-benchmark results.
  2. Fallback or corroborating source S3 for hardware price-performance, with the caveat that S3

records release price rather than realised or rental price (independently verified as a recorded confounder), so a FLOP-per-dollar series built on it ignores the cost frontier labs actually pay. S35 is named and deliberately not used: it was search-verified only, not fetched, on 2026-07-30 (independently verified as its register status), and the decline rate it reports varies by roughly two orders of magnitude depending on which capability milestone is chosen (independently verified as a property of that analysis as described). Consequence: no inference-price decline rate is reportable without naming the benchmark and the milestone, so G2i always renders per benchmark and never as one global rate.

  1. Historical data availability Constrained. S13's historical series sits behind its Commercial

tier while the free tier carries current headline indices and pricing (both independently verified from its docs). Without that licence the series begins at first ingestion and daily snapshots are the only route to history. S3 supplies hardware history independently.

  1. Update cadence Daily price snapshot, monthly recompute. S3 updates on a scale of weeks

(independently verified as its stated cadence).

  1. Expected reporting lag Days for pricing; weeks for a new model to hold a stable benchmark

result at the required reliability.

  1. Can it move backwards YES, and should be expected to: a price increase; a reasoning-heavy

model raising tokens per solved task faster than price falls; a stricter reliability rule; a benchmark version change.

  1. Known biases and confounders List price is not paid price. Token accounting differs across

providers and hidden reasoning tokens are billed inconsistently. Currency and regional pricing. Provider-side caching moves effective cost with no price change. The reliability rule sets the level, so a k-of-k series and a single-attempt series are different indicators and a mixed series is invalid.

  1. Gaming and contamination risk MODERATE. By providers, through release-timed promotional

pricing or cost moved off the per-token line. By labs, through benchmark-specific optimisation that cuts tokens on the pinned benchmark only. Contamination cuts tokens per solved task, reading as an efficiency gain.

  1. Confidence grade MEDIUM. Raised by a licence covering S13's historical series, any buyer-side

realised-price source, standardised reasoning-token disclosure, or the same series across at least three benchmarks so milestone sensitivity is visible.

  1. Composite eligibility CONDITIONAL: admissible only as a denominator, only within one pinned

benchmark and reliability rule, and only under an S13 licence permitting derived publication. Never additive.

  1. Class and lead-lag Contextual. Leading: price and token efficiency move ahead of

deployment-scale effects.


G3i. Frontier Training Compute

  1. Operational definition Estimated training compute in FLOP of the highest-compute model with a

publicly known estimate in each period, plus the trailing twelve-month growth rate. Store (period, model_id, compute_FLOP, estimate_method, error_bar_if_published, estimate_revision_count).

  1. What it measures The scale of resources committed to frontier training, as estimated by a

third party.

  1. Why it signals progress toward general capability It does not, and that is why it is included.

Treating input growth as capability growth is one of the named failure modes; G3i exists so the input trend can be displayed beside the capability trend and be seen to diverge. Rising G3i with flat A-family and C-family indicators is the signature of diminishing returns to scale, and that reading is only available if the input series is present and clearly fenced.

  1. What it does not measure Capability, efficiency, cost, whether the compute was used well, or

undisclosed training runs.

  1. Unit FLOP, log scale, plus percent per year for the growth view.
  2. Primary source S2.
  3. Fallback or corroborating source S8 for cluster-scale corroboration, S24 for actor-level

compute concentration (both independently verified as listed downloadable archives). Neither substitutes: cluster capacity bounds what a run could have used, not what it did. S8's cadence is unknown (independently verified as unestablished), so it corroborates levels rather than timing.

  1. Historical data availability Deepest series in the roster. S2 covers notable models

historically in a downloadable CSV (independently verified as fetchable).

  1. Update cadence Weekly automated sweep upstream, major models targeted within roughly two weeks

(both independently verified as stated). Monthly ingestion is adequate.

  1. Expected reporting lag Roughly two weeks for a major model, longer where nothing is disclosed

and an estimate must be constructed.

  1. Can it move backwards YES: an estimate revised downward rewrites history without notice; a

model is reclassified; a counted run is found to be a continuation of an earlier one. estimate_revision_count exists so silent drift becomes visible.

  1. Known biases and confounders The figures are estimates from parameter and token counts rather

than lab disclosures, with wide error bars revisable without notice (independently verified as a recorded confounder). Systematic undercounting of undisclosed runs. A shift of spend from pretraining to post-training or to inference-time compute lowers the series while total resource commitment rises, so it can fall for reasons unrelated to restraint.

  1. Gaming and contamination risk LOW as data integrity: no party competes on this figure in a

corrupting way. HIGH rhetorically: it is the number most often presented publicly as evidence of capability progress, by parties with a commercial interest in that reading.

  1. Confidence grade MEDIUM. Raised by lab disclosure of training FLOP, published error bars on

every estimate, or a changelog of estimate revisions.

  1. Composite eligibility EXCLUDED, unconditionally. It is an input, and admitting it would let

capital expenditure raise the capability score, the central failure this instrument exists to avoid.

  1. Class and lead-lag Contextual. Leading: inputs are committed before results appear.

G4i. Score-vs-Inference-Compute Elasticity

  1. Operational definition Self-computed, within a single model family. Fix a benchmark b at

version v, one harness version and one prompt protocol. Let c = 1..m index the configurations that the family publishes which differ only in reasoning effort level or thinking-token budget, holding base weights, tool access and sampling parameters constant. For each configuration record s_c, its score on (b, v) in the benchmark's native scale, and x_c, the inference resource spent per task instance to obtain s_c, measured as USD per task at published list price on a stated snapshot date and summing input, output and billed reasoning tokens; where no price is published, x_c is output-plus-reasoning tokens per task and the unit class is recorded as such. Fit by ordinary least squares s_c = alpha + beta * ln(x_c) + e_c and set G4i = beta_hat, the fitted slope, in score points per natural-log unit of spend, that is per e-fold increase in spend. Symbols: s_c the configuration's score; x_c its spend per task; alpha the intercept, retained for reconstruction and never reported alone; beta the population slope and beta_hat its estimate, which is the stored value; e_c the residual; m the configuration count. The semi-log form is chosen over log-log because s_c is bounded, a log-log elasticity is undefined at zero score and unstable near saturation, and a slope expressed in score points stays commensurable with a threshold distance in score points, which is what field 15 needs. A secondary fit on logit(s_c), with a continuity correction of 0.5 / n_c applied to scores at the bounds where n_c is the item count, is stored beside the primary as a form-sensitivity check. fit_form is stored per observation and the two forms never appear in one series. Fitting procedure. All thresholds and constants in this field are specification parameters chosen here, not values read from any source: the minimum configuration count, the observation window, the continuity correction, the bootstrap count, and the reported percentiles. m >= 3 distinct configurations are required, three being the smallest count at which a slope has a residual degree of freedom. At m < 3 the value is stored as insufficient_configurations and is never interpolated, imputed, or borrowed from a sibling family. All configuration scores must come from the same benchmark version inside a stated window (specification parameter, default 30 days) and all prices from one snapshot date. No pooling across families, benchmark versions, harness versions, or configuration classes: a tool-enabled variant is a different configuration class from a thinking-budget level, and mixing the two invalidates the fit. This is FD-19's variant-identity assertion applied inside the fit rather than only along the series. There is no outlier rule and no configuration may be dropped, because the choice of which to drop decides the slope: every published configuration enters or the observation is void. Where the source publishes a confidence interval per configuration score, uncertainty on beta_hat comes from a parametric bootstrap of 2,000 draws per configuration from those intervals, refitting each draw and storing the 5th and 95th percentiles (bootstrap count and percentile choice: specification parameters). Where it does not, a heteroskedasticity-robust standard error is stored and the absence of source intervals is flagged on the observation. Pinning and storage. Store (family_id, base_model_id, benchmark, benchmark_version, harness_version, config_ids, config_class, m, x_unit, price_snapshot_date, fit_form, beta_hat, beta_se, beta_ci_low, beta_ci_high, r_squared, index_methodology_version, retrieved_at). Any change of benchmark version, harness version, configuration class, or upstream index methodology version opens a new segment, and the render shows a visible break at the boundary rather than interpolating across it.

  1. What it measures How strongly one family's score on one pinned benchmark responds to inference

spend, across the configurations that family actually ships. It is a derivative, not a level.

  1. Why it signals progress toward general capability It does not, and it is the roster's only

instrument for deciding whether something else's movement did. When the frontier series is a maximum over configurations, a family that ships a larger thinking budget produces a threshold crossing with no change of model, and the tracker would otherwise record a capability step where a spending step occurred. G4i is the only indicator able to attribute a named crossing to spend rather than to a better system, because it is the only one that varies resource while holding the model fixed. G1i varies training compute across models at fixed capability; G2i reports a cost level per model at a fixed reliability rule; neither reads the within-family response and neither can be evaluated at a specific crossing. It also converts FM-06 from hygiene into evidence: FM-06 treats configuration proliferation as a nuisance to be fenced by variant-identity assertions, which keeps the series clean by discarding exactly the cross-configuration information measured here. The confusion it prevents: spend-bought capability reported as earned capability, which is the work the RESOURCE denominator exists to do, at the one place the denominator is fenced out of the headline.

  1. What it does not measure Capability level; training compute, which is G1i; realised as opposed

to list cost; cross-vendor comparability, since effort labels are vendor conventions with no shared scale; the user value of the extra spend; latency, memory, or serving footprint; and anything at all about a family that publishes a single configuration.

  1. Unit Score points on the pinned benchmark's native scale per natural-log unit of USD per task,

at a stated price snapshot date. Companion view in score points per doubling of spend, equal to beta_hat * ln 2. Where price is unavailable the unit is score points per natural-log unit of tokens per task, and the two unit classes form separate series that are never plotted on one axis.

  1. Primary source S13, which runs its own evaluations across roughly 180 models and publishes

per-model cost and score with confidence intervals (independently verified as stated properties of the source). Licence constraint, handled explicitly rather than assumed away: S13 carries a proprietary licence, attribution is required, and redistribution needs a contract (independently verified from its documentation). Consequence for this indicator, stated as a rule the pipeline enforces. beta_hat and its interval are published as derived statistics with the required attribution; the per-configuration cost table that produced them is stored locally with redistributable = false and is never republished, in whole or as a plotted point cloud from which it could be read back. Where S1 carries both the configuration scores and a token basis for the same family, the fit is recomputed from S1 under CC-BY (independently verified as its licence) and the S1 value is the published one, with the S13 fit retained privately as a cross-check and any material disagreement between them recorded as an observation in its own right. S13's index methodology is versioned, which breaks comparability across time (independently verified as a property of the source), so index_methodology_version is pinned on every observation and no series crosses a version boundary.

  1. Fallback or corroborating source S1 for scores across reasoning variants at a pinned benchmark

version, continuous to daily and CC-BY (both independently verified as stated cadence and licence). S1 is the licence-clean path and the preferred one wherever its coverage of a family's configuration set is complete; per-family completeness of reasoning-variant coverage is uneven (inferred, not verified family by family). S3 as a directional sanity check on the spend axis only, carrying its recorded limitation inline: S3 records release price, not realised or rental price (independently verified as a recorded confounder), so it bounds nothing about what serving a configuration actually costs. Cost of this arrangement, stated plainly: the licence-clean source may lack the cost basis while the cost-bearing source is not republishable, so some families yield a private value only and the public series is thinner than the internal one. That asymmetry is displayed as a coverage count per period rather than hidden by omission.

  1. Historical data availability Shallow, and shallower than G2i. Families shipping three or more

scored reasoning configurations on one benchmark version are a recent pattern, so few families have a computable historical fit (inferred). S13's 7, 30 and 90-day series sit behind its Commercial tier while the free tier carries current headline indices and pricing (both independently verified from its docs), so without that licence the spend axis begins at first ingestion and daily snapshots are the only route to history. S1 backfills the score side wherever reasoning variants are recorded. The series begins at first ingestion.

  1. Update cadence Recompute per family on any new configuration, any score revision, or any price

change, with a monthly floor. The daily price snapshot is shared with G2i rather than duplicated.

  1. Expected reporting lag Days to weeks for a new family's configuration set to be scored across

enough configurations to reach m >= 3. Unbounded for a family that publishes one configuration, which remains insufficient_configurations indefinitely.

  1. Can it move backwards YES, by several mechanisms, none of which requires a model to change. A

price move on any single configuration re-fits the slope on its own. A newly shipped cheap configuration that scores well flattens beta_hat. Benchmark saturation compresses the score range at the top, so the slope falls toward zero from a ceiling effect rather than from any change in how spend converts to score. A benchmark version or harness change voids the fit and opens a new segment. An upstream index methodology version bump re-bases the inputs. A configuration score revised on rerun re-fits the whole family retroactively, so beta_hat is a revisable quantity and every stored value carries retrieved_at and its input revision count.

  1. Known biases and confounders List price is not paid price, and provider-side caching moves

effective cost with no price change. Reasoning-token billing is inconsistently disclosed across providers, which makes x_c the least standardised quantity in the roster. The configuration axis is not a scalar: effort levels are vendor labels, so beta_hat is comparable within a family and across time and is not comparable across vendors, and the render never places two vendors' slopes on one axis. m is small, on the order of three to five for current families (inferred, not verified family by family), so intervals are wide and one configuration carries most of the leverage. The functional form is a specification choice and a different form returns a different number, which is why fit_form is stored and the logit fit is published beside the semi-log one. Selection: the family chooses which configurations to expose, so the fitted range is the range a vendor decided to publish, not the range the system supports.

  1. Gaming and contamination risk MODERATE to HIGH, and asymmetric in the direction that matters.

By providers, through promotional pricing on the high-effort tier, which flattens the apparent slope and makes a spend-bought crossing read as earned. By labs, through publishing only a maximal configuration, holding m below 3 so the field 15 gate cannot fire at all. Mitigations, both runnable: m < 3 is recorded as UNVERIFIED-ELASTICITY and never as a pass, so non-disclosure cannot purchase a clean crossing; and promotional pricing appears as a price change with no score change, which the daily snapshot already records and can be flagged automatically. Contamination acts on the score side, lifting every configuration and shrinking the fitted range, which biases beta_hat toward zero and therefore toward clearing rather than tagging a crossing.

  1. Confidence grade LOW to MEDIUM. LOW for the level of beta_hat at m = 3; MEDIUM for its

sign and for the binary question of whether a named crossing is spend-explicable, which is the only use field 15 makes of it. Raised by standardised reasoning-token accounting across providers, per-configuration score intervals from a CC-BY source so the whole fit becomes publishable, the same fit across at least three benchmarks per family, or an S13 licence covering derived publication of the cost basis.

  1. Composite eligibility CONDITIONAL, as a GATE on other indicators' threshold events and never

as a summed component. Three parts, argued. Why not a component. The quantity is the response of score to resource. The RESOURCE denominator is constitutive of the stipulated target, so admitting an elasticity as an additive term would let a family that converts spend into score efficiently score higher on generality, inverting the denominator's purpose. A steep slope is not a capability and a flat one is not a deficiency. Why not simply excluded. Exclusion is what the roster already practises for the resource dimension, and it is precisely what leaves the headline blind here: the headline is a maximum over configurations, so a spending step enters it as a capability step with nothing positioned to object. An indicator that catches that has to act on the composite rather than sit beside it, which is why the verdict is CONDITIONAL rather than EXCLUDED. The gate, precisely. A threshold crossing recorded for family f on (b, v) is tagged SPEND-BOUGHT when all four conditions hold: (i) the crossing observation's configuration is not the family's default-deployed configuration; (ii) the family's cheapest published configuration on the same (b, v) scores below the threshold T; (iii) beta_hat for that family on that (b, v) is positive with a bootstrap interval excluding zero; and (iv) beta_hat * ln(x_cross / x_0) >= T - s_0, where x_0 and s_0 are the cheapest configuration's spend and score and x_cross is the crossing configuration's spend, so that the observed spend increase alone is sufficient under the fitted slope to account for the crossing. A tagged crossing is displayed in full and is not counted as a capability event in A1, A2, or the horizon set until a configuration at or below x_0 also crosses T; when that happens the tag clears and the crossing is dated to the same-cost observation rather than to the original one. Where m < 3 the crossing is tagged UNVERIFIED-ELASTICITY and is counted, tag visible, because absence of evidence is not evidence of purchase. The gate is asymmetric by construction: it can withhold or re-date a capability event and can never create one, and it never alters any indicator's stored value.

  1. Class and lead-lag Contextual. Leading: a family's slope is computable from its published

configuration set before any threshold crossing is recorded against it.


H1. Open-Weight Lag

  1. Operational definition For a pinned benchmark version, let t_closed(v) be the first date any

closed-weight model reached score level v and t_open(v) the first date any open-weight model reached the same v. H1 is t_open(v) - t_closed(v) in months, computed at several fixed v and reported as a small set rather than one figure. Requires a per-model weights_access field in {open, closed, restricted-open}, which the primary feed does not supply.

  1. What it measures Elapsed time for a frontier capability level to become available in

downloadable weights.

  1. Why it signals progress toward general capability It does not. Both inputs are capability

observations, but their difference is a diffusion property. Its contribution is interpretation-layer: a compressing lag means capability claims about "the frontier" apply to a far larger population of actors, changing what a level means in practice without changing the level.

  1. What it does not measure Capability; usability at scale (serving cost, licence terms,

hardware); safety consequences of diffusion; commercial adoption.

  1. Unit Months.
  2. Primary source S1, scores at a pinned benchmark version.
  3. Fallback or corroborating source S9 for the qualitative gap insight, whose text lags by roughly

two months (independently verified as its stated lag), with S4 for underlying scores. The binding constraint: the S4 CSV carries no open-weight flag (independently verified) and S9 publishes no download, so the split must be joined by hand, model by model. Specified join: maintain a local weights_access mapping keyed on the model identifier used by S1 and S2, populated from each release announcement, validated for pre-2025 open models against S27's historical archive (S27 is retired and open-weights-only, independently verified, so it serves as a historical membership list only). Maintenance cost, honestly: manual work at every release, roughly a handful of entries per month, and the most fragile element of the contextual family. It also embeds a judgment, since restricted-open licences are not open weights; each classification records a reason.

  1. Historical data availability Moderate. Scores backfill from S1; the weights-access join is

reconstructed retroactively and its accuracy degrades with age.

  1. Update cadence Monthly, gated by the manual join.
  2. Expected reporting lag Weeks for scores plus the join interval, longer where licence terms are

ambiguous.

  1. Can it move backwards YES, meaning the lag lengthens: a period with no open release at the next

level; a closed model advancing v against a static open field; reclassification from open to restricted-open, which removes a model from the open series and re-lengthens a computed lag.

  1. Known biases and confounders Benchmark choice determines the lag, which differs by domain. The

access category is genuinely fuzzy. Survivorship: a model released then withdrawn distorts the first-achievement date. Contamination affects open models differentially, since their data is more often documented and their evaluation more often self-run.

  1. Gaming and contamination risk MODERATE. By labs, through a benchmark-tuned open release that

appears to close the gap. By this tracker, through classification choices at the restricted-open boundary, which is why each carries a recorded reason.

  1. Confidence grade MEDIUM. Raised by an upstream weights_access field in S1 or S4, a

maintained public access registry, or H1 computed across at least three benchmarks to expose domain sensitivity.

  1. Composite eligibility EXCLUDED. It is a difference between two capability series, not a level;

summing it would double-count the underlying scores.

  1. Class and lead-lag Contextual. Lagging: the open event is by construction the later one.

H2. Frontier Concentration

  1. Operational definition For a pinned benchmark version and a defined frontier band (all models

within a stated score distance of the current maximum), compute a Herfindahl-Hirschman index over developer organisations holding band models, plus the count of distinct organisations in the band. Store (period, benchmark, benchmark_version, band_rule, org_count, hhi). Both are always reported together: the count is the plain-language view, the index the concentration view.

  1. What it measures How many independent organisations hold a model at or near the frontier, and

how unevenly frontier presence is distributed among them.

  1. Why it signals progress toward general capability It does not. Concentration is an ecosystem

property. Its contribution is interpretation-layer: a frontier held by one organisation makes every other indicator a measurement of one organisation's engineering, and makes the roster vulnerable to that actor's disclosure and evaluation choices.

  1. What it does not measure Capability; compute concentration, which S24 covers separately; market

share; revenue; actors able to reach the frontier but choosing not to publish.

  1. Unit Index, 0 to 1, dimensionless, plus a count of organisations.
  2. Primary source S1, using raw benchmark scores.
  3. Fallback or corroborating source S4, with an explicit preference against it here. Confounder 1

applies: S4's index is a latent-trait fit re-anchored as benchmarks saturate, so historical values move retroactively and cached figures silently drift (independently verified as a recorded confounder). A concentration series on a re-anchored index would restate its own history. Raw S1 scores are therefore the specified input, with S4 used only to sanity-check band membership at a single point in time. S24 corroborates at the actor level through compute and revenue rather than capability.

  1. Historical data availability Good. Backfills as far as the chosen benchmark's score history in

S1, each observation pinned to a benchmark version.

  1. Update cadence Monthly.
  2. Expected reporting lag Weeks, following S1 ingestion of a new frontier model.
  3. Can it move backwards YES, meaning concentration rises: one organisation opening a score gap

wide enough to empty the band; an acquisition merging two organisations; a band-rule change; withdrawal of a model from evaluation.

  1. Known biases and confounders The band rule sets the answer, and a narrow band produces a

volatile series. Organisation attribution is ambiguous for joint releases, subsidiaries and compute partnerships. Coverage bias: an unevaluated frontier model reads as absent. Benchmark saturation compresses the band and manufactures apparent de-concentration.

  1. Gaming and contamination risk LOW to MODERATE. Not directly valuable to game. Indirect risk

from a benchmark-targeted release buying band membership without frontier capability.

  1. Confidence grade MEDIUM. Raised by a stable organisation-attribution taxonomy, computing H2 on

an unsaturated benchmark, or publishing band-rule sensitivity across at least three widths.

  1. Composite eligibility EXCLUDED. An ecosystem structure measure, not a capability level.
  2. Class and lead-lag Contextual. Lagging.

Note on the geopolitical split (gap G5). A China-versus-United-States frontier split is derivable from S4's country filter (independently verified as present) or from S1 joined to a developer-country mapping. It is not a separate indicator: no dedicated tracker with a machine-readable feed exists (gap G5, established), and more importantly it is largely collinear with H1, because most leading Chinese models are open-weight, so the split would restate the open-weight lag under a geopolitical label rather than add independent evidence. Available as a dashboard filter on H1 and H2.


I1. Capability-Controllability Divergence

  1. Operational definition For a pinned harness scoring the same models on both a capability suite

and a controllability suite under one prompt protocol, compute pct_rank(capability) - pct_rank(controllability) per model, per harness version, on within-suite normalised percentile ranks. Reported as a distribution across the evaluated set with the harness version and the date the harness last ran. Never added to, subtracted from, or weighted into any capability figure.

  1. What it measures Whether models scoring higher on capability suites also score higher on

controllability suites, on the same models, under the same prompting.

  1. Why it signals progress toward general capability It does not, and the architecture bars it

from doing so: safety and controllability are tracked separately and never summed. Its contribution is to make one narrative falsifiable, namely that capability and controllability advance together. A widening divergence sits beside the capability profile and may change a reader's interpretation of it, never its value.

  1. What it does not measure Real-world harm, deployment safeguards, misuse, alignment in any deep

sense, anything about a model the harness did not run.

  1. Unit Percentile-rank difference, dimensionless, reported as a distribution with spread.
  2. Primary source S16, which runs models itself with prompt-level transparency across capability

and safety boards (independently verified as its stated design), and is the only verified family evaluating both dimensions on the same models under one protocol.

  1. Fallback or corroborating source None verified. Gap G1 is adjacent: no live public leaderboard

exists for calibration or run-to-run variance, and the same absence holds for paired capability-and-controllability evaluation outside S16. Cost: no second opinion, and no recovery path if S16 stops.

  1. Historical data availability Whatever S16 has already published, which is the ceiling. Its

interactive site yielded no fetchable content on 2026-07-30 (independently verified), so history comes from the repository and extraction is manual.

  1. Update cadence None reliable. S16 entered maintenance mode on 2026-06-01 (independently verified

as a status quoted verbatim from the source), so coverage of models released after that date is not something the tracker can rely on.

  1. Expected reporting lag Unbounded. Under maintenance mode a new frontier model may never be

scored on both suites.

  1. Can it move backwards YES: a harness version change; percentile renormalisation when the

evaluated set changes, which moves every model's rank without any model changing; a capability-only release cohort shifting the distribution.

  1. Known biases and confounders Percentile ranks are relative to the evaluated set, so the

indicator is sensitive to coverage, and coverage under maintenance mode reflects maintainer capacity rather than the field. Safety suites measure refusal and behavioural compliance under test conditions, a narrow stand-in for controllability. Prompt sensitivity is high on both sides.

  1. Gaming and contamination risk MODERATE. By labs, through tuning to published safety suites,

which raises the controllability rank without changing behaviour outside the suite. The harness runs models itself with prompt transparency (independently verified), which limits self-reporting risk but not train-to-the-test risk.

  1. Confidence grade LOW. Raised by a maintained third-party harness scoring both dimensions on

the same models, a government or intergovernmental body publishing paired machine-readable results, or S16 exiting maintenance mode.

  1. Composite eligibility EXCLUDED, unconditionally and permanently. Tracked, never summed.
  2. Class and lead-lag Contextual, under the dimension I designation of tracked and never summed.

Lagging.

Consequence of the dying source, stated plainly. I1 rests on a single source family in maintenance mode as of 2026-06-01 (independently verified). Two consequences are accepted rather than papered over. It is displayed with a source-health badge and an explicit last-harness-run date, so a reader sees staleness instead of a stale value presented as current. If the harness publishes nothing covering a new frontier generation, I1 is marked DORMANT and the tracker states that the capability-controllability relationship is unmeasured for that generation. A dormant indicator saying nothing is preferable to substituting lab self-reported safety-card figures, which no third party can reproduce and which section 5 rejects for that reason.


#

Twelve measures, each rejected on a measurement ground rather than an ideological one.

Measure: capital expenditure and training spend. Why it is tempting. Documented, often disclosed in aggregate, correlated with frontier releases. Why it is rejected. An input. Spend rises with hardware prices, accounting choices and competitive positioning with no capability consequence, and it cannot fall in response to a capability plateau. What it would corrupt if included. A capital-raising cycle would register as capability progress, destroying the instrument's ability to show stagnation during a boom. Permitted alternative use. Dashboard context in a resources panel beside G3i, never summed.

Measure: company valuations and funding rounds. Why it is tempting. High-frequency, widely reported, reads as a market verdict on capability. Why it is rejected. It prices expectations of future returns, set by parties holding the asset. A forecast in the clothes of a measurement. What it would corrupt if included. Market sentiment inside the capability layer, guaranteeing the tracker moves with narrative rather than against it. Permitted alternative use. Nothing in the capability layer. S24 supplies actor-level revenue and funding for the concentration discussion only.

Measure: press attention and search interest. Why it is tempting. Free, instant, and a real signal of something. Why it is rejected. It measures salience. Attention rises on launches, controversies and litigation with no relation to measured competence. What it would corrupt if included. The tracker would become a lagging indicator of its own media environment. Permitted alternative use. Nothing, not even context: its presence invites attention to be read as progress.

Measure: parameter counts alone. Why it is tempting. Long historical availability, simple, comparable across a wide span. Why it is rejected. The parameter-to-capability relationship is unstable across architectures, training regimes, sparsity and distillation, and frontier labs have largely stopped disclosing the figure. What it would corrupt if included. Architectural fashion would look like progress, and a shift toward smaller efficient models would read as regression. Permitted alternative use. Interpretation layer only, as one input to G1i's compute estimate, since S2's estimates derive partly from parameter and token counts (independently verified as their stated method).

Measure: deployment and user numbers with no capability evidence. Why it is tempting. Concrete, occasionally disclosed, the most intuitive answer to whether AI is getting better. Why it is rejected. Adoption follows distribution, pricing, bundling and default placement. The architecture already places deployment scale outside the capability figure. What it would corrupt if included. A bundling deal would register as a capability jump, and the tracker could not report capability stagnation during rapid adoption. Permitted alternative use. Dashboard context in a diffusion panel drawing on S7, with its recorded confounder inline: S7 measures one laboratory's own user mix, so movements reflect that laboratory's product and customer changes as much as any change in the economy (independently verified as a recorded confounder).

Measure: unreplicated laboratory demonstrations and cherry-picked demos. Why it is tempting. The most vivid evidence available, arriving first, and often showing something real. Why it is rejected. Best-of-n selection with an unreported n is not a measurement. Success rate, prompt, scaffold and human intervention are typically undisclosed, and nothing is reproducible by a third party. What it would corrupt if included. The reliability axis would be replaced by a maximum-observed-performance axis, and the headline would move on marketing calendars. Permitted alternative use. Interpretation layer only, as a hypothesis generator for future roster additions, explicitly labelled a demonstration.

Measure: human-preference Elo as a capability measure. Why it is tempting. High volume, continuous updates, broad coverage, one legible number with an established audience. Why it is rejected. S18 measures preference, not capability (independently verified as a property of the source), and preference tracks formatting, verbosity, tone and refusal behaviour. Vote data and the rating pipeline are operator-controlled, the operator changed name and host during verification, and pre-release anonymous testing advantages labs that iterate against the board (all independently verified as recorded properties). A pairwise rating also has no fixed zero across time, so a level is not comparable across periods. What it would corrupt if included. Style optimisation would register as capability progress, and one operator's pipeline changes would move the headline. Permitted alternative use. Dashboard context in a preference panel, labelled preference, with the no-fixed-zero caveat displayed. The same reasoning keeps S30's robotics Elo out of F1 as a level.

Measure: expert-survey AGI timelines. Why it is tempting. Published, citable, and answers the question readers actually ask. Why it is rejected. Aggregated opinion. It responds to recent salience and to question wording, and it cannot register a capability plateau independently of respondents noticing one. What it would corrupt if included. Opinion inside the data layer, letting the headline drift with the field's mood. Permitted alternative use. Forecasting layer only, presented as opinion, with no path into any capability figure.

Measure: prediction-market prices as capability data. Why it is tempting. S23 has a documented REST API needing no key for reads, with full order-book history (independently verified), making it the easiest live feed in the register to ingest. S25's questions are widely cited. Why it is rejected. S23 is play money with thin liquidity, which leaves individual long-horizon questions poorly calibrated (independently verified as a recorded property). S25's API is gated: three probes returned HTTP 403 with an authentication-required message on 2026-07-30 (independently verified), so it is not citable as live without a token. Both price beliefs about future events rather than measure present capability, and both resolve on question wording rather than on an instrument. What it would corrupt if included. The tracker would become a belief aggregator, and its easiest feed to ingest would become its most load-bearing. Permitted alternative use. Forecasting layer only, fenced, with the play-money and liquidity caveats displayed for S23 and the gating status recorded for S25.

Measure: any composite published by a party with a commercial stake in the result. Why it is tempting. Such composites are usually the best-resourced, most current and most documented artefacts available. Why it is rejected. The party choosing the weights benefits from the outcome, and weight choice is where a composite's conclusion is decided. Methodological transparency does not remove the incentive. What it would corrupt if included. A vendor's priors laundered into the headline, the failure the brief names explicitly. Permitted alternative use. Component sub-scores are usable where independently run and separately reported; the headline index never is. This is the operative rule for S13: its per-benchmark results and pricing are admissible under licence, its composite index is not.

Measure: self-reported model-card scores no third party can reproduce. Why it is tempting. First available, covers the newest models, often the only number in existence for a frontier release. Why it is rejected. Prompt, scaffold, sampling parameters, attempt count and grader are all chosen by the reporting party, and re-running is not possible from the card. A self-reported figure cannot be promoted to an independently run series by citation. What it would corrupt if included. Every indicator's history would become a mixture of audited and unaudited observations, and the composite would favour labs that report generously. Permitted alternative use. Dashboard context in a marked self-reported column with the auditability class displayed, pending an independent run. Also usable as a discrepancy detector: a large persistent gap between a self-reported figure and an independent rerun is itself a recordable observation.

Measure: benchmark scores taken from aggregator blogs and secondary write-ups. Why it is tempting. Fast, tidy, pre-assembled into tables, easy to cite. Why it is rejected. Provenance is unrecoverable. The counter-example inherited in the brief is a briefing document whose benchmark figures trace to a commercial blog and a personal post rather than to any evaluation artefact. Two candidate aggregators in the register are already gone (S26 retired, S29 rendering no extractable data on 2026-07-30, both independently verified), so the practice is fragile as well as unsound. What it would corrupt if included. It would break the requirement that every quantitative claim carry a source and a verification class, since a laundered number has no recoverable class. Permitted alternative use. Nothing. Aggregator write-ups may point to a primary artefact, and the artefact is then cited.

Held inside the dashboard but outside the composite

Eight roster entries are displayed and never summed as terms.

  • G3i Frontier Training Compute. An input; summing it would let resource commitment raise the

capability figure.

  • H1 Open-Weight Lag. A difference between two capability series, so summing double-counts the

underlying scores. Also gated on a hand-maintained weights-access join.

  • H2 Frontier Concentration. An ecosystem structure measure, not a capability level.
  • I1 Capability-Controllability Divergence. Safety and controllability are tracked separately by

architecture, and I1 rests on a source family in maintenance mode.

  • E1 Real-Task Win Rate vs Human Expert. Summed only inside one upstream methodology-version segment

and under a licence permitting derived publication; dashboard-only otherwise.

  • E2 Benchmark-to-Reality Gap. Structurally unsummable: no quantified value exists by construction,

and it acts on the composite only by constraining interpretation.

  • F1 Simulated Embodied-Task Success Rate. Annual cadence, a suite that changes between editions,

and a simulation-versus-reality category mismatch. Shown as a standalone downward-only floor indicator with its environment class labelled, under a name that states the simulation, and with no embodied-competence claim derivable from it.

  • G4i Score-vs-Inference-Compute Elasticity. Never a term, because an elasticity of score to

resource is not a capability and adding it would let efficient conversion of spend into score raise a generality figure. It acts on the composite as a gate: a threshold crossing is withheld or re-dated when the fitted slope shows the crossing is explicable by the spend increase alone. The gate can only subtract or delay a capability event, never create one.

Source concentration across the contextual family

Stated as counts, not as a judgment. Each indicator's primary feed resolves to one upstream operator. Operators are named by the source IDs they own in the register.

IndicatorPrimary sourceUpstream operatorCorroborating sources and their operators
E1S13Artificial Analysis (S13)S20 (OpenAI)
E2S22Epoch AI (S1, S2, S3, S4, S8, S9, S22, S24, S35, S36, S37)S5 (METR), S12 (Scale AI)
F1S21Stanford BEHAVIOR (S21)none verified; S30 feed not found
G1iS1 joined to S2Epoch AIS36 (Epoch AI, literature only)
G2iS13Artificial AnalysisS3 (Epoch AI)
G3iS2Epoch AIS8, S24 (both Epoch AI)
G4iS13, published value recomputed from S1 where possibleArtificial Analysis, with Epoch AI as the licence-clean pathS3 (Epoch AI)
H1S1Epoch AIS9, S4 (both Epoch AI), S27 (Hugging Face, retired archive)
H2S1Epoch AIS4, S24 (both Epoch AI)
I1S16Stanford CRFM (S16)none verified

Counts. Ten indicators in dimensions E to I. Their primary feeds resolve to four upstream operators: Epoch AI (E2, G1i, G3i, H1, H2, and the licence-clean recomputation path for G4i), Artificial Analysis (E1, G2i, G4i), Stanford BEHAVIOR (F1), and Stanford CRFM (I1). Counting the two Stanford entities as one institution gives three. Corroboration adds four further operators (METR, Scale AI, OpenAI, Hugging Face), none of which is a primary feed for any indicator in this family. Two indicators, F1 and I1, have one operator and no verified corroborating source. Six of the ten depend on Epoch AI either as primary or as sole corroborator.

#

Scope note. This section is the methodology layer as defined in section 3.4. It contains no observation values, no benchmark scores, no dates for any capability event, and no probability of any capability event. Every numeric quantity appearing below is a specification parameter chosen here (a weight, a threshold count, a window length, an interval width), labelled as such, and is not a claim about the world. Where a property of a source is asserted, it carries a source ID and a verification class from the six-level set fixed in CRITERIA C8.

Common notation, used by all four methods. Let i index roster indicators, t index observation dates, and x_{i,t} be the raw value of indicator i at t in its own unit. Let z_{i,t} be its normalised value, w_i its weight, V_i its verification class, c_i its confidence grade under the house band convention (sBs/AGI-handbook/bible/SPINE.md), and age_{i,t} the days since the underlying source last published a new value. Let E be the set of composite-eligible indicators per section 4: unconditionally A1 and C2; conditionally B1, B2, C1, D1, E1, G1i and G2i under their stated gating rules. Let X be the excluded set: A2, C3, D2 and D3 while declared gaps G1, D4, E2, F1, G3i, H1, H2, I1. No method may move an indicator from X into a summed capability figure; the methods differ only in what they do with E, and in what role they give X.


Method A: transparent weighted composite index

  1. Formulation. A single scalar. For each release date t, compute

I_t = 100 * sum_{i in E'} w_i * z_{i,t}, where E' is E restricted to indicators whose conditional gates are open at t, z_{i,t} in [0,1] is the normalised value with sign oriented so that higher always means more capability, w_i >= 0, and sum w_i = 1 after renormalisation over E'. Sign orientation is explicit and required: A1 (headroom) and B1 (novelty gap) fall as capability rises, so they enter as 1 - z; G1i and G2i are resource denominators, so they enter only as the reciprocal-oriented efficiency term z_eff = 1 - z_cost, never additively as cost. Dimension weights are set first (a dimension budget), then split equally inside each dimension, so that a dimension with four indicators does not outvote a dimension with one. Baseline dimension budget, all figures specification parameters: A 0.25, B 0.25, C 0.30, D 0.10, E 0.10, with F, G, H and I at zero because their indicators sit in X or enter only as denominators.

  1. Required inputs. A1, C2 unconditionally. B1, B2, C1, D1, E1 when their section 4 gates open.

G1i and G2i only through the efficiency orientation above, and only inside one pinned benchmark and reliability rule for G2i. D4 is consumed as a publication-side annotation, not as a term. A2, C3, D2, D3, E2, F1, G3i, H1, H2 and I1 are not consumed.

  1. Normalisation. Two candidates, and the comparison matters more than the choice.

Min-max: z_{i,t} = (x_{i,t} - min_i) / (max_i - min_i) over a declared reference window. It produces a bounded, legible number, and it hands the entire result to the choice of window endpoints: a single early observation sets min_i for all time, and re-basing the window changes every historical value of the index. Z-score: z_{i,t} = (x_{i,t} - mu_i) / sigma_i, then squashed to [0,1] by a logistic or by clipping at plus or minus three standard deviations. It is robust to endpoint choice and destroys interpretability, because the unit becomes "standard deviations of this indicator's own history", which differs per indicator and changes as history accumulates. Neither is neutral. Specified requirement if A is used at all: publish the index under both normalisations side by side at every release, and treat any month in which the two disagree in direction as a non-result under the section 7 robustness rule.

  1. Missing data. Three admissible policies, ranked. (a) Renormalise over observed indicators

only, which is the default and which silently reweights: dropping a D-dimension indicator raises the effective weight on A and C, so the index can move purely because a source went quiet. (b) Carry forward the last observation with a staleness flag at 90 days (sBs/fow/kb/INDEX.md), which freezes the index against real movement. (c) Publish no index for that release. The specification requires (a) plus a mandatory published delta showing the index recomputed on the previous release's observed set, so reweighting effects are separated from value effects. Indicators past the 90-day staleness flag are excluded from E' rather than carried.

  1. Uncertainty. Propagated by Monte Carlo over three sampled quantities: per-indicator

measurement error (a distribution declared per indicator, widest where V_i is estimated or self-reported), the weight vector (Dirichlet draws around the baseline budget), and the normalisation choice (Bernoulli over min-max versus z-score). Ten thousand draws per release, a specification parameter. The published quantity is the 5th to 95th percentile band, never the point value alone. Confidence grade of the composite inherits the lowest grade among contributing indicators, since a composite cannot be more trustworthy than its weakest term; several core indicators carry MEDIUM at 12 months and LOW at 36 in section 4, so the composite's grade is capped there.

  1. Backward movement. Possible, and structurally muted. Every contributing indicator can fall

per section 4, so I_t can fall. But averaging is a suppression device: one dimension regressing while three advance yields a rising index, which is exactly the reading the brief forbids the instrument to produce automatically. A can therefore register regression only when regression is broad. It cannot register a narrow regression, and it has no mechanism at all for reporting that a dimension has not moved.

  1. Benchmark saturation. The failure case. When a component saturates, x_{i,t} approaches its

ceiling, z approaches 1, and the term contributes a constant. The index then keeps rising only through the unsaturated terms, while the reader sees a smooth series with no signal that a quarter of its basis has gone inert. Worse under min-max, where max_i is often the declared ceiling, so the saturated term pins at exactly 1 forever. Required mitigations if A is used: a per-term contribution-to-change decomposition published each release; automatic retirement of any term whose z has exceeded 0.95 for two consecutive quarters (specification parameter), handled by A1's documented basket rotation and chain-link splice; and publication of the count of live versus inert terms beside the index.

  1. Correlated indicators. Unhandled by construction, and this is A's deepest defect. B1 and B2

are transforms of the same paired observations. C1 and C2 are two thresholds on one logistic fit published by S5. A1 and A2 both derive from S1 headroom. Summing them counts one latent movement several times, and the weights then read as statements about importance when they are in fact statements about redundancy. Minimum required treatment: publish the pairwise correlation matrix over the shared history, collapse any pair above 0.85 (specification parameter) into a single term before weighting, and disclose the collapse. A tracker that does this has partially reinvented Method C without its uncertainty machinery.

  1. Discontinuous jumps. A jump in one term moves the index by w_i * delta z_i and nothing

distinguishes a genuine capability step from a harness change, a re-graded submission, or a basket rotation. Required guard: any single-release move exceeding two published index points (specification parameter) triggers an attribution note naming the contributing terms and their provenance changes, and the release is held until the attribution is written.

  1. Strengths. Legible in one sentence. Cheap to compute and to explain. Trivially auditable

arithmetic once weights are published. Gives the lay reader something to hold and the press something to quote. Sensitivity analysis on it is straightforward, because the levers are few and named. It also has real diagnostic value as a reference method: a headline that disagrees with a disclosed composite is a finding worth investigating.

  1. Weaknesses. The weights are unfalsifiable: no observation can show that A should have been

0.25 rather than 0.30, so the index encodes a preference and reports it as a measurement. Normalisation silently dominates, because the choice of reference window can move the series more than the underlying data does. Correlated terms double-count. Saturation is invisible. Averaging suppresses narrow regression, which is the movement most worth catching. A single scalar invites false precision, and the false precision compounds when the number is quoted without its band. It also has no place to put the excluded indicators, so I1 (safety divergence) and D4 (regression count) fall out of the headline entirely, which contradicts the layer separation in section 3.4.

  1. Manipulation susceptibility. By the maintainer, through weight and window selection: the

weight vector can be chosen after seeing the data, and no reader can detect it unless weights are version-controlled with timestamps predating the data they act on. By frontier developers, through selective submission on the A1 basket and public-split training, which raises z on the highest-weighted dimension. By benchmark operators, through version scheduling, which changes which terms are live. By a hostile commentator, through re-weighting: because A publishes the ingredients, anyone can produce an alternative index and present it as the tracker's own. That last one is a cost of transparency and should be accepted, not fixed.

  1. Subjective judgment. Degree: HIGH. Enumerated: choice of E' membership; the dimension

budget; the within-dimension split rule; sign orientation for each indicator; normalisation family; reference window endpoints; the squash function and clipping bound under z-score; missing-data policy; the staleness cut-off; the saturation retirement threshold; the correlation collapse threshold; the jump-attribution trigger; the per-indicator measurement-error distributions; the number of Monte Carlo draws; the decision to renormalise rather than suspend publication.

  1. Audit exposure. Every item in field 13 lives in one versioned parameter file in the

repository, with a semantic version and a change log entry naming the date, the reason, and the release from which it takes effect. Weights are committed before the data they act on, and the commit timestamp is the evidence. A single command recomputes the whole index from raw archived snapshots. Each release publishes the leave-one-out table, the both-normalisations pair, the term contribution decomposition, and the correlation matrix. A reader can therefore reproduce any figure and see which choice produced it.

  1. Public legibility. A non-technical reader reads I_t as a percentage of the way to AGI. That

misreading is not incidental; the form invites it. They will also read a movement from one release to the next as real change when it may be reweighting under missing data, and they will read the number as comparable across years when re-basing has changed its meaning. They will drop the uncertainty band on quotation. And they will read a rising index as evidence that nothing has regressed, because the form has no way to say otherwise.


Method B: distance-to-threshold, capability-frontier model

  1. Formulation. No aggregation of levels. Instead a pre-registered set of thresholds

T = {(i, tau_i, dir_i)}, where tau_i is a value on indicator i's own scale and dir_i says whether the threshold is met by exceeding or by falling below tau_i. Define met_{i,t} = 1 if the threshold condition holds at t, else 0. POSITION is the count P_t = sum_i met_{i,t} reported as P_t of |T|. For unmet thresholds, distance is reported in the indicator's own unit and, separately, in normalised form d_{i,t} = (tau_i - x_{i,t}) / (tau_i - x_{i,0}) against a pre-registered origin x_{i,0}, so that d runs from 1 at registration to 0 at the threshold. Dimension robustness is tested by R_t = min over dimensions of (thresholds met in that dimension / thresholds set in that dimension), published beside P_t. No sum, no weight, no single scalar. R_t as written here is the naive form and is unsafe to publish: it is undefined when a dimension's thresholds are all UNKNOWN and it is dominated by whichever dimension has the least quantified noise. Method D field 1 states the scope-gated form that governs anywhere a floor is published, and Method B's R_t is retained in this field only as the unrepaired expression it repairs.

  1. Required inputs. Thresholds are set on A1, C1, C2, D1, E1 and, when funded, D2 and D3. B1 and

B2 supply a gate rather than a threshold: no A-dimension threshold counts as met in a period where the contamination ratio has not been observed at least once, which prevents a met threshold resting on a possibly contaminated score. A2, C3 and D4 are used as instrument-health and regression qualifiers on whether a met threshold stays counted. G1i and G2i carry their own resource thresholds reported separately and never counted into P_t. F1 carries one embodiment threshold at annual granularity, flagged as simulation-class per its section 4 entry. E2, G3i, H1, H2 and I1 carry no thresholds.

  1. Normalisation. Almost none, which is the method's main advantage. Thresholds are stated in

each indicator's native unit (minutes for C1 and C2, percentage points for D1, a dimensionless ratio for A1), so no cross-indicator commensuration is performed and no normalisation window can contaminate the result. Normalisation appears only in the reported distance d_{i,t}, which is bounded by construction against the pre-registered origin and is a presentation device, not an input.

  1. Missing data. A threshold with no current observation is reported as UNKNOWN, a third state

beside met and unmet, and P_t is reported as a pair: the count met, and the count unknown. This is the method's cleanest property, because a missing observation cannot inflate the headline; it can only widen the unknown count, which is visible. A threshold whose source has passed the 90-day staleness flag reverts from met to UNKNOWN rather than staying met, so a dead source degrades the headline instead of freezing it flatteringly.

  1. Uncertainty. Per threshold, an interval on x_{i,t} is compared with tau_i and the outcome

is reported as met, unmet, or AMBIGUOUS when the interval straddles tau_i. P_t is therefore published as a range [P_low, P_high], counting ambiguous cases out and in respectively. Distances carry the same interval. No Monte Carlo is required, which makes the uncertainty statement easier for a reader to check than Method A's band.

  1. Backward movement. Native. A met threshold reverts when the observation falls back below

tau_i, and reversion is a headline event by construction rather than a smoothed contribution. R_t falls immediately when one dimension loses its only met threshold. This is the strongest regression sensitivity of the four methods.

  1. Benchmark saturation. Handled partially and honestly. A saturated benchmark still produces a

met threshold, which is correct: the capability was demonstrated. The risk is that P_t stops moving because the remaining thresholds are all far away, and stagnation of the count is then ambiguous between real plateau and thresholds set too coarsely. A2 is the required companion reading: P_t static with short saturation half-life indicts the threshold set; P_t static with lengthening half-life indicts the systems. Thresholds on retired benchmarks are migrated only through A1's documented chain-link splice, and the migration is logged as a threshold-set version bump.

  1. Correlated indicators. Substantially mitigated, because counting is not summing: two

correlated indicators crossing thresholds contribute two counts, which overstates breadth but does not compound multiplicatively as a weighted sum does. Required treatment: thresholds are set at most one per latent construct per dimension, so C1 and C2 do not both count (C2 carries the threshold, C1 is reported as context per its section 4 entry), and B1 and B2 gate rather than count.

  1. Discontinuous jumps. Absorbed well. A jump moves a distance figure and may flip one count. It

cannot move a headline scalar by an arbitrary amount, because the headline is an integer over a fixed denominator. The threshold-set version and the date of every flip are published, so a flip can be audited against a provenance change.

  1. Strengths. No weights. No commensuration. Missing data cannot flatter. Regression is a

first-class event. The reading is falsifiable in the only sense available here: a pre-registered threshold either was or was not crossed, and the record is checkable years later. It is also cheap for one maintainer, because most of the work is done once at registration.

  1. Weaknesses. The thresholds are the whole method, and setting them is a large, single act of

judgment. Coarse granularity: an integer count over a small denominator moves in visible steps and says nothing between steps, which the distance figures only partly repair. It has no rate component at all, so it cannot answer the question a progress tracker mainly exists to answer. Threshold sets age: a set registered against today's instruments becomes uninterpretable as those instruments saturate, and migration reintroduces splice judgment. Finally, an unmet threshold's distance is reported on a scale whose origin was itself chosen.

  1. Manipulation susceptibility. By the maintainer, through post-hoc threshold setting or through

quiet threshold migration that resets an inconvenient distance. This is the method's one serious attack surface and the reason for the registration discipline in field 13. By frontier developers, through targeting a published threshold: once tau_i is public, it becomes an optimisation target, and a system tuned to just clear it produces a met threshold without the general capability the threshold was meant to indicate. By benchmark operators, through version changes that alter what tau_i means on the new version line.

  1. Subjective judgment. Degree: HIGH, and concentrated at one auditable moment rather than

diffused. Enumerated: which indicators carry thresholds; the value of each tau_i; the direction of each; the pre-registered origin x_{i,0}; the denominator |T|; the one-threshold-per-construct rule and which of a correlated pair carries it; the B-dimension gating rule; the reversion rule on staleness; the ambiguity rule for straddling intervals; the migration procedure when a benchmark retires; the definition of dimension robustness R_t. The crucial procedural detail: thresholds must be set and committed BEFORE the data is examined, and versioned, or they are not thresholds at all but a post-hoc narrative device that will always find the story the maintainer already believed. Registration means a signed, timestamped commit of the full threshold set plus its rationale, and a rule that any later change is an additive version bump which publishes both the old and new counts in parallel for at least four releases (specification parameter).

  1. Audit exposure. The threshold file is the audit object: one machine-readable file, semantically

versioned, with commit timestamps that a reader can compare against the publication dates of the observations. Every flip event is logged with the observation that caused it, its source ID and its verification class. Every migration publishes parallel old and new counts. A reader who suspects post-hoc setting checks the commit date, which is the only defence that actually works.

  1. Public legibility. A reader reads "P of |T| thresholds met" as a fraction of the way to

AGI, which is worse than Method A's misreading because the denominator looks authoritative. They will also assume the thresholds are consensus definitions of AGI capability rather than one maintainer's registered guesses, and they will read a static count as "no progress" when distances may be closing fast underneath it. Required reader-facing text: the denominator is a specification choice, not a count of the requirements for AGI, and the document nowhere asserts an agreed definition of AGI.


Method C: latent-factor statistical model

  1. Formulation. Estimate one unobserved general-capability factor from many correlated

observations, preserving uncertainty. Two admissible forms. Factor-analytic: for model m on indicator i, y_{m,i} = mu_i + lambda_i * theta_m + eps_{m,i}, eps ~ N(0, psi_i), where theta_m is the latent capability of model m, lambda_i the loading of indicator i, and psi_i its residual variance; the frontier series is the maximum posterior theta_m among models observed at t. Item-response: for a binary or binarised item j, P(correct | theta_m) = g_j + (1 - g_j) * logistic(a_j * (theta_m - b_j)), with discrimination a_j, difficulty b_j and guess floor g_j, fitted jointly over models and items. Estimation by hierarchical Bayes so that theta_m carries a posterior interval rather than a point, with priors on lambda, a and b declared in the methodology file. The headline is the posterior median of frontier theta with its credible interval.

  1. Required inputs. Item-level or benchmark-level scores for the whole eligible set: A1's basket

constituents (raw per-benchmark scores from S1, not the A1 median), B1 and B2's paired splits, C1 and C2, D1, E1. G1i and G2i cannot enter the factor: they are resource denominators, and admitting them would let falling cost load onto capability, which section 4 prohibits unconditionally for G3i and conditionally for G1i and G2i. A2, C3 and D4 sit outside the model and are used to test it: D4's regression events are the natural check on whether a single monotone factor is even the right object.

  1. Normalisation. Endogenous, which is the method's strongest property. Loadings absorb scale

differences, so no min-max window and no z-score window is imposed by the analyst. The cost is that the latent scale has no natural units and no fixed zero, so it is identified only up to an affine transform, and identification requires an anchor choice (fix mu and scale at a reference model, or fix the factor's mean and variance across the estimation sample). The anchor choice is where the retroactive drift enters.

  1. Missing data. Handled properly and this is a genuine advantage: an unobserved cell contributes

no likelihood term, theta_m is estimated from what exists, and the posterior widens where coverage is thin. No imputation, no carry-forward, no silent reweighting. The requirement is that coverage is published per model as the count of observed cells and the resulting interval width, because a wide interval from thin coverage otherwise looks like genuine ambiguity about capability.

  1. Uncertainty. Native. The posterior over theta_m is the uncertainty statement, and it widens

with thin coverage, low-discrimination items and near-ceiling items. Published as posterior median plus a 90 percent credible interval (specification parameter). Two uncertainties the posterior does NOT contain and which must be published separately: uncertainty over the anchor choice, and structural uncertainty over the one-factor assumption itself. Both are model-form uncertainty and no posterior computed under the model can express them.

  1. Backward movement. Possible for a given model's theta_m when new items are added on which it

performs poorly, and possible for the frontier series when a newer model fits below an older one. But the frontier maximum is monotone by construction if it is taken as a running maximum, so the specification requires the frontier series to be the current-period maximum, not the running maximum, so it can fall. Regression at the item level is visible only in residuals, which is why D4 must be published beside theta rather than folded into it.

  1. Benchmark saturation. The method's characteristic behaviour and its fatal property for a

tracker. Saturated items carry little information about high theta, so the fit re-anchors as harder items enter, and historical theta values move retroactively. Prior art is explicit here: S4, the Epoch Capabilities Index, is a latent-trait fit with public MIT-licensed fitting code and CC-BY raw score data (S4, independently verified by live fetch 2026-07-30), and SOURCES.md confounder 1 records that it is re-anchored as benchmarks saturate, so historical values move retroactively and any cached figure silently drifts (S4, independently verified). Credit where due: S4 is the reference implementation of this method in the wild, its fitting code is inspectable, and any tracker adopting Method C should fit with that code rather than reinvent it. What a tracker must do about the drift, as a hard requirement: freeze and archive each vintage of the fit as an immutable dated artefact; never overwrite a published historical value; publish the revision as a first-class output, showing the prior vintage, the current vintage, and the per-period difference; and never quote a cached figure without its vintage identifier. A tracker that ingests a re-anchored index without vintage pinning will publish silently different history at every release and will not know it.

  1. Correlated indicators. This is what the method is for. Correlation is modelled rather than

penalised: shared movement loads on theta, indicator-specific movement lands in psi_i, and no double-counting occurs because nothing is summed. It is the only one of the four methods that addresses correlation from first principles.

  1. Discontinuous jumps. Absorbed into the fit, which is a mixed blessing. A single model's jump

raises its theta and slightly re-estimates item difficulties, so the jump is distributed rather than displayed. A tracker wanting jumps visible must publish the per-item residuals for the jumping model, because the aggregate will understate the discontinuity.

  1. Strengths. Correlation handled correctly. Missing data handled correctly. Uncertainty native

and honest. No arbitrary weights, no arbitrary normalisation window. A public, licensed reference implementation exists (S4, independently verified), so the method is not speculative. It extracts more signal per observation than any of the alternatives.

  1. Weaknesses. Retroactive drift, per field 7, which is close to disqualifying for a public

tracker whose credibility depends on a reader finding the same number they read last quarter. The deeper objection is the one-factor assumption: the method presumes the correlated observations reflect ONE coherent variable, and the inherited jagged-frontier material (sBs/AGI-handbook/chapters/03-capability-map.md) disputes exactly that, holding that capability is uneven across task families and that regression on a 30-day recheck is observed. If capability is jagged, theta is a weighted average of unlike things wearing the clothing of a measured quantity. Tests that would detect single-factor failure, all required at every release: (a) residual structure, tested as the correlation matrix of eps_{m,i} after extracting one factor, with any off-diagonal block above 0.30 (specification parameter) reported as evidence against unidimensionality; (b) a two-factor fit run in parallel, with the second factor's loadings published, and a second factor explaining more than 15 percent of common variance (specification parameter) treated as failure of the one-factor form; (c) dimension-specific divergence, tested as the sign agreement between theta movement and each dimension's own indicator movement, with any dimension diverging in sign for two consecutive releases reported as a finding. If any of the three fires, the single-factor headline is suspended, not adjusted. Further weaknesses: it is the most expensive method to maintain, the least legible to a lay reader, and the hardest for a hostile reader to check without re-running the fit.

  1. Manipulation susceptibility. By the maintainer, through item-set inclusion (which benchmarks

enter the fit is a larger lever than any weight in Method A, and it is far less visible), anchor choice, and prior specification. By frontier developers, through saturating the items that currently carry the most discrimination, which raises theta and simultaneously reduces the model's ability to resolve the next increment. By the upstream index operator, if the tracker ingests a fitted index rather than fitting from raw scores: a re-anchoring decision made upstream changes the tracker's published history without the tracker acting. Section 4 already forbids ingesting S4's fitted values for exactly this reason and restricts use to its raw score CSV.

  1. Subjective judgment. Degree: HIGH, and the least visible of the four, because the choices are

technical and buried. Enumerated: which indicators and which items enter the fit; binarisation rules for continuous scores under the IRT form; the functional form (factor-analytic versus IRT, one factor versus more); prior distributions on loadings, difficulties, discriminations and the guess floor; the identification anchor; the estimation sample (which models, which periods); the treatment of scaffold variants of one model as one item response or several; the residual-structure and second-factor thresholds in field 11; the vintage archiving cadence; the credible-interval width; whether the frontier series is a period maximum or a running maximum.

  1. Audit exposure. The fitting code is committed and the fit is reproducible from archived raw

snapshots by a single command; if the S4 code is used, its version is pinned. The item set, priors, anchor and estimation sample live in one versioned configuration file. Every vintage is archived immutably with a vintage identifier, and every release publishes the revision table against the previous vintage. The three unidimensionality diagnostics are published as standing outputs, not on request, so a reader can see the model's own evidence against itself. Posterior draws are downloadable so a third party can recompute any summary.

  1. Public legibility. A reader treats theta as an IQ-like measured quantity, because it is

scaled like one and has no natural units to contradict them. They will compare a figure across releases without noticing the vintage change, and the drift will look like real movement in whichever direction the refit went. They will read a widening interval as instability in AI capability rather than thinness of coverage. And they will not know that the number presumes capability is one thing.


  1. Formulation. Publish two numbers plus a flag, never one, and publish them only inside one

inseparable citation string (field 15).

POSITION, P_t. The count of pre-registered thresholds met out of the reported denominator, each threshold stated and evaluated in its indicator's native unit. This field is the authoritative definition of POSITION for the whole specification. Any other rendering of position anywhere in this document, including a reduced set, a partial-coverage view, an MVP subset or a single-panel summary, must be a count of pre-registered thresholds met over a stated denominator in native units, and must never be a normalised level, an index, a basket median, or an average of distances. Where another section renders position differently, this field governs and the other rendering is a specification error to be corrected, not an alternative definition.

Threshold state algebra. Every registered threshold sits in exactly one of five states at each release. Each state has a fixed, published effect on the numerator of P_t, on the denominator actually reported, and on the flag. The earlier draft of this specification defined three states and left the effect of UNKNOWN on the denominator undefined; that omission is the defect repaired here, and it is recorded as a repair rather than presented as original design.

statenumerator of P_treported denominatorflag effect
MET+1+1none
UNMET+0+1none by itself; produces a floor component of zero if it is the only in-scope threshold in its dimension
UNKNOWNexcluded, contributes nothingexcluded, contributes nothingraises co-occurring condition UNKNOWN-COVERAGE with the count; if the cause is the 90-day staleness flag it also raises STALE naming the dimension and the source ID
AMBIGUOUSexcluded from the point count; counted out of P_low and into P_high+1raises co-occurring condition AMBIGUOUS naming the threshold
PROVISIONALexcluded from the point count and from P_low; counted into P_high+1raises co-occurring condition PROVISIONAL naming the threshold and the days since first crossing

The point count is the MET count alone, so P_t = P_low by construction and P_high = MET + AMBIGUOUS + PROVISIONAL. The asymmetry is deliberate: an unresolved threshold can widen the range upward but can never move the number that gets quoted.

UNKNOWN is excluded from both numerator and denominator because the two available alternatives are both wrong. Counting UNKNOWN into the denominator turns a measurement outage into a falling capability figure. Dropping the threshold silently raises the ratio because coverage fell. Excluding it from both and publishing the excluded count separately is the only treatment that neither flatters nor libels, and it is why the count of dimensions actually measured has to appear in the headline string itself.

Weakest-link floor, with two scope gates that a dimension must pass before it enters the floor at all.

Gate 1, observation scope. A core dimension enters the floor only if it holds at least one threshold in a state other than UNKNOWN. A dimension all of whose thresholds are UNKNOWN has no defined ratio, is declared OUT OF SCOPE: MEASUREMENT for that release, and contributes no zero and no value. This is the direct fix for the failure the floor had as first drafted: a single quiet source could drive one dimension's met count to zero over a non-zero denominator, set the floor to zero, and fire headline text asserting that a capability had not moved when what had happened is that nobody published.

Gate 2, dispersion scope. A threshold enters the floor only if its dimension has a published dispersion estimate: a stated run-to-run, scaffold, grading or inter-rater dispersion for the quantity the threshold is set on, carrying its own source ID and verification class. A dimension without one is declared OUT OF SCOPE: DISPERSION and is excluded from the floor until an estimate exists. This is a general rule, stated once here and applied uniformly to every dimension, and it replaces every per-indicator carve-out in earlier drafts, including the D2-specific patch, which was ad hoc and therefore unauditable. The reason is structural: a minimum over channels is a statistic of the noisiest channel, so a dimension whose noise is unquantified would set the floor by an unbounded quantity, and in this roster the least instrumented dimensions are the ones resting on the least funded sources. That last clause is INFERRED, not verified: SOURCES.md confounder 5 establishes source mortality and confounder 8 establishes uneven scrutiny of one supplier, and neither establishes a relationship between a dimension's instrumentation and its supplier's funding. The gate stands on the dispersion argument alone, which needs no such claim.

With D_scope the set of core dimensions passing both gates and k = |D_scope|,

floor_t = min over d in D_scope of (MET thresholds in d / (MET + UNMET + AMBIGUOUS + PROVISIONAL thresholds in d))

If k < 2 the floor is suppressed and not published as a number, because a minimum over one dimension is not a weakest-link statement about breadth; the string then carries FLOOR SUPPRESSED with k and the reason. k is printed inside the headline string itself, in the same character run as the floor value, and never in a separate coverage badge below the headline, because a badge is separable and separable qualifiers are dropped on quotation.

Headline language, fixed and not at the writer's discretion. Where floor_t = 0 for a dimension that is IN SCOPE, the permitted sentence names the dimension and reads: "no threshold in dimension d is met; n of n registered thresholds in d are observed and unmet." Where a dimension is OUT OF SCOPE, the only permitted sentence is: "dimension d is not measured this release (measurement gap, source S, n days since last publication)". The phrases "has not moved", "no progress", "flat" and every synonym are reserved exclusively for observed and unmet thresholds and are forbidden for out-of-scope dimensions, by lint rule in the release script rather than by editorial care. An out-of-scope dimension is a statement about the instrument. It is never rendered as a statement about the systems.

RATE, R_t: a doubling-time estimate denominated in the horizon indicators, taken from C3, which is the ordinary least squares regression of log2(C2) on time over a trailing 24-month window with a six-observation minimum, published with its regression interval, its observation count and its residuals, and with the same regression on C1 reported beside it. Where the two thresholds imply different rates, both are published and the divergence is the finding. No fitted line is extended beyond the last observation.

FLAG, F_t: not a single value but a primary state plus a set of co-occurring conditions, emitted as one token. The primary state takes exactly one value per release from {CLEAN, REGRESSION, SATURATION, INSTRUMENT-FAILURE, STALE}. REGRESSION when D4's count exceeds its trailing-four-period median against a stable denominator. SATURATION when A2's median half-life falls below 12 months (specification parameter) or when more than one third of A1's basket has headroom below 0.10. INSTRUMENT-FAILURE when A2 is short AND P_t is static, the combination that indicts the ruler rather than the systems. STALE when any core-dimension source has passed the 90-day staleness flag. Precedence when several conditions hold, in order: INSTRUMENT-FAILURE, STALE, REGRESSION, SATURATION, CLEAN.

The condition set is separate from the primary state and is never suppressed by it. Conditions are drawn from {STALE(dimension:source,days), UNKNOWN-COVERAGE(count), AMBIGUOUS(ids), PROVISIONAL(ids,days), DISPERSION-GAP(dimensions), FLOOR-SUPPRESSED(k), SPEND-BOUGHT(ids), UNVERIFIED-ELASTICITY(ids)}, the last two raised by G4i's gate in section 4. A crossing tagged SPEND-BOUGHT is held in PROVISIONAL and counts zero until a configuration at or below the reference spend also crosses, at which point it re-dates to that observation; a crossing tagged UNVERIFIED-ELASTICITY counts, because absence of an elasticity estimate is not evidence of purchase. Both tags were defined in section 4 before this vocabulary existed, and the omission was a reconciliation failure between two concurrent revisions rather than a design choice. Every condition that holds is emitted, and precedence selects only which name leads the token. As first drafted, precedence placed INSTRUMENT-FAILURE above STALE and the flag carried one value, so an instrument-failure release could hide the staleness that had caused the floor to move, which is the second defect repaired here. A release whose primary state is INSTRUMENT-FAILURE and whose C dimension is 124 days stale emits INSTRUMENT-FAILURE+STALE(C:S5,124d), not INSTRUMENT-FAILURE; the day count in that token is an illustrative format value, not an observation.

  1. Required inputs. Position: thresholds on A1, C2, D1, E1, and on D2 and D3 when funded, gated by

B1 and B2 per Method B field 2. Rate: C3, which consumes C2 and C1. Flag: A2 and D4, plus the staleness state of every core-dimension source. Denominators and context published beside but never inside the headline: G1i, G2i, G3i. Tracked separately and never summed: E2, F1, H1, H2, I1. The recommended MVP position set is small, four to six thresholds, so that each is individually defensible. Two further inputs are required by the repairs in field 1 and are not optional. First, a dispersion estimate per core dimension, each with its source ID and verification class, without which that dimension does not enter the floor; at specification time several dimensions have none, so the floor starts narrow and says so. Second, a reproduction record per threshold crossing: the identity of the first runner, the identity of any second independent runner, the configuration release state, and the dates, which is what moves a crossing out of PROVISIONAL. Neither input is a benchmark score and neither enters the count.

  1. Normalisation. Position performs none, by inheritance from Method B: thresholds are stated in

native units. Rate performs exactly one transform, log2, applied to a quantity measured in minutes, which is a unit-preserving reparameterisation rather than a normalisation choice, and the only judgment in it is the window length and the minimum observation count. This is the smallest normalisation surface of the four methods and is the main reason the method survives sensitivity testing better than A. One consequence has to be stated because it constrains the headline choice argued after field 15. The per-threshold distance d_{i,t} = (tau_i - x_{i,t}) / (tau_i - x_{i,0}) is normalised twice over, by tau_i and by the pre-registered origin x_{i,0}, so it is a presentation device and never an input to the count. Any average, median or weighted combination of distances across thresholds is forbidden as a published figure: a mean of distances is Method A with equal weights and a second unfalsifiable constant hidden in each denominator. The distance vector is published element by element, in native units alongside the normalised form, and never collapsed.

  1. Missing data. Position reverts affected thresholds to UNKNOWN and publishes the unknown count,

so missingness degrades the headline visibly. UNKNOWN leaves both the numerator and the reported denominator per the field 1 table, and a dimension with no non-UNKNOWN threshold leaves the floor under gate 1, so a quiet source narrows the floor's scope and lowers k instead of producing a zero component. Rate publishes no value below the six-observation minimum inside the window; it publishes the observation count and the reason instead. Flag carries STALE, either as the primary state or as a co-occurring condition that a higher-precedence primary cannot suppress, naming the dimension, the source ID and the day count. The design property that matters: in this method a dead source cannot produce a flattering headline, and after the field 1 repair it cannot produce a damning one either. It produces a visibly narrower one.

PROVISIONAL, the fifth state, exists because staleness and falling back below tau_i do not cover the case that matters most at four to six thresholds, where one crossing is between 16 and 25 percent of the position (arithmetic on the specification parameter |T|, not a claim about the world). That case is a crossing that was observed and reported, whose configuration was never released, and which no second party has reproduced. A crossing enters PROVISIONAL at first observation and leaves it only on reproduction by a second independent runner on a configuration a third party can obtain, recorded with both runner identities and both dates.

The choice, stated definitely rather than left open: a PROVISIONAL crossing counts zero, not half. Half weight was rejected for two reasons. It destroys the one property that makes the count worth publishing, that the headline is an integer over a fixed denominator and therefore not a level in disguise; a position of 2.5 of 5 is a level, and readers will treat it as one. And it splits the difference on a question that is not symmetric: the harm from publishing an unreproduced crossing as capability is a false claim about the world, while the harm from withholding it is a lag, and the lag is itself published (field 5). The cost is that position trails confirmed capability by the confirmation interval, and this is disclosed in the headline block, not buried: the PROVISIONAL condition on the flag names the affected thresholds and their pending days, and P_high includes them so the range shows what position would be if every pending crossing confirmed.

  1. Uncertainty. Three separate statements, deliberately not combined. Position: the range

[P_low, P_high] from interval-versus-threshold comparison, plus the unknown count. Rate: the regression interval on 1 / slope, computed as the ordinary least squares slope interval propagated through the reciprocal, widened by a bootstrap over the observation set (2,000 resamples, a specification parameter) to reflect the small sample, and reported as an interval on months per doubling with the observation count and the residual series. Because S5 revisions move historical points, the window is recomputed from scratch each release and the observations entering and leaving are named. Flag: no interval, being categorical, but each firing condition publishes the quantity that triggered it, and each co-occurring condition publishes its own quantity as well, so an INSTRUMENT-FAILURE release still prints the staleness day count that accompanied it. Two additions carry the field 1 repairs into the uncertainty statement. The floor publishes k, the number of core dimensions that actually entered it, together with the name and the exclusion reason of every core dimension that did not, split by gate: OUT OF SCOPE: MEASUREMENT for gate 1, OUT OF SCOPE: DISPERSION for gate 2. A floor without k is not a floor, because the same value means different things over four dimensions and over two. And time-to-confirmation is published as its own series, not as a footnote to the position: per threshold, the days from first crossing observation to reproduction by a second independent runner, with never-confirmed crossings published as right-censored observations rather than dropped, and the median over confirmed crossings reported with its count. This series is not an uncertainty statement about the tracker. It is a measurement of the field: it quantifies how long a frontier capability claim stands unreproduced, and at specification time no source publishes such a quantity (class missing), so publishing it would be a finding in its own right and should be presented as one rather than as instrument plumbing. It is also the honest price tag on the choice in field 4, since the lag it measures is exactly the lag that choice imposes. Confidence grades attach per component under the house convention, decaying with horizon: position at MEDIUM for 12 months, rate at LOW for 12 months given the six-observation minimum on a revisable series per C3's own grade, flag at MEDIUM. No single composite confidence grade is published, because there is no single number to grade.

  1. Backward movement. The most sensitive of the four. Position falls when a threshold reverts.

Floor falls to zero the moment an in-scope core dimension loses its met thresholds while its thresholds remain observed, and the headline text changes with it; a dimension that loses observation rather than capability leaves the floor's scope instead, which is a different published event with different mandated wording. Distinguishing those two cases is the whole content of the field 1 repair, and the distinction is the difference between a regression finding and a coverage finding. Rate lengthens when recent frontier observations fall below trend. Flag carries REGRESSION on D4's evidence directly. Three independent routes for a regression to reach the headline, none of which averaging can suppress, and one route by which a regression cannot reach it falsely.

  1. Benchmark saturation. Made explicit rather than absorbed. SATURATION is a published headline

state, not a silent property of the arithmetic. The A1 basket rotation and chain-link splice handle threshold migration, and INSTRUMENT-FAILURE exists precisely for the case where the instruments have stopped discriminating while the position count is static, which in Methods A and C is indistinguishable from a plateau.

  1. Correlated indicators. Handled by construction rather than by correction. Position counts at

most one threshold per latent construct per dimension, so correlated pairs cannot inflate it. Rate is denominated in one dimension only, so it makes no cross-dimension commensuration and no correlation assumption at all. Flag is categorical, so correlation does not arise. The published correlation matrix is retained as a diagnostic, and any pair above 0.85 (specification parameter) that both carry thresholds is a specification error to be fixed, not a weight to be adjusted.

  1. Discontinuous jumps. Position moves by one integer at most per threshold and logs the flip with

its causing observation, and a jump arriving as a single unreproduced crossing lands in PROVISIONAL first, so the most jump-prone events reach P_high and the flag before they reach the quoted number. Rate is protected by the trailing window and by published residuals: a jump appears first as a residual outlier, which is the honest presentation, and only later as a changed slope. Any release in which one observation entering the window changes the doubling-time point estimate by more than 25 percent (specification parameter) publishes the leave-that-observation-out estimate beside it.

  1. Strengths. It answers the two distinct questions a progress tracker is asked, position and

rate, without pretending they are one quantity. It cannot suppress a narrow regression, because of the floor. It cannot hide stagnation or instrument failure, because those are published states. Missing data degrades it visibly and, after the field 1 repair, degrades it as a coverage statement rather than as a false capability statement. The qualifiers are mechanically inseparable from the numbers rather than editorially attached to them, per field 15, which is the only mitigation for quotation loss that does not depend on the reader. It has almost no normalisation surface and no weight vector, so the two largest unfalsifiable levers in Method A are absent. Its subjectivity is concentrated in one pre-registered, timestamped artefact. It is cheap for a single maintainer: a threshold file, one regression, and three flag conditions. And it retains the one genuinely useful property of the Kurzweil form. Retention of the Kurzweil log chart as the RATE view only: a log axis renders doubling behaviour as slope, so a reader sees a rate change as a bend rather than having to infer it from adjacent numbers, and a long-run log chart can carry successive pinned metric versions on one axis and so survives the saturation of any single metric. It is confined to the rate view because, per section 1.6, the log frame smuggles in a continuity assumption the data cannot supply, and because the metrics on such charts are historically selected for having been exponential, which is survivorship in the metric choice rather than a finding. The chart therefore shows observations over pinned metric versions with no fitted line extended past the last observation.

  1. Weaknesses. Three numbers are harder to communicate than one, and the press will quote

whichever is most dramatic, usually the rate. The position count is coarse and carries Method B's threshold-setting burden in full. The rate rests on C3, which rests on C2, which rests on S5, a single source with irregular cadence whose updates were observed only in March to May 2026 and whose operator self-describes limited capacity (S5, independently verified by live fetch 2026-07-30); if S5 stops, the rate half of the headline goes STALE and there is no verified substitute. The flag's thresholds are specification parameters and a reader could reasonably contest each. The method also declines to give a single answer to "how far along are we", which some readers will read as evasion. Three weaknesses are created by the field 1 repairs and are stated as costs rather than absorbed. The floor now covers fewer dimensions than there are core dimensions, and at specification time it may cover as few as one or two, since a dimension without a published dispersion estimate is excluded until one exists; k therefore appears in every citation string and a reader is entitled to say the floor is thin. That cost is the correct one to pay: a floor computed over a dimension whose noise is unquantified is a statistic of that dimension's funding rather than of its capability, and reporting a number of unknown provenance is worse than reporting a narrower number of known provenance. Where k < 2 the floor is suppressed entirely, which means the instrument's weakest-link protection can be unavailable exactly when coverage is worst; that is disclosed rather than patched, and it is why K10 exists in section 7. And the PROVISIONAL state makes position a lagging measure of confirmed capability by design, so a real crossing can sit outside the count for as long as the field takes to reproduce it, which is a quantity nobody currently publishes.

  1. Manipulation susceptibility. By the maintainer, through post-hoc threshold setting, through

choosing the rate window length after seeing the residuals, and through setting the flag conditions to avoid firing. The registration discipline and the version-bump rule are the whole defence, and they work only because the commit timestamps are public. By frontier developers, through optimising to a published threshold, and through selective non-submission, which suppresses D4's count and can hold the flag at CLEAN; the published denominator and per-developer submission counts make suppression visible. By benchmark operators, through version scheduling that changes what a threshold means. By the horizon source operator, S5, whose refit choices move both C2 and C3 and therefore the entire rate half of the headline, without the tracker acting. The repairs open two new attack surfaces and both are named rather than left to be discovered. Gate 2 makes non-publication of a dispersion estimate a way to remove a dimension from the floor, which is available to any party that would rather not be the weakest link; the defence is that the exclusion is published with the dimension named and the missing estimate named, so the omission is visible and attributable, and the dispersion register carries the date the estimate was requested. PROVISIONAL makes announcement without release a way to hold a crossing out of the count, which cuts both ways: a developer wanting the crossing counted must release a reproducible configuration, and a developer wanting to appear ahead of the count can point at an unreproduced announcement the tracker refuses to count. The time-to-confirmation series is the defence, because it makes the pattern of unreproduced announcements a published series with the announcing party named.

  1. Subjective judgment. Degree: MEDIUM to HIGH, lower than the other three and, more importantly,

concentrated. Enumerated: the threshold set and every tau_i in it; the denominator |T|; the one-threshold-per-construct assignment; the B-dimension contamination gate; the definition of the weakest-link floor and which dimensions are core to it; the rate window length of 24 months; the six-observation minimum; the log-linear functional form for the rate; the bootstrap resample count; the five flag states, their firing conditions and their numeric triggers; the precedence order among simultaneous flag conditions; the staleness cut-off of 90 days; the jump disclosure trigger of 25 percent; the decision to publish three components rather than combine them. Added by the field 1 repairs, and each as contestable as the items above: the five-state threshold algebra and specifically the decision to exclude UNKNOWN from the denominator rather than from the numerator alone; the two floor scope gates and their ordering; what counts as a published dispersion estimate and which statistic satisfies gate 2; the k >= 2 suppression rule; the decision that a PROVISIONAL crossing counts zero rather than half; what counts as a second independent runner and as an obtainable configuration; the censoring rule on the time-to-confirmation series; the condition vocabulary on the flag and the rule that conditions are never suppressed by the primary state; the exact canonical citation string format; and the refusal rule in field 15, which is a publication policy choice and not a measurement one.

  1. Audit exposure. One versioned threshold file, one versioned flag-condition file, one versioned

rate-configuration file, all semantically versioned with change-log entries naming date, reason and effective release. Threshold commits predate the observations they judge and the timestamp is the evidence. Every threshold flip, every flag transition and every observation entering or leaving the rate window is logged with its source ID and verification class. Residuals published each release. Parallel old and new counts published for four releases after any threshold-set version bump. Raw archived snapshots plus a single reproduction command. Four audit objects are added by the repairs. A dispersion register: one row per core dimension giving the dispersion statistic, its source ID, its verification class, and where absent the date it was last sought, so gate 2 exclusions are checkable and not merely asserted. A reproduction register: one row per crossing giving first runner, second runner, configuration availability, both dates and the resulting state, which is the audit object behind PROVISIONAL and behind the time-to-confirmation series. A floor scope log: per release, k, the dimensions in scope, and the gate that excluded each dimension out of scope. And the canonical citation string itself, emitted by the release script as a build artefact and archived per release, so that any string found in the wild can be matched against the archive and a doctored or truncated quotation identified. A release that cannot emit the string does not publish.

  1. Public legibility. A reader reads the position count as a fraction of the way to AGI, the same

misreading Method B invites, and the authoritative-looking denominator makes it worse. They read the rate as a forecast and will extend it mentally past the last observation even though no fitted line is drawn. They will quote the doubling time without its interval and without its observation count. They will ignore the flag, which is the component doing the most work. Required reader-facing text, in the headline block itself and not in a footnote: the denominator is a specification choice and not a count of the requirements for AGI; the rate describes observed history and projects nothing; the flag governs how the other two may be read.

Naming that failure mode is not a mitigation, and an earlier draft of this field mistook one for the other. Every headline will collapse to whichever number is shortest in every citation that is not mechanically prevented from collapsing, so the mitigation has to be mechanical. It is specified as three rules.

Rule one, the canonical atomic citation string. Exactly one string is the publishable form of the headline, in one character run with no separable fields:

AGI-PT <spec_version>/<threshold_set_version> | POSITION <P_t> of <denominator> [<P_low>..<P_high>] | FLOOR <floor_t> in <dimension> over <k> of <core_count> core dimensions | RATE <months> mo per doubling [<lo>..<hi>] n=<obs> | FLAG <PRIMARY>[+<condition>...] | VINTAGE <YYYY-MM-DD>

Where the floor is suppressed, the FLOOR segment reads FLOOR SUPPRESSED k=<k> plus the reason. Where the rate is not published, the RATE segment reads RATE NOT PUBLISHED n=<obs> of 6 required. Where a core dimension is out of scope, the FLOOR segment appends ; <dimension> NOT MEASURED (<gate>, <source>, <n>d). No segment is ever omitted, and no segment is ever rendered alone.

Rule two, single-serving. The public API returns this string as its only headline field. It does not return position, floor, rate or flag as separate scalars, and it does not return a machine-readable object from which a scalar could be lifted without the flag, because any such field is the field that will be quoted. The social-preview image is generated per release with the same string burned into the image pixels as its only text, so a link preview cannot render a number without its flag. Downloadable data remains complete and separable, per the audit rules in field 14: the restriction is on the headline surface, not on the archive, and a researcher can still compute anything from the snapshots.

Rule three, refusal. Where the flag's primary state is not CLEAN, no rendering path emits a bare position. The release script and the API both refuse: the string is served, or nothing is. This is a genuine refusal and not a fallback to a shorter form, and its cost is accepted openly. A quotation-hostile consumer will sometimes get no number rather than a number it can strip, and that is the intended behaviour, because a stripped number in an instrument-failure release is worse than silence.

Worked example, CLEAN state (illustrative format example; the values are placeholders and not observations):

AGI-PT 0.1.0/1.0.0 | POSITION 2 of 5 [2..3] | FLOOR 0.33 in D over 3 of 4 core dimensions | RATE 7.4 mo per doubling [4.1..19.6] n=8 | FLAG CLEAN | VINTAGE 2026-07-30

Worked example, REGRESSION state with a co-occurring staleness condition and one dimension out of scope (illustrative format example; the values are placeholders and not observations):

AGI-PT 0.1.0/1.1.0 | POSITION 1 of 5 [1..3] | FLOOR 0.00 in D over 2 of 4 core dimensions; B NOT MEASURED (gate 1, S3, 142d); C NOT MEASURED (gate 2, dispersion absent) | RATE NOT PUBLISHED n=4 of 6 required | FLAG REGRESSION+STALE(C:S5,124d)+PROVISIONAL(A1-t2,61d)+UNKNOWN-COVERAGE(2) | VINTAGE 2026-10-31

Read the second string as it is written: one dimension is at zero on observed and unmet thresholds, two dimensions are not measured at all, the rate is withheld for want of observations, and one crossing is pending reproduction. None of that is a statement that AI capability has or has not advanced, and the string is built so that no substring of it can be quoted as one.

Assessment and rejection of the clock as the primary form. The Doomsday Clock is rejected as the primary headline for the reason given in section 1.6: it presents expert judgment in the visual grammar of measurement, importing precision the process does not have; minutes-to-midnight implies a metric distance to an event when no such distance is defined, and this specification asserts no agreed definition of AGI to measure distance to; and no reader can recompute the position from published inputs, so the clock cannot move for a reason a reader can audit. Its one transferable property, a fixed revision cadence with an attributable owner, is kept in the governance layer.

Assessment of the proposal to invert position and distance, and the decision. The charge is that the instrument already computes d_{i,t}, a pre-registered, origin-anchored, native-unit distance that is continuous and falsifiable, and that publishing the integer count as the headline while demoting the distances to detail throws away resolution and is a legibility choice wearing the costume of an honesty choice. The charge continues that a threshold tau_i announces itself as a fact and is exactly as much a preference as a weight is. The decision is to keep the count and to reject the inversion, and the concession comes first because it is conceded entirely: tau_i is a preference. It is chosen by one maintainer, it is contestable line by line, and nothing in the data can show that it should have been set elsewhere. Any reading of this document that takes the threshold set as more objective than a weight vector is a misreading of it.

The count is nevertheless retained, for three reasons that engage the charge rather than deflect it. First, falsifiability as applied, which is a different property from being preference-free. A weight never produces a checkable event: no future observation can vindicate or refute 0.25. A pre-registered threshold produces a dated binary event that a third party can check years later against the archived observation, without access to the maintainer's parameter file and without agreeing that tau_i was well chosen. The preference is in the threshold; the record is in the crossing. Second, d_{i,t} does not remove tau_i, it divides by it, and it adds a second chosen constant, the pre-registered origin x_{i,0}, in the same denominator. A distance headline would therefore carry two unfalsifiable constants where the count carries one, and its level would move whenever either was re-registered, which is the retroactive drift objection that disqualified Method C. Third, a continuous distance headline invites the aggregation that this method exists to avoid: the only way to make a vector of distances into a headline is to average it, and an average of distances is Method A with equal weights and a hidden origin per term. The instrument would arrive back at the weighted composite by a longer road.

What the charge wins, and it is not a small concession. The resolution objection is correct: an integer over four to six thresholds says nothing between crossings, and a reader is entitled to see the movement the count cannot show. So the distance vector is published in full at every release, element by element, in native units with the normalised form beside it and an interval on each element, in the headline block and in the API record as a non-substitutable array rather than as an appendix or an on-request extra. It is not collapsed to a scalar, per field 3, and no average of it is published. The count is the claim; the distance vector is the evidence of movement between claims; and a reader who thinks the count is too coarse has everything needed to say so from the published record.


Recommendation

Method D. The argument, not the assertion.

Against A. A's weight vector is unfalsifiable, so the headline encodes the maintainer's priors and reports them in the grammar of measurement; that is the specific failure the brief names. Its normalisation choice can move the series more than the data does, its correlated terms double-count one latent movement, its saturated terms go inert invisibly, and its averaging suppresses narrow regression, which is the movement most worth catching. What is lost by not choosing A: a single quotable number, the lowest possible explanation cost, and a form the press already knows how to handle. That loss is real and D pays it deliberately. Mitigation: A is retained as a published reference computation with at least three alternative weighting presets, so a reader who wants a scalar gets one that is clearly labelled as a sensitivity artefact rather than the headline.

Against B. B is D's position component, and D adopts it wholesale. What is lost by choosing B alone is the rate: a threshold count says nothing between crossings and cannot answer whether movement is accelerating, decelerating or stopped, which is the question a progress tracker mainly exists to answer. B alone also has no mechanism for reporting instrument failure. D is B plus the two things B lacks, and it also repairs B's dimension-robustness statistic, which as written in Method B field 1 is undefined under missing data and dominated by the least instrumented dimension.

Against C. C is the methodologically strongest treatment of correlation, missing data and uncertainty, and a licensed reference implementation exists (S4, independently verified by live fetch 2026-07-30). Two reasons it is not the headline. First, retroactive drift: a latent fit re-anchors as benchmarks saturate, so published history moves (SOURCES.md confounder 1, independently verified), and a public tracker whose last-quarter figure silently changes has lost the credibility it exists to build, even with vintage archiving in place. Second, the single-factor assumption is exactly what the inherited jagged-frontier material disputes, so the headline would presume the thing most in question. What is lost by not choosing C: correct handling of correlation and missing data, native uncertainty, and the extra signal a latent fit extracts from thin observations. Mitigation: C is run as a standing diagnostic, fitted from raw scores with pinned code and archived vintages, and its three unidimensionality tests are published. Its disagreement with D is a finding, and if it disagrees persistently the threshold set is suspect.

Conditions under which I would switch. To C, if a genuinely contamination-resistant, unsaturated basket became continuously available with machine-readable history, because the drift problem is caused by saturation-driven re-anchoring and a basket that does not saturate largely removes it; the switch would additionally require that the three unidimensionality diagnostics pass for four consecutive releases. To A, only if the eligible indicator set grew large enough and independent enough that correlation collapse left six or more genuinely distinct terms, and only with both normalisations published in parallel. To B alone, if S5 died with no substitute horizon fit, since the rate component would then be unsupportable and publishing a stale rate is worse than publishing none. Away from D entirely, if the pre-registration discipline could not be sustained, because unregistered thresholds make D's position component indistinguishable from a narrative device.


Where subjectivity enters

Every point of human judgment across the whole system. In each case: the decision, who makes it, and the mechanism that exposes it. The maintainer is a single person (BRIEF, WHO), so "who" is the maintainer almost everywhere, and that concentration is itself the risk the mechanisms address.

Indicator selection. Which of the twenty-one roster indicators exist at all, which dimensions they are assigned to, and which are core, supporting or contextual. Maintainer, at specification time, constrained by the director-fixed roster. Exposed by: the section 2 selection principles stated so as to be capable of rejecting a candidate, section 5's register of explicitly rejected measures with reasons and alternative uses, and a versioned roster file whose change log records every addition and removal with its date.

Threshold definition. Each tau_i, its direction, the pre-registered origin, the denominator |T|, and which of a correlated pair carries the threshold. Maintainer, before examining the data. Exposed by: a signed timestamped commit of the full threshold set with rationale, committed before the observations it judges; the additive-version-bump rule; and parallel publication of old and new counts for four releases after any change.

Weighting. The dimension budget and within-dimension split in the retained Method A reference computation, and the capped weight on E1 imposed by its licence-constrained source. Maintainer. Exposed by: a versioned weights file committed ahead of the data, at least three published alternative presets, and the leave-one-out and weight-perturbation tables in section 7.

Normalisation. Family (min-max versus z-score), reference window endpoints, squash function, clipping bound, and the log2 transform and window length on the rate. Maintainer. Exposed by: publishing the reference computation under both normalisation families at every release, treating direction disagreement between them as a non-result, and publishing the rate under two window lengths.

Model inclusion rules. Which systems count as frontier, how scaffold variants of one model are treated (one observation or several), how a renamed or rebranded lineage is ordered, and whether the frontier series is a period maximum or a running maximum. Maintainer, using S2 release ordering whose dates are estimated for some entries (S2, estimated). Exposed by: a published inclusion rule document, per-observation records of scaffold and harness identifiers, publication of the highest and lowest scaffold result with the spread rather than the best, and a named list of excluded systems with the rule that excluded them.

Benchmark inclusion rules. Which benchmarks enter the A1 basket, the admission and retirement criteria, the ranked admission queue, the splice procedure on rotation, and which item sets enter the Method C diagnostic fit. Maintainer, against criteria fixed in A1's operational definition. Exposed by: the published admission queue ranked on headroom then auditability class, so retirements cannot be filled with a convenient replacement; publication of both baskets and the splice factor for the overlapping observation; retention of the unspliced series; and a versioned item-set file for the diagnostic fit.

Data-quality judgments. Whether an observation is admissible: verification class assignment, exclusion of unverified community submissions, the single-attempt exclusion tally in D1, refusal handling in D3, retry policy on transient errors in D2, the materiality threshold that distinguishes a D4 regression from grading noise, and separation of upstream revisions from genuine regressions. Maintainer, per observation. Exposed by: a mandatory verification class on every stored observation; published exclusion tallies rather than silent drops; the revision count published separately from the regression count; and downloadable raw snapshots so a third party can re-apply different admissibility rules.

Interpretation of missing evidence. Whether a gap means not-yet-measured or not-happening. This is the judgment most likely to be made invisibly. Specific decisions: reverting a met threshold to UNKNOWN rather than leaving it met; declining to publish a rate below the six-observation minimum; publishing D2 and D3 as DECLARED GAP G1 rather than substituting a proxy. Note the round-2 amendment: D2 is now built as a deliberately small costed probe in section 9 and carries a value from stage 1, so only D3 remains a gap with no value. Refusing to state a number for the benchmark-to-reality gap because none is published (G3, established by search 2026-07-30, class missing); and treating an absent submission as absent evidence rather than absent capability. Maintainer. Exposed by: the five-state MET, UNMET, UNKNOWN, AMBIGUOUS, PROVISIONAL reporting of Method D field 1, with published unknown, ambiguous and provisional counts, and with UNKNOWN excluded from both numerator and denominator so that a measurement outage cannot be rendered as a capability statement; the floor scope log naming the gate that excluded each out-of-scope dimension; declared gaps carrying their establishment date and gap ID in place of a value; the 90-day staleness flag on every observation; and a standing rule, checkable by a cold reader, that no proxy substitutes for a declared gap. Two judgments in this class are new with the Method D repairs and belong here explicitly: the decision that a crossing nobody has reproduced counts zero rather than half, which trades a publication lag for a false-claim risk and is argued in Method D field 4; and the decision that a dimension with no published dispersion estimate leaves the floor rather than entering it, which trades floor coverage for floor provenance and is argued in Method D field 1. Both are choices about how to treat absent evidence, both are contestable, and both are exposed by publishing the excluded counts and the excluding reason every release.


Audit mechanisms

Recommended concretely, then split honestly by what one person can sustain.

The full list. (1) A public methodology document carrying a semantic version number, cited by version in every published figure. (2) Versioned weights, thresholds and flag conditions as machine-readable files in the repository, with commit timestamps predating the data they act on. (3) Sensitivity analysis published alongside every release, per section 7, not on request. (4) At least three alternative weighting presets a reader can switch between in the interface, so the reader can see the headline's dependence on weights directly. (5) An independent advisory review with a stated cadence, reviewing the threshold set, the flag conditions and the declared gaps. (6) A published change log recording every methodology change with date, reason and effective release. (7) Downloadable raw data, including the immutable dated snapshots that make D4 possible. (8) Calculations reproducible from a single command against those snapshots. (9) Confidence intervals or ranges on every headline figure, with no point figure published alone. (10) Multiple headline figures under different assumptions, meaning the Method D headline plus the Method A reference computation plus the Method C diagnostic, all three visible. (11) The canonical atomic citation string of Method D field 15 emitted as a build artefact and archived per release, with the API and the social-preview image serving that string only, and the refusal rule enforced in the release script rather than in editorial practice. (12) The dispersion register, the reproduction register and the floor scope log of Method D field 14, published each release, so that every exclusion from the floor is attributable to a named gate and a named missing input.

What one maintainer can realistically sustain. Items 1, 2, 6, 7, 8 and 9 are cheap once and cheap thereafter, because they are file discipline and pipeline discipline rather than recurring analytical labour: versioning, change logging, snapshot retention, a reproduction script, and a house rule against bare point figures. Item 3 is sustainable if and only if it is automated as part of the release script, so that the sensitivity tables are a build artefact and a release cannot be produced without them; run manually it will be skipped within two quarters. Item 4 is sustainable because presets are a small amount of configuration and interface work done once. Items 11 and 12 are sustainable for the same reason as items 1 and 2: the citation string, the social-preview image, the two registers and the floor scope log are emitted by the release script from files the pipeline already holds, so they cost build engineering once rather than analytical labour per release. The recurring human cost in item 12 is one honest line per release when a dispersion estimate is still missing, and the register makes the omission visible whether or not the maintainer writes about it.

What is aspirational, stated as such. Item 5, independent advisory review, requires other people's unpaid time on a fixed cadence and is the item most likely never to happen; the honest minimum substitute is a public standing invitation to challenge the threshold set, plus a published register of challenges received and their disposition, which is weaker because it is reactive and self-graded. Item 10 in full is aspirational too: the Method C diagnostic requires maintaining a fit, archiving vintages and publishing three unidimensionality diagnostics every release, which is the largest recurring analytical cost in this list; at MVP the honest position is to publish Method D plus the Method A reference computation, and to mark Method C as specified but not yet running rather than to imply it exists. The recurring probe cost behind D3 is unfunded at specification time and D3 is declared a gap rather than promised. D2 is no longer in that category: its cost was calculated in section 9 and a deliberately small probe is admitted to the MVP, so D2 is a funded measurement from stage 1.


#

This is the analysis, specified as a build step. All numeric parameters below are specification parameters.

Tests that must run on every release. Automated in the release script; a release that cannot produce these tables does not publish.

  1. Leave-one-indicator-out. Recompute position, floor, rate and flag with each contributing indicator removed

in turn. Report, per removed indicator: the position count and range, the floor, the floor scope count k, the doubling-time point and interval, and the full flag token including its co-occurring conditions. The table's purpose is a single question, answered per release: does the headline movement survive the removal of every single indicator? Removing an indicator can change the floor by changing k rather than by changing any ratio, and that case is reported separately, because a floor that moves on scope is not a finding about capability.

  1. Weight perturbation. On the retained Method A reference computation, 1,000 Dirichlet draws around the

baseline dimension budget at two concentrations, one tight and one diffuse, reporting the 5th to 95th percentile of the index and of its release-over-release change, plus the share of draws in which the change reverses sign. Any sign reversal share above 10 percent is reported in the release text.

  1. Alternative normalisation. Recompute the reference index under min-max and under z-score, over two

reference windows: the full history and a trailing 36-month window. Four series. Direction disagreement between any two is reported as a non-result for that release.

  1. Alternative threshold sets. Three pre-registered alternative sets committed at the same time as the primary

set: a strict set (each tau_i harder by a pre-registered margin), a lenient set (easier by the same margin), and a maximally different set built by an adversarial pass whose brief is to produce the least flattering defensible thresholds. Publish position and floor under all four. Divergence in the direction of movement between the primary and any alternative is reported in the release text.

  1. Source-swap test. For each indicator, recompute using its fallback source in place of its primary, and

recompute with the primary declared dead and no fallback available. Given empirically high source mortality, with four of the checked sources retired, dormant or in maintenance mode and two having moved host or frozen a version line (SOURCES.md confounder 5, independently verified), the dead-source case is not hypothetical. The test's output is a per-indicator statement of what the headline becomes when that source dies, which doubles as the continuity plan. Note the honest limit: several indicators have no verified fallback (C1, C2, C3 all rest on S5, and D2 and D3 have none), so for those the swap test can only report the degradation, not an alternative value.

How to report a result that is not robust. The rule, applied without discretion. A headline movement that does not survive leave-one-out is reported as NOT ROBUST and the movement is not claimed. Operationally: if removing any single indicator reverses the sign of the release-over-release change in position, floor or rate, the release publishes the level with its range and states that the change is not robust, naming the indicator whose removal reverses it. No directional language is used in the release text for a non-robust movement, in either direction. The same rule applies to a movement that reverses under alternative normalisation, or under any of the three alternative threshold sets. A non-robust movement is still published; what is withheld is the claim, not the data.

Uncertainty taxonomy. Four kinds, kept separate because they have different remedies.

Measurement uncertainty. Error in a single observation given its source: grading noise, scaffold variance, run-to-run nondeterminism, estimated human baseline times. Remedy: intervals, repeated measurement, published verification classes. Partially quantifiable.

Source uncertainty. Error introduced by the source itself rather than the measurement: laboratory self-reporting, estimated rather than disclosed figures, revisable historical values, retroactive re-anchoring, release price standing in for realised price, a vendor measuring its own user mix. All six are recorded as established confounders in SOURCES.md (independently verified). Remedy: corroborating sources, raw rather than fitted inputs, vintage pinning. Bounded but rarely quantifiable.

Aggregation uncertainty. Error introduced by the method: weights, normalisation, thresholds, functional form, missing-data policy, correlation treatment. Remedy: the five sensitivity tests above, which exist to make this uncertainty visible. This is the only one of the four the tracker can fully enumerate, because every lever is a choice recorded in a versioned file.

Structural uncertainty about whether the target is being measured at all. Whether the roster measures progress toward general capability rather than progress on the evaluations that happen to exist. This one is the largest, and the reason is specific rather than rhetorical: no agreed definition of AGI exists and this document asserts none, so there is no criterion against which construct validity could be established; the benchmark-to-reality gap is argued to exist but has no published quantification (G3, established by search 2026-07-30, class missing), and the horizon source's own operator has stated it does not know the gap's size (G3, independently verified by search); the jagged frontier argument holds that capability is not one coherent variable, which undercuts any single-figure reading; and the reliability dimension has no public source at all (G1, established by search 2026-07-30), so one of eleven core indicators is a declared gap rather than a measurement. Measurement uncertainty is bounded by repetition, source uncertainty by corroboration, aggregation uncertainty by sensitivity analysis. Structural uncertainty is bounded by nothing available to this instrument, which is why the flag exists, why the weakest-link floor exists, and why the headline is three components rather than one.

Pre-registered kill conditions. Committed with the threshold set, before the data. Each names a numeric or logical condition and the specific consequence. This set is the criterion that stops the instrument becoming a confirmation device: the conditions are written so that they can fire against the maintainer's expectations, and firing is a publication event, not a private decision.

  • K1. Headline invalid: two consecutive releases in which the flag reads INSTRUMENT-FAILURE. Consequence: the

position headline is suspended and the release publishes the instrument diagnosis instead of a position.

  • K2. Headline invalid: more than one third of the pre-registered thresholds in state UNKNOWN. Consequence: publish

the unknown count as the headline, no position claim.

  • K3. Rate retired for the release: fewer than six qualifying observations in the 24-month window, or the horizon

source past 180 days without a release. Consequence: no rate published, staleness stated in the headline block.

  • K4. Method A reference computation retired permanently: three consecutive releases in which min-max and z-score

disagree in the direction of change. Consequence: the reference index is withdrawn as uninformative and the withdrawal is published.

  • K5. Method C diagnostic retired, or downgraded to explicitly multi-factor: any of its three unidimensionality

tests failing for two consecutive releases. Consequence: no single-factor figure is published, and the failing diagnostic is published in its place.

  • K6. Threshold set retired and re-registered: primary and adversarial alternative sets disagreeing in the direction

of position movement for three consecutive releases. Consequence: the primary set is declared unfit, and the re-registration is published with both series in parallel.

  • K7. Indicator retired: an indicator whose removal reverses the sign of the headline change in four of six

consecutive releases. Consequence: it is either given its own published series outside the headline or dropped, with the decision recorded.

  • K8. Whole-tracker kill: fewer than three core dimensions carrying at least one live, non-stale, non-gap indicator.

Consequence: the tracker publishes its source-mortality state and no headline at all until coverage is restored.

  • K9. Contamination kill on the B gate: the contamination-controlled ratio unobserved for four consecutive releases.

Consequence: A-dimension thresholds revert to UNKNOWN, because a met depth threshold with no contamination check is not evidence.

  • K10. Floor retired for the release: fewer than two core dimensions passing both floor scope gates, that is k < 2.

Consequence: no floor is published, the citation string carries FLOOR SUPPRESSED with k and the excluding gate per dimension, and the release states that the weakest-link protection is unavailable this release. Two consecutive releases at k < 2: the position headline is suspended and the release publishes the coverage diagnosis instead, on the same logic as K1.

  • K11. Reproducibility finding published, and the threshold set reviewed: any crossing held in PROVISIONAL for three

consecutive releases, or a median time-to-confirmation exceeding two release intervals over the trailing four releases. Consequence: the time-to-confirmation series is published as a standing finding about the field with the announcing party named per crossing, and the affected threshold is reviewed for whether it was set on a quantity no second party can reproduce, which is a specification error rather than a fact about capability.

Confidence-band convention. HIGH, MEDIUM or LOW per the house convention (sBs/AGI-handbook/bible/SPINE.md), attached to a stated horizon and decaying as the horizon lengthens. Applied per component rather than to a composite: each of position, floor and rate carries its own grade at 12 and at 36 months, and no aggregate grade is published, because there is no aggregate figure to grade. A component's grade is capped by the lowest grade among its contributing indicators. Every stored observation additionally carries the 90-day staleness flag (sBs/fow/kb/INDEX.md), evaluated against the date its source last published a new value, not against the date the tracker last fetched: a fetch of an unchanged file does not refresh staleness. A stale observation is excluded from the position count, excluded from the rate window, and sets the flag to STALE. Epistemic tags ESTABLISHED, PROBABLE, CONTESTED, SPECULATIVE apply to written interpretation, not to observations, which carry verification classes instead.

#

Scope. This register enumerates the ways this instrument can produce a number that is wrong, stale, or unfalsifiable. It is written against the 20-indicator roster (A1, A2, B1, B2, C1, C2, C3, D1, D2, D3, D4, E1, E2, F1, G1i, G2i, G3i, H1, H2, I1) and the verified source register (S1 to S37, gaps G1 to G5, confounders 1 to 7).

Reading conventions. Every quantitative claim carries a source ID and one verification class from independently verified / self-reported / estimated / expert judgment / inferred / missing. No URL is authored anywhere in this section. Where no runnable detector exists, the field says so explicitly rather than naming an aspiration. Detector IDs (FD-nn) are defined once in the inventory at the end of this section and referenced by ID above it.

The register's own bias. Twenty-seven modes are catalogued and the honest count of those with a detector that can be executed on a scheduled run is smaller than twenty-seven. The gap between those two counts is the instrument's real error budget, and it is published rather than hidden.

FM-01. Goodhart's law

  1. Distortion Once the tracker is public and its indicator list is fixed, the indicators stop

    describing progress and start describing what laboratories optimise. A1 (Frontier Eval Headroom) is the most exposed, because a published headroom figure names precisely which evaluations are worth buying gains on. C1 and C2 (task time horizons) are next, because a published horizon target defines the task-length distribution worth training against. The composite floor in the section 6 Position-and-Rate method inherits every distortion in its binding axis.

  2. Likelihood LIKELY. The mechanism requires only that the tracker be read by anyone with a

    training budget, and the tracker's stated purpose is to be read.

  3. Severity HIGH. Goodharting does not add noise, it adds a consistent upward bias to

    position and rate simultaneously, which is the direction Victor is least able to detect from inside the instrument.

  4. Detection FD-06 (held-out versus public delta) and FD-05 (saturation half-life). A metric

    under targeted optimisation shows a widening public-minus-held-out gap on B1 while A1 headroom closes faster than the historical A2 half-life predicts. Neither detector proves intent, and no detector available to this tracker can; they detect the signature only.

  5. Mitigation Keep at least one indicator in each core dimension sourced from a withheld set

    (S10 private, S11 private, S12 held-out 858 tasks, S37 unpublished tiers; all independently verified as withheld per SOURCES.md), so a targeted gain on a public set is visible as a divergence rather than as progress. Publish the indicator list but never publish a weighting that would let an actor compute the marginal value of a specific eval.

  6. Verdict KEEP WITH CAVEAT. A1 keeps its CORE class but is never reported without its

    paired B1 divergence figure on the same view. C1 and C2 keep. The weakest-link floor is retained precisely because it refuses to let one optimised axis carry the headline.

FM-02. Benchmark gaming

  1. Distortion Distinct from FM-01: gaming is per-submission rather than per-training-run.

    Best-of-n sampling, retry loops, verifier-guided reranking, and answer-format tuning raise a reported score without changing the underlying competence. This hits A1 directly, E1 (Real-Task Win Rate) through S6 self-submitted artefacts, and D1 (pass@1 to pass@k Spread) in the perverse direction, because a heavily reranked pass@1 already contains hidden k.

  2. Likelihood NEAR-CERTAIN. S6's submission model is open and submitter-driven

    (S6, independently verified: results are submitted in-tree with metadata.yaml and trajectories), so the incentive and the mechanism are both present today.

  3. Severity HIGH. It inflates A1 and E1 while simultaneously suppressing D1, which makes the

    system look both more capable and more reliable at once.

  4. Detection FD-04 plus FD-14. FD-14 (scaffold-disclosure completeness check) asserts that

    each observation carries a declared sampling budget, retry policy, and verifier; observations failing it are quarantined rather than scored. S6 has required reasoning traces since July 2024 (S6, independently verified), so trajectory presence is machine-checkable even when trajectory content is not machine-adjudicable.

  5. Mitigation Schema-level: the section 10 record requires a scaffold block, and the pipeline

    refuses an observation whose scaffold block is absent. Prefer sources where the evaluator runs the model itself (S10 verified scores, S12, S13, S16 historic, S37) over sources where the submitter runs it (S6, S14).

  6. Verdict KEEP WITH CAVEAT. A1 and E1 keep, restricted to evaluator-run observations for the

    headline and submitter-run observations relegated to a secondary series. D1 keeps but is reported only against observations with a declared sampling budget.

FM-03. Benchmark contamination

  1. Distortion Public evaluation items enter pretraining corpora through ordinary web

    crawling, so the score measures recall rather than reasoning. Every indicator built on a public set is exposed: A1, B2 (Contamination-Controlled Score Ratio) by construction, E1, and the S1-sourced score feed behind A2 and G1i.

  2. Likelihood NEAR-CERTAIN. It requires no intent; S11's public set was finalised on

    2025-04-03 (S11, independently verified) and has been crawlable since, and S15 is password-gated on GitHub explicitly to resist scraping (S15, independently verified), which is itself evidence that the operators treat the risk as live.

  3. Severity CRITICAL. Contamination produces the single most misleading possible reading:

    apparent generalisation that is memorisation, on the axis (breadth and depth) the headline weights most.

  4. Detection FD-06. Compute the ratio of performance on a withheld set to performance on the

    matched public set from the same family: S11 public versus S11 private, S12's 731 public versus 858 held-out (S12, independently verified as counts), S10 public tasks versus S10 semi-private and private. A ratio that falls over successive model generations is the contamination signature. The detector is runnable only where the operator publishes both halves; where only the withheld half is published, FD-06 degrades to a level check with no ratio.

  5. Mitigation B2 exists for this purpose and is CORE for this reason. The novelty discount in

    the operational target is applied as a multiplier derived from B1 and B2, so a contaminated gain is discounted rather than counted. Every observation pins a model release date and an evaluation-set publication date in the schema, so an item-set-newer-than-model check is possible per row.

  6. Verdict KEEP. B1 and B2 keep as CORE and are load-bearing. A1 keeps with its

    contamination discount applied and never reported raw in the headline.

FM-04. Training on evaluation data

  1. Distortion The deliberate variant of FM-03: evaluation items, or near-duplicates of them,

    included in a training or fine-tuning set. Same indicator exposure as FM-03 (A1, B2, E1) but with a sharper profile, because a deliberately trained-on set can be saturated in a single release, which corrupts A2 (Saturation Half-Life) and therefore the rate half of the headline.

  2. Likelihood POSSIBLE. Distinguishable from FM-03 only with training-corpus access, which

    this tracker does not have and will not have; the honest position is that its rate cannot be estimated from outside.

  3. Severity CRITICAL. It breaks the rate signal, and rate is the half of the headline that

    claims to detect acceleration.

  4. Detection No direct detector exists and none can be built from outside a laboratory. State

    this plainly. The only observable proxies are FD-06 (a public-versus-withheld gap that widens discontinuously at one release rather than gradually) and the canary approach: S5 carries a "DO NOT TRAIN ON THIS DATA" canary string on its page (S5, independently verified), and canary recall can be probed only by whoever can query the model, which the tracker can do only for models it has API access to. Treat canary probing as a manual, opportunistic check, not a scheduled detector.

  5. Mitigation Weight withheld-set sources for the headline (S10 private, S11 private, S12

    held-out, S37 unpublished tiers) and treat public-set movement as corroboration only. Record a per-indicator contamination_control field in the schema with values withheld / rolling / public / unknown, and refuse to compute the novelty discount from any indicator marked unknown.

  6. Verdict KEEP WITH CAVEAT for A1 and E1, both flagged with contamination_control at row

    level. B2 keeps as CORE. The caveat is published as an acknowledged blind spot rather than mitigated away.

FM-05. Selective reporting by laboratories

  1. Distortion A laboratory reports the evaluations it did well on and omits the rest, so the

    observed distribution of scores is a maximum rather than a sample. This hits E1 and A1 wherever the feed is submission-driven, and it hits D4 (Generation-over-Generation Regression Count) hardest, because a regression is exactly the result least likely to be submitted. S6's scores are self-submitted and depend on a laboratory choosing to submit at all, so absence of a score is not absence of capability (S6, independently verified as submission-driven).

  2. Likelihood NEAR-CERTAIN. Submission is voluntary on S6 and S14, and reporting on S20 is

    laboratory self-reported by class (S20, class LAB per SOURCES.md).

  3. Severity HIGH. It biases position upward and suppresses D4 toward zero, which makes the

    instrument structurally unable to show regression, the exact failure the BRIEF names.

  4. Detection FD-09 (submission-coverage census). For each model in S2's notable-models table,

    assert presence or absence of a matched observation in each submission-driven feed, and publish the coverage fraction alongside the indicator. A falling coverage fraction with a rising score is the selective-reporting signature. This is fully runnable: S2 exposes model CSVs and S6 exposes an in-tree directory structure (both independently verified as machine-readable).

  5. Mitigation Publish coverage as a first-class field next to every submission-driven

    indicator, and state in the interpretation layer that missing is not zero and not absent capability. Compute D4 by diffing successive S1 dumps rather than from any submission feed, per gap G2's recommendation.

  6. Verdict KEEP WITH CAVEAT. E1 keeps with a published coverage fraction. A1 keeps. D4 keeps

    but is explicitly self-computed, not sourced, because gap G2 establishes that no live feed tracks regression.

FM-06. Cherry-picking model variants

  1. Distortion A family ships many configurations (reasoning effort levels, thinking budgets,

    tool-enabled variants, checkpoints) and only the strongest configuration appears in the record, so the tracked entity is a best-of-family rather than a shipped product. This corrupts C1, C2 and C3 (horizons), G2i (Inference Cost per Solved Task at Fixed Reliability) severely, since the strongest variant is usually the most expensive, and H1 (Open-Weight Lag) by comparing a maximal closed variant against a default open one.

  2. Likelihood NEAR-CERTAIN. Variant proliferation is already the norm, and S13 tracks on the

    order of 180 models (S13, independently verified) precisely because the variant space is large.

  3. Severity HIGH for G2i (it can invert the sign of an efficiency trend) and MODERATE for the

    horizon set.

  4. Detection FD-19 (variant-identity assertion). Require a fully qualified model identifier

    plus configuration fields in the schema, then assert that a family's series does not switch configuration class between consecutive observations. Switches are flagged, not silently accepted. Runnable, because it is a self-consistency check over the tracker's own records.

  5. Mitigation Define, per indicator, whether the tracked variant is best-available or

    default-deployed, and never mix the two in one series. G2i uses default-deployed. A1 and the horizon set use best-available and say so.

  6. Verdict KEEP WITH CAVEAT for all of C1, C2, C3, G2i and H1, each carrying an explicit

    variant-policy declaration in its section 4 row.

FM-07. Private evaluations that cannot be audited

  1. Distortion The withheld sets that defend against FM-03 and FM-04 cannot be inspected by

    anyone outside the operator, so their difficulty, item quality, and grading consistency are taken on trust. S10, S11, S12 and S37 all hold withheld sets (all independently verified per SOURCES.md). Indicators exposed: B1 and B2 (which depend on withheld halves by construction), A1 where its headroom comes from S37, and E1 where it comes from S12.

  2. Likelihood NEAR-CERTAIN. It is not a risk but the current structural condition of every

    contamination-resistant source in the register.

  3. Severity HIGH. It does not bias the number in a known direction; it removes the tracker's

    ability to say why the number moved, which is worse than a known bias for a defensibility goal.

  4. Detection Partially detectable only. FD-18 (human-baseline presence check) asserts that a

    withheld set publishes a human baseline and item counts; S10 publishes a human baseline board for its third generation (S10, independently verified by search, not fetched) and S12 publishes its split counts. What cannot be detected is grader drift or item-quality change inside a set the tracker cannot see. There is no detector for that and there will not be one.

  5. Mitigation Per-indicator disclosure rather than resolution. Each section 4 row carries an

    auditability field with the SOURCES.md class (IND / LAB / AGG / GOV) and a free-text withheld-set note. Where a withheld source's governance is contested, the contest is disclosed in the row. At S37 the funding arrangement is a known controversy that the verifier could not confirm from the page (S37, missing); it is recorded as unconfirmed and disclosed, not asserted.

  6. Verdict KEEP WITH CAVEAT, all four of B1, B2, A1 and E1. This is the central trade of the

    whole instrument: contamination protection is bought with lost external auditability, the tracker cannot resolve the trade, and it therefore discloses which side of it every indicator sits on.

FM-08. Test-set saturation

  1. Distortion As scores approach the ceiling, the metric loses resolution and further

    capability growth becomes invisible, so a saturating evaluation reports a plateau that is a property of the ruler. A1 is directly affected. A2 exists to measure the phenomenon, so A2 is the instrument's own saturation sensor rather than a victim of it. S4's latent index is re-anchored as benchmarks saturate (SOURCES.md confounder 1, independently verified), which propagates saturation effects into anything built on S4, including H1 via S9.

  2. Likelihood NEAR-CERTAIN. Saturation is the observed life cycle of every benchmark in the

    register that has been live for more than a short period, and S27's retirement and S16's maintenance mode are partly downstream of it (S16, independently verified: maintenance mode from 1 June 2026).

  3. Severity CRITICAL for the headline, because a saturation plateau and a genuine capability

    plateau are indistinguishable at the level of a single score, and the BRIEF's exit condition requires the instrument to tell them apart.

  4. Detection FD-05 (saturation half-life, that is A2 itself) plus FD-20 (ceiling-proximity

    flag). FD-20 asserts, per evaluation, the distance between the frontier score and the published ceiling or human baseline, and marks any evaluation inside a declared proximity band as resolution-exhausted. Both runnable from S1's dump, which was observed updated on 2026-07-30 (S1, independently verified).

  5. Mitigation A1 is defined as headroom rather than score, so it degrades gracefully as

    saturation approaches: headroom near zero is a statement about the evaluation, not the model. The pipeline retires an evaluation from A1's basket when FD-20 fires and records the retirement with a methodology version bump, so the retirement is visible in history.

  6. Verdict KEEP. A1 keeps with mandatory FD-20 flagging. A2 keeps as CORE and is promoted in

    emphasis, since it is the discriminator the exit condition depends on.

FM-09. Changing evaluation difficulty

  1. Distortion An operator revises a set (harder items, removed items, corrected labels) and

    the series silently mixes two difficulty regimes, so a score drop reads as capability regression and a score rise reads as progress. S14 has moved 1.0 to 2.0 to 2.1 (S14, independently verified) and S6 carries five splits (S6, independently verified). Affected: A1, A2, E1, D4 (a difficulty change manufactures false regressions), C1 and C2 wherever the task suite behind S5 changes.

  2. Likelihood NEAR-CERTAIN. Two of the register's most-used sources have already done it.
  3. Severity HIGH, and specifically dangerous because it produces false regression, which is

    the reading Victor would be least inclined to doubt given the BRIEF's insistence that the instrument be able to show regression.

  4. Detection FD-07 (version-pin mismatch check). Every observation records benchmark version;

    the pipeline refuses to compute a delta across differing versions and instead opens a new series. S14 submissions must pin a Harbor dataset version (S14, independently verified), so the pin is available at source. Runnable and cheap.

  5. Mitigation Schema requires benchmark_version as a required field per C11. Cross-version

    comparison is permitted only through an explicitly published bridge with overlapping models evaluated on both versions, and the bridge is itself versioned.

  6. Verdict KEEP for A1, A2, E1, C1, C2. D4 KEEP WITH CAVEAT: a regression is counted only

    when benchmark version and methodology version are both held constant, otherwise it is logged as a version artefact and excluded from the count.

FM-10. Inconsistent scaffolding and prompting

  1. Distortion The same model scores differently under different agent harnesses, system

    prompts, tool sets, and context budgets, so a cross-model or cross-time comparison confounds model capability with harness engineering. Affected: E1 and A1 across submission-driven feeds, C1 to C3 (horizon results are harness-dependent by nature), F1 (embodied policies are inseparable from their controller stack), and G2i (cost depends on harness as much as model).

  2. Likelihood NEAR-CERTAIN. Agentic evaluation is harness-mediated everywhere in the

    register; S16's distinguishing property is prompt-level transparency (S16, independently verified), which implies its absence elsewhere.

  3. Severity HIGH. It adds a large, non-random, and improving-over-time component to every

    agentic series, which inflates rate.

  4. Detection FD-14 (scaffold-disclosure completeness check), plus FD-21 (harness-constancy

    assertion) which flags any series where the harness identifier changes between consecutive observations. Runnable where the source discloses a harness; where it does not, the detector reports undisclosed, which is itself the finding.

  5. Mitigation Prefer sources that fix the harness centrally (S10's own runner, S12 where

    Scale runs the evals, S13 which runs its own evaluations across roughly 180 models, S37 which Epoch runs; all independently verified as operator-run). Record harness and prompt-policy fields in the schema. Report agentic indicators as bands rather than points, per the section 6 uncertainty-band requirement.

  6. Verdict KEEP WITH CAVEAT for A1, E1, C1, C2, C3, G2i. F1 DEMOTE TO CONTEXT: harness

    confounding plus gap G4 (no robotics benchmark with both unstructured tasks and a machine-readable series) leaves F1 unable to support a headline claim, so it remains a supporting indicator reported at annual granularity from S21 and never enters the floor.

FM-11. Hidden human assistance in autonomous results

  1. Distortion A result presented as autonomous includes human intervention: hand-written

    scaffolds tuned per task, retries triggered by a human observer, human-authored intermediate plans, or human grading that supplies signal. This attacks the operational target directly, since the target counts tasks completed "without human assistance". Affected: C1, C2, C3 (the whole horizon dimension), E1, and F1 where teleoperation or hand-tuned resets are involved.

  2. Likelihood POSSIBLE. The register gives no evidence of it, and none is alleged; the

    exposure is structural, because most autonomy results are produced by the party reporting them and the boundary between harness engineering and human assistance is not sharply defined anywhere in the register.

  3. Severity CRITICAL if present, because it falsifies the constitutive claim of the C

    dimension rather than biasing it.

  4. Detection Partial only. FD-22 (trajectory-availability assertion): S6 has required

    reasoning traces since July 2024 and stores trajs/ in-tree (S6, independently verified), so trajectory presence is machine-checkable. Trajectory content is not machine-adjudicable at this tracker's resource level, so a spot audit of a sampled trajectory is manual and occasional. There is no automated detector for hidden assistance and there is unlikely to be one.

  5. Mitigation Require an autonomy_class field per observation (unassisted / harness-assisted

    / human-in-loop / undeclared) and exclude undeclared observations from C1 to C3. Prefer S5, whose analysis code is published (S5, independently verified), so the assistance boundary is at least inspectable in code.

  6. Verdict KEEP WITH CAVEAT for C1, C2, C3, with autonomy_class mandatory and undeclared

    observations excluded rather than downweighted.

FM-12. Unequal compute budgets

  1. Distortion Inference-time compute is now a free variable, so a score difference can be

    purchased rather than earned. Two systems compared at unequal token, sample, or wall-clock budgets are not comparable. Affected: A1, E1, C1 and C2 (longer horizons can be bought with longer runs), H1 (open-weight models are frequently evaluated at smaller budgets), and G2i, which is the indicator designed to hold this variable fixed.

  2. Likelihood NEAR-CERTAIN. Variable inference budget is a standard product feature and the

    register contains no source that normalises it across all its rows.

  3. Severity HIGH. Unnormalised, it makes the denominator of the operational target

    (per unit of resource consumed) unenforceable, which converts the whole tracker into a capability-at-any-price measure.

  4. Detection FD-15 (cost-per-solved-task at fixed reliability, that is G2i itself) and FD-23

    (budget-disclosure assertion). FD-15 is runnable where a source publishes price and score together: S13 publishes pricing alongside headline indices on its free tier (S13, independently verified). FD-23 is runnable as a schema completeness check. What is not runnable is recovering an undisclosed budget after the fact.

  5. Mitigation G2i is CONTEXTUAL but load-bearing per the BRIEF's denominator ruling, and no

    capability indicator is reported in the headline without its matched G2i figure on the same view. Where budget is undisclosed, the observation is admitted with a wider uncertainty band rather than excluded, and the widening is recorded.

  6. Verdict KEEP. G2i keeps as the denominator. A1, E1, C1, C2, H1 KEEP WITH CAVEAT, each

    carrying a budget-disclosure flag that widens the published band when unset.

FM-13. Best-case versus median performance

  1. Distortion A single reported score is usually a single run or a maximum over runs, so it

    describes the best case rather than the distribution. Users experience something closer to the median, and a tracker built on best cases overstates position and, worse, understates variance. D2 (Rerun Variance) is the indicator that would catch this and D3 (Calibration Error) is its companion. Affected by the absence: A1, E1, C1, C2, and every reliability weight in the operational target.

  2. Likelihood NEAR-CERTAIN. Single-run reporting is the register's default; only S13 is

    verified as publishing confidence intervals (S13, independently verified).

  3. Severity HIGH. The operational target is explicitly reliability-weighted, so a missing

    distribution weakens the definition itself, not just one row.

  4. Detection Under-detected today, and this must be stated plainly. D2 and D3 have no live

    source: gap G1 establishes that no live public leaderboard exists for calibration or run-to-run variance, with partial coverage only from S16 (now in maintenance mode) and S13's confidence intervals. The consequence is that the failure modes D2 and D3 exist to catch are currently under-detected. FD-24 (interval-presence assertion) is runnable and reports how many observations carry any dispersion estimate at all, which measures the size of the hole rather than closing it.

  5. Mitigation Self-compute where API access allows: repeated runs on a small fixed task

    sample, published as an own-measurement series with an explicit self-reported verification class. Where that is not affordable, D2 and D3 report status unavailable rather than a number, per gap G1's instruction that a reliability indicator must be self-computed or declared unavailable.

  6. Verdict D2 and D3 KEEP WITH CAVEAT, both shipping in an explicit unavailable-or-

    self-computed state and excluded from the weakest-link floor until a dispersion estimate exists, so a missing reliability number cannot silently read as a passing one. A1, E1, C1, C2 KEEP WITH CAVEAT with a best-case-versus-median flag per row.

FM-14. Reliability obscured by pass@k

  1. Distortion A pass@k figure counts a task as solved if any of k attempts succeeds, which

    answers a different question from the one the operational target asks. Reported without k, or with k varying across the series, it converts an unreliable system into a capable-looking one. D1 (pass@1 to pass@k Spread) exists specifically to detect this. Say so. Affected: A1, E1, and C1 and C2 wherever horizon success is scored with retries.

  2. Likelihood LIKELY. pass@k reporting is common in code and agentic evaluation, and the

    register does not verify a source that forbids it.

  3. Severity HIGH. It attacks the reliability axis directly, and reliability is a constitutive

    axis of the target rather than a footnote.

  4. Detection FD-04, which is D1 operationalised: for every evaluation where both pass@1 and

    pass@k are available, publish the spread; where only pass@k is available, refuse to admit the observation to a reliability-weighted computation. Runnable where both figures exist, and where they do not, the detector reports the absence.

  5. Mitigation Schema requires k and pass_at_k fields per observation, with k = 1 asserted

    rather than assumed. The reliability weight in the operational target is computed from pass@1-equivalent figures only.

  6. Verdict KEEP. D1 keeps as CORE and is the named detector for this mode. A1 and E1 KEEP

    WITH CAVEAT, admitted to the headline only at declared k.

FM-15. Model release selection bias

  1. Distortion The population of models that appear in the record is not the population of

    models that exist. Laboratories choose whether to submit, whether to permit evaluation, and whether to publish at all, so the observed frontier is a survivor set. S6's scores are self-submitted and depend on a laboratory choosing to submit at all, so absence of a score is not absence of capability (S6, independently verified). Affected: A1, E1, H2 (Frontier Concentration, which is computed over whoever is visible), H1, and D4, since an unsubmitted weaker successor is an invisible regression.

  2. Likelihood NEAR-CERTAIN. Voluntary submission plus withheld internal models is the current

    structure, not a hypothetical.

  3. Severity HIGH for H2 and D4, MODERATE for A1, where the bias mostly affects the tail

    rather than the frontier estimate.

  4. Detection FD-09 (submission-coverage census) run against S2's notable-models list and

    S24's company data, both machine-readable (S2 and S24, independently verified). Publish the fraction of known-released models with a matched observation per feed. A falling fraction is the signature. Runnable on every scheduled run.

  5. Mitigation Publish coverage next to every affected indicator, and state in the

    interpretation layer that the tracker measures the visible frontier, not the frontier. Compute D4 from S1 dump diffs per gap G2, which captures any model S1 covers regardless of voluntary submission elsewhere.

  6. Verdict KEEP WITH CAVEAT. A1, E1, H1, H2 keep with a published coverage fraction. D4 keeps

    as a self-computed diff, with its known blind spot (models never evaluated anywhere) stated.

FM-16. Non-comparable historical results

  1. Distortion Historical rows were produced under different versions, harnesses, graders and

    metric definitions, so a long series is a chimera and its slope is partly an artefact of its own construction. This attacks the rate half of the headline, which is the half that would detect acceleration or stagnation. Affected: A2, C3 (Horizon Doubling Time), G3i, and any backfilled portion of A1 and E1.

  2. Likelihood NEAR-CERTAIN. S14's three live versions, S6's five splits, and S4's retroactive

    re-anchoring (all independently verified per SOURCES.md) each produce it independently.

  3. Severity CRITICAL. A wrong rate is worse than a wrong level for this instrument, because

    rate is what a public commentator is tempted to extrapolate from.

  4. Detection FD-07 (version-pin mismatch), FD-08 (re-anchor drift monitor: recompute a stored

    historical value from the current source and alert when it has moved), and FD-25 (vintage diff: compare today's fetched history against the archived vintage of the same series). All three runnable, and FD-08 is specifically required because S4 re-anchors historical values retroactively.

  5. Mitigation Vintage archiving, and never overwriting history. Every fetch is stored as an

    immutable dated vintage; the current view is computed from the latest vintage while prior vintages remain queryable, so a retroactive source revision appears in the tracker as a revision event rather than as a change in the past. The schema's revision fields carry this per C11.

  6. Verdict KEEP WITH CAVEAT. A2, C3, G3i keep, each reported with its methodology version and

    a visible revision history. Any pre-vintage backfilled segment is drawn distinctly and excluded from rate computation.

FM-17. Overreliance on synthetic evaluations

  1. Distortion Simulated or constructed tasks stand in for real work, and the gap between them

    is unmeasured, so a rising simulated score is read as a rising real-world capability. F1 is entirely simulated: S21 is simulation, not unstructured reality (S21, independently verified). E2 (Benchmark-to-Reality Gap) is the indicator meant to bound this. E1 is partly exposed wherever its task set is constructed rather than sampled from real work.

  2. Likelihood NEAR-CERTAIN for F1 given gap G4. LIKELY for E1.
  3. Severity HIGH for F1, MODERATE for E1, because S22 establishes that economic-value

    benchmarks measure isolated tasks rather than end-to-end situated work (S22, independently verified), which bounds the interpretation without quantifying the error.

  4. Detection FD-26 (realism-class assertion): every observation declares its task provenance

    as sampled-from-real-work / expert-constructed / simulated, and the tracker publishes the mix behind each indicator. Runnable as a schema check. What is not runnable is a numeric gap: gap G3 establishes that no published quantification of the benchmark-to-real-world gap exists, and METR's own note states they do not know its size (G3, independently verified by search), so any numeric gap claim here would be invented and none is made.

  5. Mitigation E2 is specified as a bounded, framed indicator rather than a measured number,

    per gap G3. F1 is annual-granularity supporting only. The interpretation layer carries S22's isolated-task finding verbatim as a standing caveat on E1 and E2.

  6. Verdict F1 DEMOTE TO CONTEXT, reported annually from S21 and never in the floor. E2 KEEP

    WITH CAVEAT as a qualitative bound with no numeric value, explicitly labelled missing rather than estimated. E1 KEEP WITH CAVEAT with a published realism mix.

FM-18. Capability overhang

  1. Distortion Capability exists inside laboratories before any evaluation observes it, so

    every indicator is a lower bound with an unknown lag, and a genuine acceleration can appear as a plateau until it is released and evaluated. Affected: all of A1, A2, B1, B2, C1, C2, C3, E1. The direction is systematic: the tracker will always be late, never early.

  2. Likelihood NEAR-CERTAIN. It follows from the fact that evaluation happens after training,

    and S2's own cadence records major models arriving within roughly two weeks of release (S2, independently verified), which bounds the release-to-record lag but says nothing about the internal-to-release lag.

  3. Severity MODERATE for position, HIGH for the interpretation of a plateau, since an overhang

    period and a stagnation period look identical from outside.

  4. Detection No runnable detector exists for the internal-to-release lag; the tracker cannot

    observe unreleased systems and should not pretend otherwise. What is runnable is FD-27 (release-to-record lag monitor), which measures the second, observable half of the lag from model release date to first observation date, per indicator. That bounds the tracker's own latency without bounding the laboratory's.

  5. Mitigation Publish the observable lag as a standing metadata figure on every view, and

    state in the interpretation layer that all positions are lower bounds. Never describe a plateau as stagnation on position alone: require corroboration from A2 and D4 per the plateau detectors below.

  6. Verdict KEEP for all affected indicators, with a mandatory lower-bound label on position

    and a published release-to-record lag. No indicator is dropped, because overhang is a property of the world, not a defect in a row.

FM-19. Deployment lag

  1. Distortion Distinct from FM-18: capability exists and is measured, but is not deployed,

    so any indicator drawing on deployment or usage data moves later and more slowly than capability. Affected: E1 where its task mix reflects deployed usage, H1 (Open-Weight Lag, which is a deployment-timing measure by construction), and anything drawing on S7, whose releases are roughly quarterly and tightening (S7, independently verified).

  2. Likelihood NEAR-CERTAIN. Quarterly and annual cadences on the adoption-side sources make

    lag structural: S17 is annual only (S17, independently verified), S33 carries a currently-being-updated banner (S33, independently verified).

  3. Severity MODERATE. It distorts timing rather than level, and the BRIEF already forbids

    summing deployment into capability.

  4. Detection FD-28 (cadence-versus-lag ledger): per source, publish declared cadence, observed

    last-update date, and computed reporting lag. Runnable on every run and shared with FD-01.

  5. Mitigation Deployment-side indicators are never summed into the capability headline, per

    the BRIEF's eight-way resolution. H1 is CONTEXTUAL. Views that mix capability and deployment series carry different time axes or an explicit lag annotation.

  6. Verdict KEEP. H1 keeps as CONTEXTUAL with a published lag. E1 keeps as SUPPORTING with its

    deployment-sensitive components separated from its capability components.

FM-20. Conflating economic adoption with intelligence

  1. Distortion Adoption and revenue series are read as capability, when they move with

    pricing, product decisions, sales motion, and customer mix. The sharpest case is S7: it measures one laboratory's own user mix, so it moves with that laboratory's product and customer changes, not with the economy (SOURCES.md confounder 4, independently verified). Affected: E1 and E2 if adoption data leaks into them, and any headline that borrowed an adoption series for momentum.

  2. Likelihood LIKELY. Adoption data is the most abundant and most quotable data in the

    register, which makes it the most likely to be misused by a reader even if the tracker is disciplined.

  3. Severity HIGH for credibility, because it is the specific error the BRIEF names as a way

    the tracker could embarrass its author, and because an adoption boom during a capability plateau would make the instrument report acceleration that is not there.

  4. Detection FD-29 (layer-provenance assertion): every field in the capability computation is

    tagged with its layer, and the pipeline fails if any adoption-class source (S7, S17, S24, S31, S32, S33, S34) appears in a capability-layer computation. Runnable as a static check over the pipeline configuration, and it is the cheapest high-value detector in this register.

  5. Mitigation The BRIEF's separation ruling is enforced in code, not in prose: economic

    usefulness and deployment scale are tracked on their own axis and never summed. S7 is labelled in the UI as one laboratory's own user mix and self-reported by class.

  6. Verdict DEMOTE TO CONTEXT for every adoption series including S7-derived figures: they

    remain visible as context panels, are never inputs to the capability position or rate, and E1 keeps as SUPPORTING with its task-performance component only.

FM-21. Conflating compute growth with capability growth

  1. Distortion Rising training compute is read as rising capability, which turns an input into

    an outcome. G3i (Frontier Training Compute) uses S2, whose compute values are Epoch estimates derived from parameter and token counts rather than laboratory disclosures, with wide error bars revisable without notice (SOURCES.md confounder 2, estimated). The consequence deserves stating precisely: the estimate can be revised, which means G3i can move without any change in the world. G1i (Compute to Reach a Fixed Capability Threshold) is exposed through the same estimates, and G2i through S3, which records release price rather than realised or rental price (SOURCES.md confounder 3, independently verified).

  2. Likelihood NEAR-CERTAIN. Compute series are the most available, the most extrapolated, and

    the most frequently presented as progress in public discussion.

  3. Severity HIGH. Treating the denominator as the numerator inverts the operational target,

    which measures capability per unit of resource, not resource.

  4. Detection FD-12 (estimate-revision flag on S2): store each fetched compute value as a

    vintage and alert when a historical value changes between vintages; publish the revision alongside the series. Plus FD-29 (layer-provenance assertion), which fails the build if a resource-class source is used as a capability input. Both runnable, and FD-12 is necessary because revision without notice is documented behaviour for this source.

  5. Mitigation G1i, G2i and G3i are CONTEXTUAL denominators per the BRIEF and are rendered on

    a separate axis with an explicit estimated verification class. G3i carries a visible revision log. Any FLOP-per-dollar figure derived from S3 is labelled release-price-based, since the cost frontier laboratories actually pay is not in the register.

  6. Verdict KEEP as CONTEXTUAL, all three of G1i, G2i, G3i, each tagged estimated, each with a

    published revision history, and none admitted to the capability position or the weakest-link floor.

FM-22. Treating expert judgment as objective data

  1. Distortion Forecasts, survey medians, preference votes, and expert ratings are imported and

    rendered with the same visual authority as measured scores, so opinion enters the instrument wearing measurement's clothes. Exposed: I1 (Capability-Controllability Divergence), whose controllability side has no measured feed and rests substantially on judgment; anything drawn from S23 (play money, thin liquidity, poorly calibrated on individual questions, per SOURCES.md, expert judgment); S18, which measures preference rather than capability and whose Elo pipeline is operator-controlled (S18, independently verified); and S25, which is API-gated and from which no community median could be confirmed, so none is cited (S25, missing).

  2. Likelihood LIKELY. The pull toward a quotable forecast number is strong precisely when the

    measured picture is ambiguous.

  3. Severity HIGH for credibility, because a sceptical reader will find the judgment input

    first and will treat it as evidence of the whole instrument's discipline.

  4. Detection Under-detected. D4 has no live source at all: gap G2 establishes that no live eval

    or leaderboard tracks generation-over-generation regression, so the regression check that would independently corroborate a judgment-based claim must be self-computed, and until it is, judgment inputs are unchecked. What is runnable is FD-30 (verification-class census): count and publish, per view, how many rendered figures carry the expert-judgment class. That measures the exposure; it does not validate the judgments.

  5. Mitigation The six-level verification class is rendered, not just stored: expert-judgment

    figures are drawn in a visually distinct, deliberately subordinate style, and the four-layer separation puts forecasting in its own layer with no path into the data or methodology layers. No AGI date and no probability of AGI appears anywhere in the instrument.

  6. Verdict I1 KEEP WITH CAVEAT, TRACKED and never summed per the BRIEF, with its

    controllability side labelled expert judgment. S23-derived and S18-derived figures DEMOTE TO CONTEXT. No indicator sources a headline value from S25.

FM-23. False precision in the headline

  1. Distortion A headline figure rendered to more digits than its inputs support implies a

    resolution the instrument does not have, and invites comparison of movements that are inside the noise. The Position number is most exposed, and any composite would be worse, since a composite's apparent precision grows as its inputs are averaged even when its uncertainty does not shrink. Affected transitively: every indicator feeding the floor, and A1, C1, C2, C3 in particular.

  2. Likelihood LIKELY. Precision inflation is the default behaviour of arithmetic, and it takes

    a deliberate rule to stop it.

  3. Severity HIGH. It is the specific way a defensible instrument becomes indefensible in a

    single screenshot.

  4. Detection FD-31 (significant-figure gate): assert that no rendered figure carries more

    significant figures than its least precise input, and that every headline figure ships with an interval. Runnable as a rendering-time assertion, and it fails the build rather than warning.

  5. Mitigation The section 6 recommendation is a two-number Position-and-Rate profile with a

    published uncertainty band, not a single composite digit, and the band is mandatory rather than optional. The document nowhere asserts an agreed definition of AGI, so a single-number "progress score" has no denominator to be precise against, and none is offered.

  6. Verdict KEEP for the Position-and-Rate headline, with an interval always rendered and the

    significant-figure gate enforced at build time. DROP the single composite index as a headline form, retained only as a discussed and rejected alternative in section 6.

FM-24. Politically or commercially motivated data manipulation

  1. Distortion Sources sit inside commercial and institutional structures that create incentive

    exposure around the numbers they publish. Two specific structural vectors exist in this register, and naming them is not an allegation that any named party has done anything wrong. First, S13 operates under a proprietary licence with attribution required and redistribution requiring a contract (S13, independently verified), which means the tracker's access to a feed it depends on is contractually revocable and the terms of that access are set by the party whose numbers are being tracked. Second, at S37 the funding arrangement is a known controversy that the verifier could not confirm from the page (S37, missing), which means a withheld-set evaluation the tracker relies on has a governance question the tracker cannot resolve. Affected: A1 and G2i via S13, A1 and B2 via S37.

  2. Likelihood POSSIBLE as to actual distortion, and it is deliberately rated no higher,

    because the register contains no evidence of it and none is claimed. NEAR-CERTAIN as to structural exposure, which is the thing being registered.

  3. Severity CRITICAL if it occurred, since it would be undetectable from outside and would

    affect a headline input.

  4. Detection No detector for manipulation exists and the tracker should say so rather than

    imply oversight it does not have. Two runnable proxies: FD-13 (licence and terms watcher), which alerts on any change to a source's stated licence or access terms, and FD-32 (cross-source concordance check), which compares indicators derivable from two independent sources (for example an S13-derived figure against an S1-derived figure for the same evaluation) and flags divergence beyond a published tolerance. Divergence identifies a discrepancy, never a cause.

  5. Mitigation Disclose the structural exposure per indicator in neutral terms in the section 4

    row, including the funding-arrangement question at S37 as unconfirmed. Prefer at least two independent sources for any headline input where the register permits it. Never build a headline input on a single proprietary feed without a named fallback.

  6. Verdict KEEP WITH CAVEAT for A1, B2 and G2i, each carrying a disclosed structural-exposure

    note and, where available, a concordance figure from a second source. No indicator is dropped on the basis of an unconfirmed controversy, and no wrongdoing is alleged.

FM-25. Excessive dependence on a small number of laboratories

  1. Distortion Two dependencies, both real. On the source side, a large share of this

    register's machine-readable spine is one organisation's output: S1, S2, S3, S4, S8, S9, S22, S24, S35, S36 and S37 are all Epoch-operated (independently verified per SOURCES.md), so a single organisational change could take out most of the pipeline at once. On the measured side, H2 (Frontier Concentration) is computed over a frontier produced by a handful of laboratories, so the tracker's capability signal is a small-sample statistic and one laboratory's release schedule can move the rate. Affected: A1, A2, C3, G3i, H2, and structurally every indicator drawing on the Epoch spine.

  2. Likelihood NEAR-CERTAIN, as a present structural fact rather than a future risk.
  3. Severity CRITICAL for continuity, HIGH for interpretation. The register already records

    four dead or dormant sources and two that moved host, so single-provider dependence is not a theoretical concern here.

  4. Detection FD-33 (provider-concentration census): compute and publish the share of headline

    inputs by operating organisation, and alert when any single organisation exceeds a declared share. Runnable as a static check over the source configuration. FD-10 (frontier-actor census) does the same on the measured side using S24's company and chip-owner data (S24, independently verified as machine-readable).

  5. Mitigation Name a fallback source per indicator in the schema, and prefer indicators with

    at least two independent operators for the headline. Publish the concentration figure as part of the instrument's own metadata, so a reader can see the dependency rather than discover it during an outage. Where no fallback exists, record fallback as none and treat the indicator as fragile.

  6. Verdict KEEP for A1, A2, C3, G3i with a named fallback recorded per indicator and a

    published provider-concentration figure. H2 KEEP as CONTEXTUAL with an explicit small-sample caveat, never in the floor.

FM-26. Methodological drift over time

  1. Distortion Definitions, fits and versions change under the series, so a comparison across

    time is a comparison across methodologies. The register documents this three times over: S14 has moved 1.0 to 2.0 to 2.1 (independently verified), S6 carries five splits (independently verified), and S13's index methodology is versioned (independently verified). Confounder 1 is the sharpest case: S4's latent index re-anchors as benchmarks saturate, so historical values move retroactively and any cached figure silently drifts (SOURCES.md confounder 1, independently verified; the named anchors observed there are index values, not scores, and are not reproduced as current figures here). Affected: A1, A2, C3, H1 via S9, and every rate computation.

  2. Likelihood NEAR-CERTAIN. It has already happened in at least three of the register's live

    sources.

  3. Severity CRITICAL. Retroactive movement in history is the failure that makes a rate

    uninterpretable and, unlike most modes here, it changes numbers the tracker already published.

  4. Detection FD-08 (re-anchor drift monitor) and FD-25 (vintage diff). On every run, recompute

    each stored historical value from the current source and diff it against the archived vintage; any non-zero diff is a revision event that is logged and surfaced, not absorbed. FD-11 (methodology-version changelog gate) additionally refuses to publish a series whose methodology version changed without a changelog entry. All three runnable.

  5. Mitigation Vintage archiving, and never overwriting history. Each fetch is stored immutably

    under its fetch date; the schema's methodology_version and revision fields per C11 carry the version at observation time; recomputation always produces a new vintage rather than mutating an old one. Rate is computed only within a constant methodology version, and cross-version rate requires a published bridge.

  6. Verdict KEEP WITH CAVEAT for A1, A2, C3, H1, each carrying a methodology version, a

    visible revision history, and rate computed only within version. Any S4-derived figure is labelled as re-anchoring and is never quoted from cache.

  1. Distortion A source dies, moves, or freezes, and the pipeline keeps returning the last

    value it saw, so the indicator flatlines. A frozen indicator reads as a plateau, and a plateau is the single most consequential reading this instrument can produce. This mode is not hypothetical and was not in the original mode list: it was added because verification found it empirically. Of the sources checked on 2026-07-30, four are retired, dormant, or in maintenance mode (S16 entered maintenance mode on 1 June 2026; S26 is retired and now redirects; S27 is archived; S28 appears dormant, promising monthly refreshes while its README release date reads 2025-04-25; all independently verified), and two moved host or froze a version line (S18 from lmarena to arena, S19 pinned to a v1 host; both independently verified as redirects observed on the check date). Every indicator whose feed dies is affected; the highest exposure sits on A1, A2, E1, C1, C2 and H1, since their feeds include the sources already trending dormant.

  2. Likelihood NEAR-CERTAIN. Six of the register's sources have already exhibited it, measured

    on a single day of checking.

  3. Severity CRITICAL, and it belongs among the most severe modes in this register. A dead

    source silently freezes an indicator at its last value, the frozen indicator reads as a plateau, and a plateau is exactly the reading a public commentator would be tempted to interpret as capability stagnation. The instrument would then report the opposite of a measurement failure: it would report a finding.

  4. Detection FD-01 (per-source liveness and freshness probe), run on every scheduled run, and

    it must fail loudly rather than warn. For each source it asserts three things: that the endpoint resolves, that the payload validates as the expected content type and shape, and that the source has moved within its own declared cadence, comparing observed last-change date against the cadence recorded in the source register. A source that has not changed within its declared cadence is marked stale and every indicator downstream of it is marked stale in the same run. FD-02 (content-fingerprint validator) is a required companion, because an HTTP status check alone is not a liveness detector: the US Census server behind S32 returns HTTP 200 for fabricated filenames under its downloads path, including one the verifier invented, so it soft-404s (S32, independently verified). Content must therefore be validated, not just fetched: assert an expected schema, a row-count floor, and a changed-content fingerprint before accepting a payload as live. FD-34 (redirect-and-host watcher) records the final resolved host per source and alerts on any change, since two sources have already moved host.

  5. Mitigation Never hardcode a URL in the pipeline: every source is a configuration record

    with an ID, an endpoint, a declared cadence, an expected content fingerprint, and a named fallback, so a move is a configuration edit rather than a code change. Stale sources propagate a staleness flag to every dependent indicator, and the UI renders a stale indicator visibly differently from a flat one, so a frozen series can never be mistaken for a measured plateau. Sources in the register's Tier 3 are never cited as live. The 90-day staleness flag inherited from the house methodology applies per indicator.

  6. Verdict KEEP for every indicator, with mandatory staleness propagation. Any indicator whose

    only feed is a Tier 3 source is DROPPED from the headline rather than carried on a frozen value, and appears in the instrument as an explicit unavailable state with the date its source was last confirmed live.

Detector inventory

Every detector below is runnable as described. Cadence "per run" means on each scheduled pipeline execution. "Automated" means it executes without a human and can fail the build; "manual" means it requires a person and is therefore occasional rather than guaranteed.

IDDetectorGuardsCadenceMode
FD-01Per-source liveness and freshness probe: endpoint resolves, payload validates, source has moved within its declared cadenceall sources S1 to S24; FM-27per runautomated
FD-02Content-fingerprint validator: expected schema, row-count floor, changed-content fingerprint. Required because S32 soft-404s with HTTP 200all fetched payloads; FM-27per runautomated
FD-04pass@1 to pass@k spread (D1 operationalised)D1, A1, E1; FM-02, FM-14per runautomated
FD-05Saturation half-life (A2 operationalised)A2, A1; FM-01, FM-08per runautomated
FD-06Held-out versus public delta across matched set halvesB1, B2, A1; FM-01, FM-03, FM-04per run where both halves publishautomated
FD-07Version-pin mismatch check: refuse cross-version deltasS6, S14, S5; FM-09, FM-16per runautomated
FD-08Re-anchor drift monitor: recompute stored history, alert on movementS4 and S9-derived series; FM-16, FM-26per runautomated
FD-09Submission-coverage census against S2 and S24 model listsE1, A1, H2, D4; FM-05, FM-15per runautomated
FD-10Frontier-actor census from S24 company and chip-owner dataH2; FM-25monthlyautomated
FD-11Methodology-version changelog gateA1, A2, C3, H1; FM-26per runautomated
FD-12Estimate-revision flag on S2 compute valuesG3i, G1i; FM-21per runautomated
FD-13Licence and terms watcherS13, S5, S11, S14; FM-24weeklyautomated
FD-14Scaffold-disclosure completeness checkA1, E1, C1 to C3; FM-02, FM-10per runautomated
FD-15Cost per solved task at fixed reliability (G2i operationalised)G2i; FM-12per runautomated
FD-18Human-baseline presence check on withheld setsB1, A1, E1; FM-07monthlyautomated
FD-19Variant-identity assertion: no silent configuration switch within a seriesC1 to C3, G2i, H1; FM-06per runautomated
FD-20Ceiling-proximity flag: mark resolution-exhausted evaluationsA1, A2; FM-08per runautomated
FD-21Harness-constancy assertionA1, E1, C1 to C3, F1; FM-10per runautomated
FD-22Trajectory-availability assertion (presence only, not content)C1 to C3, E1; FM-11per runautomated
FD-23Budget-disclosure assertionA1, E1, C1, C2, H1, G2i; FM-12per runautomated
FD-24Interval-presence assertion: how many observations carry any dispersion estimateD2, D3; FM-13per runautomated
FD-25Vintage diff: today's fetched history against the archived vintageall series; FM-16, FM-26per runautomated
FD-26Realism-class assertion and published task-provenance mixE1, E2, F1; FM-17per runautomated
FD-27Release-to-record lag monitorall capability indicators; FM-18per runautomated
FD-28Cadence-versus-lag ledger per sourceH1, E1, S7 and S17-derived context; FM-19per runautomated
FD-29Layer-provenance assertion: no adoption or resource source in a capability computationwhole capability layer; FM-20, FM-21per runautomated
FD-30Verification-class census per rendered viewI1, all judgment inputs; FM-22per runautomated
FD-31Significant-figure gate and mandatory interval at render timeheadline Position and Rate; FM-23per buildautomated
FD-32Cross-source concordance check beyond a published toleranceA1, B2, G2i; FM-24per run where two sources existautomated
FD-33Provider-concentration census over headline inputswhole pipeline; FM-25per runautomated
FD-34Redirect-and-host watcher: final resolved host per sourceS18, S19, all sources; FM-27per runautomated
FD-35Canary-recall probe on models with API accessA1, B2; FM-04opportunisticmanual
FD-36Sampled trajectory audit for assistance boundaryC1 to C3; FM-11quarterly, sampledmanual
FD-37Withheld-set governance review and disclosure refreshA1, B2, E1; FM-07, FM-24quarterlymanual

Failure modes with NO runnable detector today. This list is the honest weakness of the instrument and is published rather than buried.

  • FM-04 (training on evaluation data): no detector is possible from outside a laboratory. FD-06

    detects a signature, FD-35 is manual and depends on API access, and neither establishes the fact.

  • FM-11 (hidden human assistance): FD-22 asserts trajectory presence, not trajectory content.

    Adjudicating assistance requires FD-36, which is manual and sampled, so most observations are never audited.

  • FM-13 (best-case versus median) and FM-22 (expert judgment as data): under-detected by

    construction. D2 and D3 have no live source per gap G1, and D4 has none per gap G2, so the reliability and regression checks that would catch these modes are self-computed or absent. This is stated plainly rather than papered over: the reliability dimension of a reliability-weighted target is currently the least instrumented dimension in the tracker.

  • FM-18 (capability overhang): the internal-to-release lag is unobservable. FD-27 bounds only the

    tracker's own latency.

  • FM-24 (motivated manipulation): FD-13 and FD-32 detect terms changes and discrepancies. Neither

    detects manipulation, and no detector for it is claimed.

  • FM-01 (Goodhart) and FM-02 (gaming): the signature is detectable, intent is not. No detector

    distinguishes a legitimate capability gain from a targeted one.

Plateau, saturation and regression detectors

A tracker that cannot tell "progress stopped" from "our ruler melted" is worthless in exactly the moment it matters most, because the two readings recommend opposite actions and look identical on a chart. Four named detectors carry that discrimination. No plateau claim is published unless all four have reported, and their combined verdict, not any single one, determines whether the instrument states stagnation.

PD-1. Saturation-versus-stagnation discriminator (A2 plus FD-05 and FD-20). A2 measures how quickly evaluations are being saturated. The discrimination is directional. If the frontier score is flat while A2's half-life is short and FD-20 flags the evaluation as inside its ceiling proximity band, the flatness is the ruler: resolution has been exhausted and the evaluation must be retired from A1's basket with a methodology version bump. If the frontier score is flat while A2's half-life is lengthening and FD-20 reports ample headroom, the flatness is in the world: new evaluations are not being saturated and old ones are not being finished. Both branches are computable from S1's dump, which was observed updated on 2026-07-30 (S1, independently verified), so the discriminator runs on every scheduled run rather than on request.

PD-2. Regression detector by vintage diff (D4 plus FD-25, self-computed per gap G2). Gap G2 establishes that no live eval or leaderboard tracks generation-over-generation regression, and recommends self-computation by diffing successive S1 dumps as cheap and viable. D4 does exactly that: for each model family, diff the current vintage against the prior vintage and count evaluations where a later generation scores below an earlier one, with benchmark version and methodology version held constant per FD-07 and FD-11. The discrimination comes from where the diff lands. A regression concentrated in one family across many evaluations is a capability finding about that family. A regression appearing simultaneously across unrelated families on the same evaluation is a measurement event: a grader change, a harness change, or a set revision. A regression that appears in the tracker's own stored history without any new observation is neither, and is a source revision caught by FD-08. Note the honest limit: D4 has no live source, so its verification class is inferred from the tracker's own diffs, and it cannot see a model that was never evaluated anywhere per FM-15.

PD-3. Source-freshness plateau veto (FD-01 plus FD-02 and FD-34). Before any flat series is interpreted, the freshness probe must have confirmed that every feed behind it moved within its own declared cadence. A source that has not moved within its cadence marks the indicator stale, and a stale indicator is rendered as unavailable rather than flat, which removes the possibility of reading a dead feed as a plateau. This detector exists because six sources in the register already exhibited mortality on one day of checking (S16, S26, S27, S28 retired, dormant, or in maintenance mode; S18 and S19 moved host or froze a version line; all independently verified). FD-02 is inseparable from it: an HTTP 200 is not evidence of liveness, since the S32 host returns 200 for fabricated filenames (S32, independently verified), so the payload's schema, row count and fingerprint must all be checked before the series is treated as fresh. The discrimination is therefore blunt and reliable: if the feed did not move, the tracker says nothing about the world.

PD-4. Rate-integrity check on the Rate half of the headline (C3 plus FD-07, FD-08 and FD-11). Position can plateau for measurement reasons; Rate can plateau for methodology reasons. Before a slowdown in C3 (Horizon Doubling Time) is published, the check asserts that the series contains no version boundary, no re-anchoring event, and no methodology change over the window being fitted, and that the underlying feed's cadence is satisfied. S5 is self-described as operating at limited capacity with irregular updates (S5, independently verified), so a flattening C3 is first a cadence question and only then a capability question. If the window is clean and the feed is fresh, a lengthening doubling time is reported as a rate finding with its uncertainty band. If any assertion fails, the rate is withheld and the failure named.

Together these four give the instrument the property the exit condition requires: it can report stagnation, and it can report that it is unable to report, and it cannot confuse the two. The residual exposure is stated in the detector inventory above, and the largest piece of it is that the reliability dimension (D2, D3 per gap G1) is the least instrumented part of a target defined as reliability-weighted.

#

10.1 The four layers, made concrete

Directories in one git repository. The separation is a filesystem fact, not an intention.

/registry/            sources.yaml            source IDs, feed URLs, licence, cadence, class
/data/                                        LAYER 1: DATA. Observations only.
  /snapshots/<source_id>/<retrieved_at>/      immutable fetched artefacts plus .sha256
  /observations/<indicator_id>.jsonl          append-only observation records
  /vintages/<indicator_id>/<vintage_date>.jsonl   archived upstream vintages
/methodology/                                 LAYER 2: METHODOLOGY. Rules, versioned.
  methodology.yaml                            version, definitions, thresholds, weights, anchors
  /versions/<semver>.yaml                     every superseded methodology version, retained
  CHANGELOG.md
/interpretation/                              LAYER 3: INTERPRETATION. Removable.
  /<yyyy-qq>/commentary.md                    dated, attributed, no new numbers
/forecasting/                                 LAYER 4: FORECASTING. Quarantined.
  /<yyyy-qq>/notes.md                         labelled, sourced separately
/pipeline/                                    fetch, validate, compute, render scripts
/site/                                        generated static output

Enforcement of the no-editorial rule in the data layer: notes is the only free-text field in an observation record, and validation rejects it if it matches a forbidden-token list held in /methodology/forbidden_tokens.txt. That list covers predictive language (future-tense modals, any bare year expressed as a target, any percentage attached to a future state) and evaluative language (adjectives grading a reading as good, bad, surprising, or unremarkable). The list lives outside this document so the tokens themselves do not appear in the data or methodology layers, and it is versioned with the methodology. Enforcement of removability: the build runs a test that renders the site with /interpretation/ and /forecasting/ deleted and asserts that every number in the output is byte-identical to the full build. If a number changes, an editorial layer was load-bearing, and the build fails.

10.2 The observation schema

One record is one observation of one indicator at one date from one source. Derived indicators store their inputs by observation id.

FieldTypeReqDescription
observation_idstring (ULID)yesStable, immutable, never reused. Primary key
record_statusenumyesobservation or schema_illustration. Illustrations are excluded from every computation
indicator_idstringyesSection 4 ID. The MVP set is A1, A2, C1, C2, C3, D2, D4, G1i, G3i; every other section 4 ID is permitted for a deferred indicator's gap record
system_idstringyesModel or system identity as the source names it
system_checkpointstringyesCheckpoint, snapshot date, or API version string. unspecified-by-source where the source gives none
developerstringyesDeveloper organisation as named by the source, before any attribution normalisation
developer_normalisedstringnoOrganisation after the H2 attribution mapping, with the mapping version
valuenumber or nullyesThe observed value. Null only where verification_class is missing
unitstringyesUnit from the section 4 field 5 for that indicator
uncertainty_typeenumyesinterval, stderr, dispersion, none-published
uncertainty_lownumbernoRequired when uncertainty_type is interval
uncertainty_highnumbernoRequired when uncertainty_type is interval
observation_dateISO dateyesThe date the observation refers to, not the fetch date
source_idstringyesSOURCES.md ID. The only permitted way to name a source
source_urlstringyesURL as fetched, captured at fetch time from /registry/sources.yaml
retrieved_atISO datetimeyesUTC timestamp of the successful fetch
content_sha256stringyesSHA-256 of the fetched artefact, matching a file under /data/snapshots/
verification_classenumyesindependently-verified, self-reported, estimated, expert-judgment, inferred, missing
evidence_tierinteger 1 to 6yesEvidence tier from the six-level source hierarchy
auditability_classenumyesIND, LAB, AGG, GOV, from the register
benchmark_idstringyesBenchmark or suite identity. not-applicable for indicators with no benchmark
benchmark_versionstringyesVersion or split as the operator names it. Never inferred
scaffold_idstringyesHarness or scaffold identity. unspecified-by-source where absent
scaffold_versionstringyesAs above
tool_accessenumyesnone, retrieval, code-execution, browser, full-agentic, unspecified-by-source
sampling_methodstringyesTemperature, attempt count, decoding, or unspecified-by-source
autonomy_classenumyesunassisted, harness-assisted, human-in-loop, undeclared. Required by the FM-11 mitigation in section 8. undeclared is the value used when the source does not state the assistance boundary, and it is never silently mapped to unassisted. Validation refuses any observation carrying undeclared as an input to C1, C2 or C3, so an undeclared assistance boundary excludes the observation from the horizon dimension rather than degrading it
grader_protocol_versionstringrequiredIdentifier and version of the grading or judging protocol that produced the value, for example an LLM-judge rubric version or a human inter-rater protocol. The literal not-applicable where the benchmark is scored programmatically. In the comparability key because a change of grader silently changes the measured quantity
environment_classenumconditionalsimulated, structured-physical, unstructured-physical, unspecified-by-source. Only simulated is populated by any verified source today, so an F1 record in any other class is a validation error until a physical suite exists. Required when indicator_id is F1, per check 4 in 10.6; optional and defaulting to absent for every other indicator. It exists in the schema before F1 exists as an indicator so that the gap G4 slot cannot later be filled with a structured-environment benchmark unnoticed
estimate_revision_countintegeryesCount of times the UPSTREAM estimate for this observation has been revised across the vintage archive, 0 at the first record. Distinct from revision_number, which counts tracker-side record revisions: an extraction error the tracker fixes increments revision_number and not this field, while an upstream re-estimate increments both. Asserted non-decreasing by check 5 in 10.6. This is the revision counter the G3i construction requires
compute_budget_flopnumbernoWhere known. Absent is not zero
compute_budget_sourcestringnoSource ID for the compute figure when present
compute_budget_bucketenumyesBucketed training compute, derived deterministically from compute_budget_flop by the rule below, never hand-entered. One of unknown, lt-1e23, 1e23-1e24, 1e24-1e25, 1e25-1e26, gte-1e26. It is an element of the comparability key in 10.3, so it is required on every record, including records for which the FLOP figure is absent
licence_statusenumyescc-by, cc-by-sa, mit, apache-2.0, proprietary, not-stated
methodology_versionsemver stringyesMethodology version under which the value was computed
comparability_groupstringyesDeterministic hash of the comparability key (section 10.3)
superseded_bystring or nullyesobservation_id of the record that replaces this one, else null
revision_numberintegeryes0 for the first record of an observation, incrementing thereafter
revision_reasonstring or nullyesupstream-revision, extraction-error, methodology-change, licence-withdrawal, or null at revision 0
vintage_dateISO dateyesThe date of the upstream vintage this value was read from
staleness_flagbooleanyesTrue when retrieved_at exceeds the source's declared cadence at computation time
notesstringyesNon-editorial. Factual provenance remarks only. Validated against the forbidden-token list. Empty string permitted

10.3 The comparability rule, made mechanical

The comparability key is the tuple (indicator_id, benchmark_id, benchmark_version, scaffold_id, scaffold_version, tool_access, sampling_method, compute_budget_bucket, methodology_version, system_checkpoint, grader_protocol_version). grader_protocol_version is the schema field of that name declared above; it is in the key because a change of grader silently changes the quantity. comparability_group is a deterministic hash of that tuple. Observations differing in any element sit in different groups.

**The compute_budget_bucket rule, since the key depends on it and a hand-entered bucket would be a judgment inside a hash.** The bucket is computed by the pipeline from compute_budget_flop, in half-open decade bands [lower, upper): below 1e23 gives lt-1e23; 1e23 up to but excluding 1e24 gives 1e23-1e24; and so on to gte-1e26 for 1e26 and above. Where compute_budget_flop is absent the bucket is unknown, which is a distinct bucket and not a wildcard, so an observation with no compute figure is never grouped with one that has a figure. Decade bands are deliberately coarse, so that the S2 estimate error bars rarely straddle a boundary. Where an upstream re-estimate does move a figure across a boundary, the bucket changes, which changes the comparability key, which by rule 1 below would split the series; that case is therefore handled as a comparability event: the new record opens a new group, estimate_revision_count increments, and a splice object is required before the two groups may be shown as one line. A re-estimate can never silently regroup a series.

Enforcement:

  1. The compute step refuses to aggregate across groups. Every series query is scoped to a single

    comparability_group. A request spanning more than one group raises and the build fails. There is no flag to override it.

  2. Splices are explicit objects, not silent joins. A splice record in /methodology/ names the

    two groups, the overlapping observation used, the ratio, and the date. Section 4 already requires this for A1 basket rotation, and it retains the unspliced series. The pipeline stores both and the default published series is the unspliced one, with the spliced series available as an alternate view.

  3. **unspecified-by-source is a distinct value, not a wildcard.** Two observations both carrying

    unspecified-by-source for scaffold are in the same group only if every other element matches, and the resulting series carries an unpinned-scaffold warning flag.

  4. Dashboard behaviour at a boundary. The line breaks. A visible vertical rule marks the boundary,

    labelled with what changed and the date. No interpolation, no bridging segment, no dotted continuation implying equivalence. Where a splice exists, a toggle shows the spliced version with the splice factor printed beside it. Tooltips on either side of the boundary state that the values are not comparable. This is the same treatment section 4 specifies for E1's methodology-version boundaries.

10.4 Revision and correction procedure

Append-only. Nothing is ever overwritten. A correction is a new record.

  1. The new record carries revision_number incremented, a revision_reason, and the same

    indicator_id, observation_date and comparability_group as the record it replaces.

  2. The superseded record is not deleted or edited except for one field: its superseded_by is set to

    the new observation_id. That is the only permitted mutation in the data layer, and the validation script asserts that no other field of an existing record has changed by comparing against the previous commit.

  3. The prior vintage is archived: the fetched artefact already sits under

    /data/snapshots/<source_id>/<retrieved_at>/ with its hash, and the value as it stood is written to /data/vintages/<indicator_id>/<vintage_date>.jsonl. A vintage file is immutable once written.

  4. Published views default to the latest vintage. A vintage selector lets a reader reproduce any

    previously published headline exactly, which is the practical test of the archive.

  5. The change log records: date, indicator, observation id, old value, new value, revision reason,

    source vintage dates on both sides, and whether the change moved the headline. Headline-moving corrections additionally get an interpretation-layer note.

The S4 retroactive-drift case, specifically. S4's ECI is a latent-trait fit that is re-anchored as benchmarks saturate, so historical values move retroactively and any cached figure silently drifts (SOURCES.md confounder 1, independently verified). Three rules follow.

  • No fitted S4 index value enters any indicator. Section 4 already restricts S4 to raw scores for

    A1, D4 and H2. The pipeline enforces this by refusing to load the fitted-index columns at all.

  • **Where S4 raw scores are used, each is stored with its own vintage_date.** A re-anchoring that

    changes a historical raw score produces a new record with revision_reason: upstream-revision, and the previous vintage stays readable. An upstream re-anchoring can never overwrite a stored vintage, because the writer has no update path.

  • Re-anchoring is reported, not absorbed. D4's construction already separates revisions from

    regressions and publishes a revision count. The same separation applies to every indicator: a period's report states how many stored values changed because upstream changed, distinct from how many changed because the world changed. A quarter in which upstream revisions exceed genuine movements is itself a finding about the instrument.

10.5 Benchmark-version and methodology-version handling

Benchmark version is pinned at fetch: the expected version string per benchmark lives in /registry/sources.yaml and the validator asserts that the fetched artefact reports the expected version. An unexpected version is a hard failure, not a warning, and the affected observations land in a new comparability_group on the maintainer's review. Section 4's requirement that every observation pin benchmark version is thereby a schema constraint rather than a habit.

Methodology version is semantic, held in /methodology/methodology.yaml, and stamped onto every observation at computation time. Superseded versions are retained under /methodology/versions/.

When a benchmark version changes: the series closes at the boundary, a new series opens, and the headline recomputes on the new group only after the minimum-history rule in section 9.4 is met for that group, or after a splice object exists. In the interim the affected dimension shows its last comparable value with a boundary badge and is excluded from the floor. Excluding a dimension raises the floor, so the coverage badge states the exclusion explicitly to prevent the rise being read as progress.

When the methodology version changes: the entire published history is recomputed under the new version and published beside the old, both labelled with their version. The headline never mixes methodology versions within one series. A breaking methodology change (section 11.2) blocks publication until the recomputation and the side-by-side comparison exist.

10.6 Data-validation checks

Runnable as pipeline/validate.py, exit non-zero on any failure, wired as a publication gate.

Exemption rule for schema illustrations, stated as part of the specification rather than left as an oversight. A record carrying record_status: schema_illustration documents the schema; it measures nothing. Such records are exempt from check 1 (content validation of a fetched artefact), from the content_sha256 clause of check 10 (referential integrity), from check 8 (cross-source disagreement), and from the range checks in check 3 where the illustrative value is outside a bound. They are not exempt from check 2, so every type, enum, required field and conditional requirement still applies, and that is the entire point of the record. Three further rules make the exemption safe rather than a hole: an illustration must carry verification_class: missing and a notes string declaring that it is an illustration and must not be cited; illustrations live in /methodology/examples/ and the validator **fails the build if any record with record_status: schema_illustration appears anywhere under /data/observations/**; and the compute step refuses illustration records at load, so they are barred from every computation, every series, and every published figure. An illustration whose hash resolves to no archived artefact is therefore expected, not a validation failure, because it was never permitted to be a data point.

  1. Content validation, not status validation. HTTP 200 is not evidence of a valid artefact. The

    register establishes that the US Census server returns HTTP 200 for fabricated filenames under its downloads path, including one the verifier invented, so it soft-404s (S32, independently verified). Every fetch is therefore validated by: expected content type; a non-trivial minimum byte size; successful parse in the expected format; presence of an expected column or key set; a plausible row count against the previous fetch; and a SHA-256 recorded on every artefact. A fetch that returns HTML where CSV was expected fails, whatever the status code.

  2. Schema validation. Every observation record validates against a JSON Schema covering types,

    enums, required fields, and the conditional requirements (uncertainty_low and uncertainty_high present when uncertainty_type is interval; value null only when verification_class is missing; revision_reason non-null when revision_number is above zero).

  3. Range checks. Per-indicator bounds from section 4 field 5: A1 in [0, 1]; A2 positive months;

    C1 and C2 positive minutes; C3 positive months with a published interval; D4 a non-negative integer with a positive denominator; G1i and G3i positive FLOP; H2 HHI in [0, 1] with an organisation count of at least one. Out-of-range fails rather than clamps.

  4. Structural constraint checks. E2 numeric fails (section 4 makes a numeric E2 invalid by

    construction). F1 requires environment_class. B2 requires a joint B1 record. Any indicator whose record carries value null must carry verification_class: missing, and per the gap rule in 9.11 a DECLARED GAP ID on an indicator does not by itself require a null value: D2 and D4 both sit inside declared gaps and both carry self-computed values, so the check tests the record, not the gap ID. autonomy_class: undeclared on an observation whose indicator_id is C1, C2 or C3 fails, per the FM-11 mitigation. BG's diagnostic record may never carry an indicator_id of B1 or B2.

  5. Monotonicity checks where appropriate. Applied only where the quantity is monotone by

    construction, never where it would launder an assumption: revision_number strictly increasing per observation; retrieved_at non-decreasing per source; cumulative snapshot count non-decreasing; estimate_revision_count non-decreasing. Capability values are explicitly not monotonicity checked, since section 4 requires every core indicator to be able to move backwards.

  6. Staleness check per source's own declared cadence. For each source, compare now against the

    latest retrieved_at with content change, using the cadence recorded in /registry/sources.yaml. Flag at one cadence period, demote to DORMANT at twice, and apply the 90-day flag to S5 as section 4 specifies. Where a cadence is UNKNOWN (S8 in the register), the default is 90 days and the default is recorded as an assumption, not as the source's cadence.

  7. Duplicate-observation check. No two non-superseded records may share

    (indicator_id, system_id, system_checkpoint, benchmark_id, benchmark_version, observation_date, comparability_group). A duplicate is a pipeline bug, not a data point.

  8. Cross-source disagreement check. Where two sources cover the same system and benchmark version

    (S1 against S4 raw scores, S1 against S13 where licensed), compute the absolute difference and fail above a materiality threshold of 2 percentage points, matching D4's threshold (specification parameter). A disagreement is recorded as a disagreement object with both source IDs and both values, published rather than resolved.

  9. Append-only integrity check. Diff /data/observations/ against the previous commit; fail if any

    field other than superseded_by changed on an existing record, or if any line was deleted.

  10. Referential integrity. Every source_id resolves to a registry row; every superseded_by

    resolves to an existing observation_id; every derived indicator's input ids resolve; every content_sha256 matches a file present under /data/snapshots/.

  11. Schema-drift check on upstream. The fetched artefact's column or key set must equal the pinned

    expected set. Additions fail as loudly as removals, because a silent addition often accompanies a silent semantic change.

  12. Licence-gate check. No observation with licence_status: proprietary is present in the

    published build unless a licence record exists in the repository. No artefact from a not-stated source is present under a public path.

  13. Layer-separation check. Forbidden-token grep over the data and methodology layers, and the

    removability test in section 10.1.

10.7 Technology stack

ChoiceJustification against the maintenance constraint
One public git repository, plain files, no databaseEvery value is diffable and reviewable in a pull request; the revision history is the audit trail; there is no server to patch, back up, or pay for. A database would put the audit trail somewhere the maintainer cannot read with git log
JSONL for observations, CSV for published series, YAML for registry and methodologyJSONL appends without rewriting; CSV is what a critic can open; YAML is what a human edits. All three diff readably
Python 3 standard library plus requests, pandas, jsonschema, PyYAML, and epochaiA small, boring dependency set. epochai is the operator's own client (pip install epochai, Airtable-backed, independently verified as offered), which removes a scraping surface for the most load-bearing source
Roughly six scripts: fetch.py, validate.py, compute.py, snapshot_diff.py, render.py, healthcheck.pyEach runnable alone from a terminal. The maintainer can reproduce any published number with one command. No orchestration layer to learn or debug
Scheduled CI (GitHub Actions or equivalent) with the fetch job weekly and the compute job monthlyFree at this scale, logs retained, failures notifiable by email. The dead-man's-switch banner is generated from the last successful run timestamp committed by CI
Static site, no client-side data fetching, generated by render.py into /site/Nothing to keep running; the page cannot break because an API is down; the whole site is a diffable artefact. Aligns with the section 12 render
Snapshot archive in-repo with Git LFS only if artefact size forces itReproducibility depends on retaining the fetched bytes; D4 depends on it absolutely

Deliberately not used. A hosted database (Postgres, Supabase): adds an availability dependency and hides the audit trail. An orchestrator (Airflow, Dagster, dbt): more maintenance surface than the nine indicators justify. A JavaScript framework: the register already documents three sources that are unusable because they are JavaScript shells, and building the tracker as one would be the same mistake. Headless-browser scraping: the leading cause of silent breakage, and the reason B1, E1, F1 and I1 are deferred rather than hacked in. Any LLM inside the pipeline itself, meaning extraction, parsing, grading or commentary generation: a non-deterministic extraction step would make the archive unreproducible. The distinction that matters, since the MVP does call model endpoints: in the D2 probe the model is the object being measured, its raw per-run outputs are scored by a deterministic auto-grader and stored as observations, and no model output is ever used to produce, transform or describe another indicator's value. Containers: unnecessary for a stdlib-plus-five-packages job.

10.8 A complete sample record

The record below is a schema illustration. Its value is not a measurement and its system_id is not a real model. It exists to demonstrate that every required field can be populated coherently. record_status: schema_illustration excludes it from every computation, verification_class: missing marks the value as unsourced, and the notes field states this. The source_url is the S1 download page as recorded in SOURCES.md. Its content_sha256 is a placeholder that resolves to no archived artefact, which is permitted by the schema-illustration exemption stated at the head of 10.6 and would otherwise fail check 10; the record lives under /methodology/examples/, never under /data/observations/. autonomy_class is undeclared here, which is legal on an A1 record and would be rejected on a C1, C2 or C3 record by check 4.

{
  "observation_id": "01J0ZQ7X9K8M3N4P5R6S7T8V9W",
  "record_status": "schema_illustration",
  "indicator_id": "A1",
  "system_id": "ILLUSTRATIVE-MODEL-A",
  "system_checkpoint": "illustrative-checkpoint-0000",
  "developer": "ILLUSTRATIVE-DEVELOPER",
  "developer_normalised": "ILLUSTRATIVE-DEVELOPER",
  "value": 0.42,
  "unit": "dimensionless ratio on [0,1]",
  "uncertainty_type": "dispersion",
  "uncertainty_low": 0.31,
  "uncertainty_high": 0.55,
  "observation_date": "2026-07-30",
  "source_id": "S1",
  "source_url": "https://epoch.ai/benchmarks/use-this-data",
  "retrieved_at": "2026-07-30T09:00:00Z",
  "content_sha256": "0000000000000000000000000000000000000000000000000000000000000000",
  "verification_class": "missing",
  "evidence_tier": 6,
  "auditability_class": "AGG",
  "benchmark_id": "ILLUSTRATIVE-BENCHMARK-BASKET",
  "benchmark_version": "illustrative-v0",
  "scaffold_id": "unspecified-by-source",
  "scaffold_version": "unspecified-by-source",
  "tool_access": "unspecified-by-source",
  "sampling_method": "unspecified-by-source",
  "autonomy_class": "undeclared",
  "grader_protocol_version": "not-applicable",
  "environment_class": "unspecified-by-source",
  "estimate_revision_count": 0,
  "compute_budget_flop": null,
  "compute_budget_source": null,
  "compute_budget_bucket": "unknown",
  "licence_status": "cc-by",
  "methodology_version": "0.1.0",
  "comparability_group": "cg_illustrative_0000000000000000",
  "superseded_by": null,
  "revision_number": 0,
  "revision_reason": null,
  "vintage_date": "2026-07-30",
  "staleness_flag": false,
  "notes": "Schema illustration only. The value 0.42 is not a measurement of any system and must not be cited. record_status is schema_illustration and verification_class is missing, so this record is excluded from all computation and from every published series. content_sha256 is a placeholder and resolves to no archived artefact."
}

#

11.1 Who decides what

One maintainer. No committee exists and the spec does not pretend otherwise. The governance mechanism is therefore not a body; it is a set of constraints on one person that a stranger can check without the maintainer's cooperation.

DecisionWhoConstraint that binds them
Indicator admission or removalMaintainerMust satisfy the section 9.7 tests, be recorded in the change log with the test results, and be committed before the first observation of that indicator is published
A1 basket rotationMaintainerDraws from a publicly posted admission queue ranked on headroom then auditability class (section 4). The queue is committed before the rotation, so the choice cannot be made after seeing the effect
Thresholds, weights, anchorsMaintainerPre-registered per section 11.4. Cannot be set or changed in the same release as the data that tests them
Stage promotion or demotion (9.10)MaintainerTriggers are pre-registered in /methodology/methodology.yaml before Stage 0 publishes, so a stage cannot be promoted to suit a reading. A promotion that publishes a dimension's first position must satisfy that dimension's history rule in 9.4; a demotion publishes the condition that caused it
Registering BG as the minimum admissible K9 observation (9.11)MaintainerDone once, in writing, before any BG value exists, with both readings of K9 published side by side and a change-log entry naming 9.11. Superseded automatically if a matched-split source becomes machine-readable
Methodology version bumpMaintainerSemantic rules in section 11.2, with recomputation and side-by-side publication for a breaking change
CorrectionsMaintainerAppend-only per section 10.4. A correction is never a deletion, and a submitted correction gets a public disposition per section 11.6
Interpretation commentaryMaintainer, named and datedConfined to /interpretation/, must contain no numbers the data layer has not already published, and must survive deletion without changing any number
Freezing or retiring the trackerMaintainerConditions pre-stated in section 11.7

The single-maintainer weakness that this does not fix: the maintainer can still choose which indicators to build first, and that choice shapes what the instrument can see. It is mitigated only by the deferred-indicator table in section 9.2, which states the admission condition for every excluded indicator in advance, so a critic can ask why a met condition has not produced an admission.

11.2 Methodology versioning

Semantic versioning on the methodology, independent of any code version, held in /methodology/methodology.yaml.

  • MAJOR increments on a change that alters a published historical value or the meaning of a

    series. Examples: changing the A1 normalisation; changing the floor construction; changing the weakest-link dimension set; changing D4's materiality threshold; promoting a fallback source to primary; changing a comparability key element.

  • MINOR increments on an additive change that leaves existing values intact. Examples: admitting a

    new indicator; adding a corroborating source; adding a new view or a new uncertainty component reported alongside the existing band.

  • PATCH increments on a change that alters no value and no meaning. Examples: fixing a typo in a

    definition; correcting a field description; refactoring a script with a numerically identical output, demonstrated by a byte-identical rebuild.

A breaking change is defined by effect, not by intention: any change after which a rebuild from the same archived artefacts produces a different published number is MAJOR, whatever the author believed they were doing. The build detects this automatically by rebuilding on the previous methodology version and diffing the outputs, so the version bump is not a judgment call. A MAJOR bump blocks publication until (a) the full history is recomputed under the new version, (b) both versions are published side by side with their version labels, and (c) a change-log entry states what moved and why.

11.3 Change log

One file, /methodology/CHANGELOG.md, append-only, newest first. Every entry carries all nine fields:

### <ISO date> | <methodology version> | <MAJOR|MINOR|PATCH>
- change: one sentence, what changed
- scope: indicator ids and source ids affected
- reason: why, in the author's own words
- effect on published values: none | listed diffs | full recomputation
- effect on the headline: none | direction and magnitude
- pre-registration reference: commit hash where the parameter was registered, or "not applicable"
- author: name
- reviewer: name, or "none"
- correction origin: internal | external submission id | upstream revision

A release with no change-log entry fails the publication gate. The change log is the only permitted account of why a number differs from the number a reader saw last quarter.

11.4 Pre-registration discipline

The rule: every parameter that can move the headline is published before the data that tests it. That covers thresholds, weights, floor construction, basket admission and retirement rules, the admission queue order, materiality thresholds, comparability key elements, staleness limits, and the uncertainty band construction.

Mechanism, checkable after the fact by a stranger:

  1. Parameters live in /methodology/methodology.yaml, in git. A parameter's registration date is the

    commit date of the commit that introduced it, and git log -p on that file shows every value the parameter has ever held.

  2. Every observation stamps methodology_version. Comparing an observation's retrieved_at against

    the commit date of the methodology version it names shows whether the rule predated the data. The validator runs this comparison and fails on any observation whose methodology version was committed after the observation's retrieved_at, which makes retrofitting a parameter a build failure rather than a discovered embarrassment.

  3. Commits to /methodology/ are signed, and the signature plus the commit timestamp is the audit

    record. Where signing is unavailable, the alternative is a public timestamp: the commit hash is posted to an append-only public location (an issue thread, a mailing-list post, or a dated public note) on the day of the commit, so the ordering is attested outside the maintainer's own repository.

  4. The publication gate refuses a release in which a headline-affecting parameter changed in the same

    commit as new observations. Parameter changes and data updates are separate commits, always.

11.5 Conflict of interest

Declaration requirement. A standing /CONFLICTS.md file, dated and revised on change, declaring: every financial relationship with any organisation named in the data, including consulting, speaking fees, equity, grants, and paid platform access; every free or discounted API, licence, or data grant received from a source operator; every institutional affiliation; and every commercial interest in AI training, advisory, or commentary work.

The maintainer's own exposure, named. The maintainer writes and speaks publicly about AI. That creates a direct incentive for the tracker to agree with his published commentary, and the incentive runs in both directions: a tracker that contradicts a paid talk is inconvenient, and a tracker that confirms one is a marketing asset. This is the sharpest conflict in the project and it is structural, not hypothetical.

The mechanisms that constrain it, in order of strength:

  1. Pre-registration is the primary constraint. A parameter committed and publicly timestamped

    before the data cannot be tuned to reach a preferred reading. The validator's methodology-date check in section 11.4 makes tuning a build failure. This is the only mechanism that binds without requiring the maintainer's honesty at the moment of temptation.

  2. The interpretation layer is quarantined and removable. Commentary is attributable, dated, and

    contains no numbers the data layer has not published. A reader who distrusts the maintainer's reading can delete /interpretation/ and get the identical numbers, and the build test in section 10.1 proves that is true.

  3. Mandatory publication of disagreeing evidence. The disagreement objects from check 8 in

    section 10.6, the D4 regression count, the A2 half-life, the BG ratio and its unobserved state, the D2 rerun variance with its probe parameters, the kill-condition state table from 9.11, and the coverage badge are all published unconditionally. The instrument's plateau and regression detectors are not optional views the maintainer can decline to render.

  4. Commentary citation discipline. Any public writing or speaking by the maintainer that cites the

    tracker cites a specific methodology version and observation date, so a reader can rebuild the figure quoted. A quoted figure that cannot be rebuilt is a correction-worthy error.

  5. Disclosure at the point of use. Where an indicator depends on a source whose operator has a

    relationship with the maintainer, the dependency is disclosed on the indicator's own page, not only in /CONFLICTS.md.

What none of this fixes: the maintainer chooses what to build and what to write about. The honest statement is that pre-registration constrains the numbers and nothing constrains the choice of subject.

11.6 Independent review path

What a critic can do without permission. Clone the repository; read every observation, its source id, its hash, and its methodology version; fetch the archived artefact and verify the hash; rerun compute.py and confirm the published series byte for byte; rebuild any previously published headline from its vintage; and inspect git log -p on /methodology/ to see when every parameter was set. Reproducibility is the review path, and it does not depend on the maintainer answering.

How a correction is submitted. A public issue or a pull request against the repository, using a template that requires: the observation id or the published figure disputed; the claimed correct value or the claimed error in the rule; the source id and URL supporting the claim; and the submitter's conflict declaration. Email submissions are accepted and are transcribed into a public issue by the maintainer, so no correction exists only in private correspondence.

The commitment. Every submission receives a public disposition: accepted, accepted in part, rejected with a stated reason, or open pending upstream confirmation. Target acknowledgement within 14 days and disposition within 60 days (specification parameters, and the honest constraint is one person's hours). An accepted correction produces an append-only revision per section 10.4 and a change-log entry naming the external submission. A rejected correction stays visible with its reason, because a public record of rejected criticism is part of what makes the accepted corrections credible. Corrections that moved a published headline are additionally listed on a standing corrections page, so the count of the maintainer's own errors is public and cumulative.

11.7 Retirement conditions for the whole tracker

Stated in advance, so the decision is not made under pressure by a person who has become attached to the artefact. Two end states: FROZEN, where the site remains as a dated archive with a banner and no further updates, and RETIRED, where the site is replaced by a notice and the repository is archived read-only. In both cases the data, the methodology history, and the change log stay published.

Freeze when any of these holds:

  1. Maintainer capacity falls below the steady-state estimate for two consecutive quarters. The

    honest act is a dated freeze, not a slowly rotting dashboard. The dead-man's-switch banner makes this visible before the decision is taken.

  2. The core dimension set collapses. If dimensions A and C cannot both be published inside their

    staleness windows, the floor covers one dimension and is no longer a weakest-link floor. Section 9.3 already caps what a two-dimension floor may claim; a one-dimension floor may claim nothing.

  3. S5 is dead and no horizon fit replaces it. Dimension C rests on a single source with no verified

    fallback (section 4). Its death removes the extended-horizon clause of the operational target, which is constitutive, not supporting.

Retire when any of these holds:

  1. The instrument can no longer register the readings it exists to register. If the plateau,

    saturation and regression detectors are all dormant at once, the tracker cannot show stagnation, and a tracker that can only show progress is worse than no tracker.

  2. A better-resourced instrument supersedes it. If an independent body publishes a

    version-pinned, openly licensed, uncertainty-carrying capability profile over the same axes, maintaining a one-person duplicate is vanity. The honest act is to retire and point at it.

  3. Licence or legal position becomes untenable. A source withdrawal that removes a constitutive

    dimension with no substitute, or an unresolved licence dispute over a load-bearing feed.

  4. The maintainer's conflicts become unmanageable. If a commercial relationship arises that the

    mechanisms in section 11.5 cannot constrain, and no independent party will take over the methodology, retirement is preferable to a compromised instrument carrying a credibility it no longer earns.

Procedure on either decision: a dated notice on the site stating which condition was met; a final change-log entry; the archive left readable with every vintage intact; and an explicit statement that values are historical and must not be cited as current. No silent abandonment, which is the most common way a public tracker ends and the only ending this document rules out in advance.

#

This section comes last because the visual form is downstream of the measurement, and designing it first is how trackers end up with a beautiful number nobody can defend. Everything below is constrained by decisions fixed in sections 1 to 11 and adds no new quantity, threshold, or claim. Out of scope, deliberately: naming, logo, colour palette, typographic identity, and branding. Those are chosen after the instrument is settled and none of them is load bearing for credibility.

12.1 The headline block

One bounded visual unit holds every element needed to read the headline, rendered as a single card at a fixed aspect, so the card and not the page is the smallest shareable piece. Contents in reading order:

  1. State flag. A chip carrying the word (CLEAN, REGRESSION, SATURATION,

    INSTRUMENT-FAILURE, STALE), a distinct shape, and one clause on what the flag governs. The word is always in text. Precedence follows section 6 Method D field 1, which selects only which name LEADS: the flag is a primary state plus a set of co-occurring conditions, and the card renders the full condition set beside the primary name, never only the leading name. Suppressing a condition in the rendering would defeat the mechanism that stops an instrument failure hiding the staleness that caused it. Beneath the chip sits the quantity that fired it, with its indicator ID.

  2. Position. A fraction in native counting units, illustrated here as 2 of 2 thresholds countable, 2 met (a format

    placeholder, not an observation; the registered set is four to six thresholds and UNKNOWN ones are excluded from the numerator and the reported denominator alike, per Method D's state table, so the countable denominator is smaller than the registered set whenever coverage is incomplete), with [2 to 2] adjacent at the same type size and weight, and 2 unknown beside it. Never a percentage, never a decimal, never a share of AGI.

  3. The floor sentence. One sentence naming the weakest core dimension and its component, for

    example Weakest core dimension: C, 1 of 2 met, k=1 of 4 core, with k in the same character run as the floor value per Method D field 15, never in a separable badge, and with any out-of-scope dimension named with its gate and its reason. Where the component is zero it states that no claim of general advance is made. Text, not a badge, because a badge is easier to ignore.

  4. Rate. Horizon doubling time: 7 months [4 to 19], n = 7, 24-month window, with the second

    threshold's regression underneath when the two disagree. The point estimate is rounded to the precision the interval supports and no further. Interval, observation count, and window length sit in the same line of type as the number, so a crop that keeps the number keeps them.

  5. Provenance strip. Release date, methodology version, count of stale sources, and the three

    reader-facing clauses required by section 6 Method D field 15.

Deliberately not shown: any composite score, any percentage complete, any arrow or trend glyph, any sparkline, any date, any probability, any figure from the interpretation or forecasting layers, and the resource panel, which sits adjacent and arithmetically separate.

The screenshot problem. The headline gets screenshotted and pasted without context, so the uncertainty cannot live in a caption, a footnote, a tooltip, or a second row a crop can remove. The interval is part of the numeral's own typographic run at the same size and weight as the point estimate, and the flag sits above the numbers rather than below, because casual crops take the top of a card. A shaded band from P_low to P_high renders behind the position fraction, inside the number's bounding box, so removing it means cropping the number. The social preview image is a separate surface from the card: per Method D field 15 it carries the canonical atomic citation string burned into the image pixels as its ONLY text, because that method permits no other payload there. The card is what a visitor sees on the page, and it renders the same string in full beneath itself, so the uncontrolled case (a manual screenshot of the card) still carries the qualifiers a link share carries. The card and the preview differ deliberately, and neither can be reduced to a bare number. Where MVP figures are bounds rather than estimates, per section 9.3, the card says upper bound or lower bound in the same run as the value.

12.2 The clock, restated as a UI decision

A clock face and a single needle are rejected. A needle has one degree of freedom and nowhere to render an interval, so uncertainty must be drawn outside the dial and a crop removes it. A dial also implies a bounded distance to a defined endpoint, and this specification asserts no agreed definition of AGI, so that endpoint does not exist. A dial's position cannot be recomputed from published inputs either, which fails the test in 12.8. The clock's one transferable property, a fixed revision cadence with an attributable owner, is carried by the release header and the change log.

12.3 The four layers as visible zones

The four layers of section 3.4 render as four zones in a fixed vertical order, each with a persistent label and a distinct treatment.

  • Data. Neutral field, tabular type, every figure hyperlinked to its indicator page.
  • Methodology. Same field, monospaced parameter values, each showing its version and registration

    date.

  • Interpretation. Inset panel with a visible left rule, a named author, a date, an epistemic tag,

    and no numeral the data layer has not already published, which the section 10.1 build test enforces.

  • Forecasting. Separately headed, distinct background, its own standing notice, reached by an

    explicit link rather than rendered inline with the headline.

A navbar toggle, Data and methodology only, removes the interpretation and forecasting zones, and its state is reflected in the URL so it can be linked and cited. With the toggle off the page is still complete: headline block, coverage badge, indicator panel, rate view, provenance strip, downloads, and change log all present and readable, with no dangling reference to removed prose. That the stripped page reads as a finished document is the test of whether the separation is real, and it pairs with the byte-identical-numbers assertion in section 10.1.

12.4 The indicator panel

Twenty-one indicators become a wall of numbers at equal weight, so the panel is grouped and graded.

  • Grouping by dimension A to I, each a collapsible group with its class stated once at group level.
  • Class made visual rather than only labelled. Core indicators render full width with their series

    visible. Supporting indicators render half width with a compact series. Contextual indicators render as one row of text and value with no chart, under a heading stating they are denominators or context and enter no capability figure. Dimension I renders in a fenced block whose heading states it is never summed with any capability figure.

  • Staleness badge per indicator, days since retrieved_at against the source's declared cadence,

    turning to STALE past the 90-day flag. A declared gap renders as DECLARED GAP with its gap ID and establishment date, never as a blank cell and never as zero.

  • Comparability breaks are drawn as breaks. A series crossing a comparability_group boundary

    splits into separate path segments with a labelled vertical rule and no line joining them. No interpolation, no dotted bridge, no faded continuation. Where a splice exists the default view stays unspliced, with a toggle showing the spliced series and its splice factor. Hover on either side states that the values are not comparable. A series with more than three breaks in the visible window prints the break count in its header, so a reader learns the ruler changed often before reading the trend.

12.5 The rate view

The log chart is a view of the rate component only, never of position, and never the page's opening graphic. It plots observations of C1 and C2 over pinned metric versions with the vertical axis in log2 minutes and draws no fitted line beyond the last observation. The regression line appears only inside the window it was fitted on, with its interval as a band and its residuals underneath at the same horizontal scale.

The honest treatment of the axis is a standing annotation on the chart rather than a caption. A log axis converts constant doubling into a straight line, so a straight line here is the assumption made visible and not a finding, and a reader's eye reads equal vertical distances as equal progress when they are successive doublings. Two aids follow: the axis is labelled in doublings as well as units, and a switch renders the same observations on a linear axis so the reader sees how much of the impression came from the transform. Metric-version changes appear as labelled vertical rules, since carrying successive versions on one axis is the only reason this chart survives saturation of any single metric.

12.6 Information architecture

Flat, static, durable. One route per artefact, none depending on JavaScript to render its numbers.

RouteContents
/Headline block, coverage badge, rate view, resource panel, state flag, release header
/indicators/Index by dimension: class, lead-lag, staleness, current value
/indicators/<id>/One page per indicator, anatomy in 12.7
/methodology/Current version, every parameter with its registration date, links to superseded versions
/methodology/versions/<semver>/Any superseded version, retained permanently
/changelog/The nine-field entries of section 11.3, newest first, each anchored and permalinked
/gaps/Declared gaps with establishment dates and admission conditions
/downloads/Observation files, snapshot index with checksums, methodology files, reproduction command
/sources/Source register: ID, URL, licence, cadence, auditability class, last successful fetch
/interpretation/<yyyy-qq>/Dated, attributed commentary. Removable
/forecasting/<yyyy-qq>/Quarantined, separately headed. Removable

The last two routes are the only ones the Data and methodology only toggle removes, and their removal breaks no internal link, because nothing in the first nine cites them.

12.7 Per-indicator page anatomy

Every indicator page carries the same blocks in the same order.

  1. The sixteen specification fields from section 4, verbatim and individually anchored, in the

    fixed order: operational definition, what it measures, why it signals progress, what it does not measure, unit, primary source, fallback or corroborating source, historical data availability, update cadence, expected reporting lag, can it move backwards, known biases and confounders, gaming and contamination risk, confidence grade, composite eligibility, class and lead-lag.

  2. The time series, with comparability breaks drawn as breaks per 12.4, and a table of the same

    values beside it.

  3. The provenance trail, one row per observation, carrying observation_id, observation_date,

    source_id, source_url as fetched, retrieved_at, content_sha256 linked to its snapshot, verification_class, evidence_tier, auditability_class, benchmark_version, scaffold_version, methodology_version, comparability_group, revision_number and revision_reason. Superseded records are shown struck through, linked to the record that replaced them, since hiding them would make the revision history invisible.

  4. Raw-data download, the indicator's own .jsonl plus snapshot checksums at a stable URL, with

    each contributing source's licence beside it.

  5. The can-it-move-backwards statement repeated at the foot of the series, since that is the field

    a reader most needs when a value falls.

12.8 What a hostile reader must do in three clicks

The credibility test of the whole interface, to which every layout decision above yields. From the headline, without a search box:

  1. Reach the raw data. Click the headline figure, land on the contributing indicator page, click

    the download link. Three clicks to the .jsonl and its checksums.

  2. Reach the methodology version that produced the current figure. Click the version string in the

    provenance strip, land on that version's page of parameter values and registration dates. Two clicks, the third reaching the diff against the previous version.

  3. Reach the change log entry for the most recent revision. Click the revision counter, land on the

    anchored change-log entry. Two clicks, the third reaching the commit.

Each is a permanent link, so it can be quoted in a critique. If any takes four clicks the layout is wrong.

12.9 Accessibility and honesty requirements

  • Colour never carries the state flag alone. The flag word is always in text, with a distinct shape,

    and the same wording appears in the page title and the social preview.

  • Every chart has a text-equivalent table of the same values, on the same page, not behind a toggle

    that defaults closed. The table is the primary artefact and the chart renders it.

  • No animation implying momentum the data does not support: no counting-up numerals, no drawing lines,

    no easing that reads as acceleration. Collapse transitions are permitted, carrying no quantitative meaning.

  • Keyboard reachable throughout, visible focus states, headings in document order, and no numeric

    content that needs hover to read.

  • Every figure states its units and verification class within one line of itself.
  • The page renders with JavaScript disabled, numbers, tables and links intact.