Benchmarks

Capability is measured, not claimed

The platform evaluates itself with the same discipline it applies to research: defined batteries, measured baselines, pre-registered analysis and replication before acceptance.

How evaluation works

Versioned batteries

Capability tests are versioned and frozen before use — results reference an exact instrument.

Baselines first

No improvement is asserted without a measured baseline to compare against.

Worst-case floors

Long-horizon evaluations report the floor across independent seeds — not just the best run.

Ablation

Component contributions are established by removal, not by narrative.

Pre-registered analysis

Metrics, margins and victory conditions are locked in the record before collection.

Replication

Findings are re-run independently — different seeds, fresh state — before being treated as stable.

Benchmark design follows the platform ethics: saturated instruments are reported as saturated, ties are reported as ties, and uncertainty is quantified rather than hidden.

Reporting principles

  • Internal scores are not marketing numbers — they carry method, margin and uncertainty.
  • Component-level effects are reported by dimension, not averaged away into a single figure.
  • Null and negative findings are published in the record as first-class results.
  • External validation of performance happens under controlled conditions, with results made available to partners.