Benchmarks
Capability is measured, not claimed
The platform evaluates itself with the same discipline it applies to research: defined batteries, measured baselines, pre-registered analysis and replication before acceptance.
How evaluation works
Versioned batteries
Capability tests are versioned and frozen before use — results reference an exact instrument.
Baselines first
No improvement is asserted without a measured baseline to compare against.
Worst-case floors
Long-horizon evaluations report the floor across independent seeds — not just the best run.
Ablation
Component contributions are established by removal, not by narrative.
Pre-registered analysis
Metrics, margins and victory conditions are locked in the record before collection.
Replication
Findings are re-run independently — different seeds, fresh state — before being treated as stable.
Benchmark design follows the platform ethics: saturated instruments are reported as saturated, ties are reported as ties, and uncertainty is quantified rather than hidden.
Reporting principles
- Internal scores are not marketing numbers — they carry method, margin and uncertainty.
- Component-level effects are reported by dimension, not averaged away into a single figure.
- Null and negative findings are published in the record as first-class results.
- External validation of performance happens under controlled conditions, with results made available to partners.
