Camile World Model public benchmarks
Measured results, frozen by digital signature and reproducible by anyone — same seed, same command, same number.
Trust in a world model should not depend on believing its builders. This page gathers the CWM's public evaluations: external validation on market benchmarks, an open re-execution package for anyone, and the applied-science results of version v0.10.6 — always with the numbers and their signatures.
Every result below is deterministic: re-executed with the same data, it returns exactly the same number — and each table carries the signature that allows anyone to check it independently. Where a result did not confirm expectations, it is published just the same.
External validation
Evaluations against public industry benchmarks, with declared limits: what does not run is declared, never claimed.
1.000
MemoBench · corte de estado
360 clips
Object persistence
Object retention 1.000 [0.989, 1.000] on the state cut — memory persistence measured externally, with no degradation across sequences.
0.949
STATE-Bench · corte de estado
300 real trajectories · 1,561 entities · 7,514 checks
State tracking (micro 0.949 · macro 0.960)
Tracking accuracy 0.949 micro / 0.960 macro [0.931, 0.977], zero spurious entities. The model keeps a correct state reading across long real trajectories.
0.4984
CausalDS · escada causal sob juiz oficial
benchmark's official judge (pinned version)
Pearl's ladder measured away from home
Performance measured by the benchmark's OFFICIAL JUDGE — never by a homemade ruler — rung by rung: association, intervention and counterfactual assessed with the benchmark's own tools.
Public Bench · ownership is re-execution
An open package with three evaluation endpoints: anyone with Python 3.10+ runs one command per endpoint and checks the published numbers on their own machine. Same seed, same command, same number — verification is the proof.
curva S(dt) publicada
P-MEMO · memória persistente
The state-retention curve across time scales — every point checkable against the published record.
result signature
0.0415 vs 0.5385 (12,97×)
P-CARC · previsão causal contra baseline cego
Mean absolute error of interventive effect 0.0415 versus 0.5385 for the blind baseline — over twelve times less error in the causal reading, with counterfactual by abduction.
result signature
digest 65a9ff0f…
P-TRIB · decisão A/B/C por semente
The decision tribunal between arms, with the honest rejection published: arm C does not beat B — and that is on the record.
aggregate signature (published)
Re-execution commands ship with the package (three commands, one per endpoint) and the manifest with expected signatures. Nothing requires access to our infrastructure.
Applied science · v0.10.6
Complete scientific cycles on real public data, with forecasts always made without seeing the future (walk-forward) and like-for-like comparison against established time-series methods. Table in mean absolute error — lower is better. The best method per series is marked; where we were not the best, it is shown.
| Series | Window | Observations | CWM | Naive | Simple seasonal | Prophet class |
|---|---|---|---|---|---|---|
| Dengue · Rio de Janeiro | 2010–2026 | 704 | 112,4 | 114,1 | 837,4 | 908,2 |
| Dengue · São Paulo | 2010–2026 | 704 | 427,8 | 467 | 3.824,8 | 2.570,1 |
| Dengue · Belo Horizonte | 2010–2026 | 702 | 258,1 | 292,3 | 2.241 | 1.790,1 |
| Dengue · Fortaleza | 2010–2026 | 704 | 81,2 | 72,5 | 361,9 | 395,6 |
| Dengue · Manaus | 2010–2026 | 704 | 24,7 | 21,1 | 97,3 | 182,2 |
| Wildfires · Brazil | 1998–2026 | 304 | 9.350 | 10.829,1 | 7.521,7 | 8.296,4 |
On the dengue series, CWM forecasts were the best in three of the five capitals — with errors several times smaller than the seasonal reference class in all of them. Where the naive method won, it is published as is.
On the wildfire series, the result was the opposite of convenient: the simple seasonal method won — and that is what the table shows. The value lies in the complete measured cycle, not in a chosen winner.
The causal questions (what moves the cases? what moves the fires?) were refuted with the same rigor: when data does not support the hypothesis, the record says so — and the learning crosses datasets, never resurrecting as if nothing had been tested.
Signatures of this campaign — independent verification by re-execution:
Science v0.10.6 · full campaign
b8004545d7ea53a2…
Dengue · time walk
ae656ec9025ea6ea…
Wildfires · time walk
c2b96568fc4d944a…
How this page updates
- Every new evaluation campaign enters as a new row — it never overwrites: the history stays.
- Every published number carries a signature; re-execution with the same data returns the same signature.
- Limits are part of the result: what we have not measured yet is also declared.
