Pular para o conteúdo
Camile AI — Built to Exist

Abrir busca rápida

Buscar páginas, documentação e ações…

CWM · Benchmarks

Camile World Model public benchmarks

Measured results, frozen by digital signature and reproducible by anyone — same seed, same command, same number.

Trust in a world model should not depend on believing its builders. This page gathers the CWM's public evaluations: external validation on market benchmarks, an open re-execution package for anyone, and the applied-science results of version v0.10.6 — always with the numbers and their signatures.

Every result below is deterministic: re-executed with the same data, it returns exactly the same number — and each table carries the signature that allows anyone to check it independently. Where a result did not confirm expectations, it is published just the same.

External validation

Evaluations against public industry benchmarks, with declared limits: what does not run is declared, never claimed.

1.000

MemoBench · corte de estado

360 clips

Object persistence

Object retention 1.000 [0.989, 1.000] on the state cut — memory persistence measured externally, with no degradation across sequences.

0.949

STATE-Bench · corte de estado

300 real trajectories · 1,561 entities · 7,514 checks

State tracking (micro 0.949 · macro 0.960)

Tracking accuracy 0.949 micro / 0.960 macro [0.931, 0.977], zero spurious entities. The model keeps a correct state reading across long real trajectories.

0.4984

CausalDS · escada causal sob juiz oficial

benchmark's official judge (pinned version)

Pearl's ladder measured away from home

Performance measured by the benchmark's OFFICIAL JUDGE — never by a homemade ruler — rung by rung: association, intervention and counterfactual assessed with the benchmark's own tools.

Public Bench · ownership is re-execution

An open package with three evaluation endpoints: anyone with Python 3.10+ runs one command per endpoint and checks the published numbers on their own machine. Same seed, same command, same number — verification is the proof.

curva S(dt) publicada

P-MEMO · memória persistente

The state-retention curve across time scales — every point checkable against the published record.

result signature

0.0415 vs 0.5385 (12,97×)

P-CARC · previsão causal contra baseline cego

Mean absolute error of interventive effect 0.0415 versus 0.5385 for the blind baseline — over twelve times less error in the causal reading, with counterfactual by abduction.

result signature

digest 65a9ff0f…

P-TRIB · decisão A/B/C por semente

The decision tribunal between arms, with the honest rejection published: arm C does not beat B — and that is on the record.

aggregate signature (published)

Re-execution commands ship with the package (three commands, one per endpoint) and the manifest with expected signatures. Nothing requires access to our infrastructure.

Applied science · v0.10.6

Complete scientific cycles on real public data, with forecasts always made without seeing the future (walk-forward) and like-for-like comparison against established time-series methods. Table in mean absolute error — lower is better. The best method per series is marked; where we were not the best, it is shown.

SeriesWindowObservationsCWMNaiveSimple seasonalProphet class
Dengue · Rio de Janeiro2010–2026704112,4114,1837,4908,2
Dengue · São Paulo2010–2026704427,84673.824,82.570,1
Dengue · Belo Horizonte2010–2026702258,1292,32.2411.790,1
Dengue · Fortaleza2010–202670481,272,5361,9395,6
Dengue · Manaus2010–202670424,721,197,3182,2
Wildfires · Brazil1998–20263049.35010.829,17.521,78.296,4

On the dengue series, CWM forecasts were the best in three of the five capitals — with errors several times smaller than the seasonal reference class in all of them. Where the naive method won, it is published as is.

On the wildfire series, the result was the opposite of convenient: the simple seasonal method won — and that is what the table shows. The value lies in the complete measured cycle, not in a chosen winner.

The causal questions (what moves the cases? what moves the fires?) were refuted with the same rigor: when data does not support the hypothesis, the record says so — and the learning crosses datasets, never resurrecting as if nothing had been tested.

Signatures of this campaign — independent verification by re-execution:

Science v0.10.6 · full campaign

b8004545d7ea53a2…

Dengue · time walk

ae656ec9025ea6ea…

Wildfires · time walk

c2b96568fc4d944a…

How this page updates

  • Every new evaluation campaign enters as a new row — it never overwrites: the history stays.
  • Every published number carries a signature; re-execution with the same data returns the same signature.
  • Limits are part of the result: what we have not measured yet is also declared.