Verified confidence across eight independent scenarios and a new service adjustment
Confidence indicators are now published with the full distribution of results, and the service gains a documented adjustment parameter.
Camile AI has published the verification of the Camile World Model's confidence indicators across eight independent scenarios — and this time, not just the average of the results: the full distribution, including what did not turn out as expected. The choice is deliberate. In systems that work with uncertainty, a single number conveys more certainty than exists; a range shows the real behavior.
In the same version, the CWM service gained a new adjustment to the confidence measure, with its own contract and documentation, and went into production with this capability stable for those already integrating and more controllable for those operating it.
What changed
The CWM's confidence indicators are now presented with the full distribution of results obtained across the eight independent scenarios, rather than only the average value. This makes the variation between scenarios visible: where performance was consistent and where there was greater dispersion.
The service received a new option for adjusting the confidence measure, available as a configuration parameter, with complete contract and documentation. The new version is in production, documented, and has no impact on existing integrations.
What this enables
Publishing the full range allows those evaluating the CWM — engineering teams, partners, researchers — to judge the system's reliability from observed behavior, not from an aggregated summary. Confidence ceases to be a claim about the system and becomes a result that can be examined.
For those operating it, the new parameter makes it possible to calibrate how the service expresses and uses its confidence measure, according to the type of application and the degree of tolerance for uncertainty in each context. Adjusting behavior without rewriting integrations is the difference between operating a service and adapting it.
Why it matters
World models are useful to the extent that their confidence is credible. A system that represents states, makes predictions, and communicates uncertainty requires that communication to be verifiable — otherwise, those deciding have no way to distinguish signal from noise. By publishing the distribution of results and offering explicit control over the confidence measure, the CWM makes its behavior auditable and adjustable, two practical conditions for use in real decisions.
In practice
For a team building on the CWM, this means two concrete things: evaluating performance from published data rather than an average, and calibrating the degree of conservatism with which the service treats its own estimates. Applications that require caution can operate with a more demanding confidence measure; exploratory applications may prefer a more permissive adjustment, without changing the rest of the integration.
The service remains in production, documented, and ready for use, which sustains continuous evolution without disruption for those already depending on it.
Limits
The verification covers eight independent scenarios; it does not claim that the behavior generalizes to any context, nor does it replace one's own evaluation in specific domains. The published distribution includes results that were not perfect — and that is exactly the point: an honest reading of performance is more useful than the impression of precision.
Conclusion
Confidence, in systems that operate with uncertainty, is not a claim: it is a measure that is published, verified, and left open to adjustment. The advance in this version lies less in an isolated result and more in the standard it establishes — complete evidence, explicit control, and transparency about limits.
- confidence
- verification
- world model
- uncertainty
