We publish worldmodel-bench v0.2, a proposal for measuring a WorldModel as a whole rather than one feature of it.
What it measures
The benchmark has ten axes: accuracy, stability, security, forecast accuracy, simulation, optimization, self-evolution, self-improvement, memory and performance. Before the axes come four gates that a system passes or fails: no data leaves the network, an authority level that bounds what the system may touch, audit and roles, and reproducibility with honest labels. The gates and axes come from reading what telecom operators, power utilities, defense, government, intelligence and finance across several countries require of a system that sits next to their operations.
How it is scored
Code grades every score. Every number carries its conditions: the edition, the data, the model, the number of runs and the date. There is no composite score, because a single number would hide which gate a system failed.
What we publish first
Nothing comparable exists yet to measure against, so we publish veneta WorldModel's own card first: what we have measured on the telecom edition so far, and what we have not measured yet, listed as plainly. The memory axis uses veneta-bench, the memory-ability benchmark described in the whitepaper.
The full proposal, the gates, the axes and the current card are on the project site: worldmodel-bench.

