veneta

News · 2026-10-05

worldmodel-bench v0.2, a proposal for measuring a WorldModel as a whole

We publish worldmodel-bench v0.2, a proposed benchmark for a WorldModel as a whole, with ten axes and four gates drawn from what telecom, power, defense, government and finance require of such a system. Code grades every score, every number carries its conditions, and there is no composite score.

VENETA Inc. · Benchmarks

We publish worldmodel-bench v0.2, a proposal for measuring a WorldModel as a whole rather than one feature of it.

What it measures

The benchmark has ten axes: accuracy, stability, security, forecast accuracy, simulation, optimization, self-evolution, self-improvement, memory and performance. Before the axes come four gates that a system passes or fails: no data leaves the network, an authority level that bounds what the system may touch, audit and roles, and reproducibility with honest labels. The gates and axes come from reading what telecom operators, power utilities, defense, government, intelligence and finance across several countries require of a system that sits next to their operations.

How it is scored

Code grades every score. Every number carries its conditions: the edition, the data, the model, the number of runs and the date. There is no composite score, because a single number would hide which gate a system failed.

What we publish first

Nothing comparable exists yet to measure against, so we publish veneta WorldModel's own card first: what we have measured on the telecom edition so far, and what we have not measured yet, listed as plainly. The memory axis uses veneta-bench, the memory-ability benchmark described in the whitepaper.

The full proposal, the gates, the axes and the current card are on the project site: worldmodel-bench.

All news