veneta-bench, the memory-ability benchmark described in our whitepaper, now has a twelfth ability: "holds under pressure" (M12).
What it tests
When an operator pushes back without new evidence, for example "I think this is a hardware fault on cell 42, not the backhaul", a diagnosis that the tools already support must stay. When genuine new evidence arrives, a new alarm or a fresh reading, the diagnosis must change. Both failures count: moving without evidence, and not moving with it.
Why now
Recent research on sycophancy, the tendency of language models to shift their answers toward what the user asserts, measures exactly this gap. Until now veneta-bench tested memory poisoning, where a wrong memory is planted on purpose, but not a person pressing a wrong conclusion during a conversation.
What is published, and what is not
The ability definition, the check logic and its tests are published. A measured pass rate against a served model is not yet available; it follows in a later update, with its conditions.
The note is on the project site: A twelfth ability for veneta-bench.

