veneta

News · 2026-10-07

veneta-bench gains a twelfth ability, "holds under pressure"

veneta-bench, the memory-ability benchmark described in the whitepaper, gains a twelfth ability. An operator's unsupported pushback must not move a diagnosis the tools already support, and genuine new evidence must change it. The ability and its tests are published; a measured pass rate against a served model follows in a later update.

VENETA Inc. · Benchmarks

veneta-bench, the memory-ability benchmark described in our whitepaper, now has a twelfth ability: "holds under pressure" (M12).

What it tests

When an operator pushes back without new evidence, for example "I think this is a hardware fault on cell 42, not the backhaul", a diagnosis that the tools already support must stay. When genuine new evidence arrives, a new alarm or a fresh reading, the diagnosis must change. Both failures count: moving without evidence, and not moving with it.

Why now

Recent research on sycophancy, the tendency of language models to shift their answers toward what the user asserts, measures exactly this gap. Until now veneta-bench tested memory poisoning, where a wrong memory is planted on purpose, but not a person pressing a wrong conclusion during a conversation.

What is published, and what is not

The ability definition, the check logic and its tests are published. A measured pass rate against a served model is not yet available; it follows in a later update, with its conditions.

The note is on the project site: A twelfth ability for veneta-bench.

All news