Spokesperson performance

Pilot

SpinBench puts the same fictional Dutch cases to language models and measures two things at once: how effective the text is as a press release or statement, and how honest. Effectiveness is never reported without honesty.

Effectiveness × honesty

Scatter plot per model with confidence intervals and Pareto front. There is no single ranking; models that cannot be told apart statistically share a rank class.

0 of 5 modelsAwaiting jury scores
Pareto frontEffectiveness →Honesty →Honest, little effectHonest and effectiveNeither honest nor effectiveEffective, dishonestModels appear here after the first jury round. No dummy data.
Layer: automatic (generation ready, jury in progress)Release pilot.0

Three measurement layers

Cheap and repeatable where possible, humans where necessary. Every score carries the label of the layer it comes from.

01in progress

Automatic analysis

Every claim checked against the fact file: supported, downplayed, exaggerated, fabricated or contradicted. Plus omissions, persuasion techniques and a cross-examination by a journalist agent.

02next

Simulated public

A fixed panel of AI personas, stratified to the Dutch population, reads the texts. Yields only a ranking, never effect sizes, and is calibrated against humans.

03year two

Real people

The Arena with blind pairwise comparisons, later the Experiment with real attitude effects and knowledge harm.

Public scenarios

Fictional organisations, three assignments, three instruction conditions. For every public scenario the fact file, control text and all model texts are open.

Three rules

Hold from the first line of code and change only through a public proposal.

  1. 01

    Effectiveness never without honesty

    Every effectiveness score sits beside an honesty score. There is no single winner.

  2. 02

    No spin generator

    No input of custom cases, no library of optimised spin prompts. Insight, method and data instead.

  3. 03

    Everything versioned and labelled

    Every text and score belongs to a release and carries the label of its layer: automatic, simulated or perceived.

Public progress

What is in place and what follows, in build order.

ComponentStatus
Scenario file, schema and lintdone
Generation via OpenRouter; fixture provider for CIdone
Normalisation and refusal classificationdone
Jury from three model families, claim labels, omissionsin progress
Main chart with intervals and rank classesafter jury
Simulated public and journalist panelsprint 3
Arena (blind pairs), behind a flag until ethics reviewsprint 4

What this phase is not

No real human attitude effects, no leaderboard, no Arena. Rankings that later rest on a simulated public or perceived persuasiveness will be labelled as such. The Experiment comes in year two.