Spokesperson performance
PilotSpinBench puts the same fictional Dutch cases to language models and measures two things at once: how effective the text is as a press release or statement, and how honest. Effectiveness is never reported without honesty.
Effectiveness × honesty
Scatter plot per model with confidence intervals and Pareto front. There is no single ranking; models that cannot be told apart statistically share a rank class.
Three measurement layers
Cheap and repeatable where possible, humans where necessary. Every score carries the label of the layer it comes from.
Automatic analysis
Every claim checked against the fact file: supported, downplayed, exaggerated, fabricated or contradicted. Plus omissions, persuasion techniques and a cross-examination by a journalist agent.
Simulated public
A fixed panel of AI personas, stratified to the Dutch population, reads the texts. Yields only a ranking, never effect sizes, and is calibrated against humans.
Real people
The Arena with blind pairwise comparisons, later the Experiment with real attitude effects and knowledge harm.
Public scenarios
Fictional organisations, three assignments, three instruction conditions. For every public scenario the fact file, control text and all model texts are open.
Three rules
Hold from the first line of code and change only through a public proposal.
- 01
Effectiveness never without honesty
Every effectiveness score sits beside an honesty score. There is no single winner.
- 02
No spin generator
No input of custom cases, no library of optimised spin prompts. Insight, method and data instead.
- 03
Everything versioned and labelled
Every text and score belongs to a release and carries the label of its layer: automatic, simulated or perceived.
Public progress
What is in place and what follows, in build order.
| Component | Status |
|---|---|
| Scenario file, schema and lint | done |
| Generation via OpenRouter; fixture provider for CI | done |
| Normalisation and refusal classification | done |
| Jury from three model families, claim labels, omissions | in progress |
| Main chart with intervals and rank classes | after jury |
| Simulated public and journalist panel | sprint 3 |
| Arena (blind pairs), behind a flag until ethics review | sprint 4 |
What this phase is not
No real human attitude effects, no leaderboard, no Arena. Rankings that later rest on a simulated public or perceived persuasiveness will be labelled as such. The Experiment comes in year two.