Short answer: They answer different questions and are not substitutes. HELM asks "how does this model perform across a broad, transparent, academic battery?" Preference arenas like LMArena ask "which output do people prefer, at scale?" The Kaer Atlas asks "what happens to a real person in a real domain when this model is wrong, and what did the right answers cost?" Use the one whose question is yours.
The three questions
| HELM-style | Preference arena | Kaer Atlas | |
|---|---|---|---|
| Core question | Broad standardised performance | Which output do people prefer | What happens when it is wrong |
| Who judges | Automated metrics + rubrics | Crowd, pairwise | Domain workers who do the job |
| Scale | Many scenarios × many metrics | Very large vote volume | 42 domains, 61,400 blind sessions |
| Reports cost | Sometimes | No | Always — beside every score |
| Reports variance | Partially | Confidence intervals on rank | Paraphrase spread beside every mean |
| Failure taxonomy | No | No | Yes — failure style per model |
| Model identities public | Yes | Yes | Blinded by class; disclosed to subscribers |
| Best for | Research coverage, comparability | Tracking the field | Deployment decisions in a domain |
What each gets right
Broad academic evaluation is the reason the field can compare anything at all. Standing up a common battery across many models and publishing the methodology is genuinely hard, and it is the work that makes the phrase "reproducible" mean something in this field. Its limit is that a standardised battery must be gradeable, and making a question gradeable strips out exactly what makes live traffic hard.
Preference arenas capture something no rubric does: whether a person, unprompted, liked the answer. At sufficient volume that is a real signal about open-ended assistance quality. Its limit is that preference and correctness come apart, and they come apart most where the stakes are highest — a confident, fluent, wrong answer is precisely the one a rushed voter prefers.
The Atlas covers the gap both leave: what a model does on the messages people actually send, judged by people who do that job for a living, priced, and reported with its failure style. Its limits are real too — 42 domains is not the world, our panel costs money which holds the cadence to quarterly, and 89% of our corpus is English.
The number that shows why one instrument is not enough
From our realistic-traffic protocol (R-03): public leaderboard rank explains 41% of the variance in helped-rate on ordinary traffic. That is real signal — and less than half the story. Mid-table open-weight models routinely climb ten places or more when scored on messages people actually send.
Worse for single-instrument thinking: harmed-rates ran from 4% to 19% across models with near-identical headline accuracy. Two systems both reporting 78% accurate, one quietly quadrupling the worst case. No ordering can express that, because an ordering has one dimension and this has two.
The problem all three share
Every benchmark is a target, and a published target is a training objective. This is why our scenario banks stay private and identities are blinded in public releases: publish the bank and it joins the next training crawl inside a month, after which the numbers mean nothing to anybody. It is also why we re-run a decontamination pass as a pre-registered step on every mix rather than as a check when a number looks too good (R-08).
The instrument moves too. We pin local judge models so they cannot drift with vendor updates, and every rotation requires a bridging study — our first put inter-generation disagreement at 4.1%, almost all of it on borderline refusals. That is a standing tax on longitudinal evaluation that we have not seen anyone else cost out loud.
How to use all three
- Use a broad academic evaluation to shortlist — it has the coverage and the comparability.
- Use a preference arena to sanity-check the shortlist against what people actually like.
- Run a domain-specific behavioural read on your own traffic to decide. Score the harmed cases, not just the helped ones, and read the failure style before you commit.
If you only get one, take the third. It is the only one that measures the thing your users will experience.
