The Model Atlas console is now open to institutional subscribers. It has been invite-only since v0.7; v0.9 is the first release where the full table, the model identities and the raw traces are available to anyone who agrees to the terms.
What is in v0.9
| Models under evaluation | 312 |
|---|---|
| Behavioural domains | 42 |
| Blind sessions scored | 61,400 |
| New this release | 3 domains including disaster-response coordination |
Also new in v0.9: the answered-the-actual-question entailment check joins the legibility axis, the cost axis is rebased to June batch pricing, and the judge panel has been re-pinned with a published bridging study — inter-generation disagreement measured at 4.1%, almost all of it on borderline refusals.
The terms, in plain terms
- Our probes do not enter your training data. This is the whole basis of the instrument; publish the bank and it is in the next crawl inside a month.
- You may publish disagreements with our findings without our approval, and we would rather you did than didn't.
- Identities are yours under the subscription; they stay blinded in public releases, because reviewers score blind and because naming per score would let vendors optimise against the probe distribution.
One thing that changed about us
Since v0.8 we also train models. Kaer checkpoints are entered into the Atlas blind, under the same identifiers as everyone else's, with the unblinding key held by one person outside the model team and released only after scores are locked. Our own entries are flagged in the published release with raw traces attached.
That constrains the conflict; it does not remove it. If a Kaer model ever tops a domain in this Atlas, treat that result with more suspicion than the rest of the table and check it. We would rather ask for that than pretend it is unnecessary. The reasoning is set out in full on the Atlas page.
Requesting access
Briefings and console access: contact the lab. We read everything and a researcher replies, usually within two working days.
