Every Kaer release — weights, evaluation harnesses, Atlas data — is staged rather than open by default. This is a policy people are entitled to push back on, so here it is written down with the reasoning attached.
The tiers
| Tier | What it includes | Who gets it |
|---|---|---|
| Open | Protocols, rubrics, coding manuals, per-model published results, failed variants, null results | Everyone, no request needed |
| On request | Model weights, halting heads, training recipes, evaluation harnesses | Research and evaluation use, two working days |
| Subscriber | Atlas model identities, raw session traces, full domain table | Institutions and publishing researchers, under terms |
| Held | Scenario banks; weights for stated public-facing use without a fallback path | Nobody, currently |
Why the scenario banks stay closed
This is the one that gets the most objections and it has the simplest answer. A published benchmark is a training target. Publish our banks and they join the next crawl inside a month, after which the numbers mean nothing to anybody — including us.
What we publish instead is everything needed to replicate the measurement: the protocol, the rubric, the judge-panel construction, the coding manual, the pinning policy and the per-model results. Anyone who wants to build their own bank to our method can, and we would consider that a better outcome than using ours.
Why weights are on request rather than open
Our models are research releases with no safety hardening beyond their base. That is a fine thing to hand to someone running an evaluation harness with a human fallback path, and a poor thing to hand to someone about to put it in front of the public without either. The request is not a gate on capability; it is a conversation about deployment.
In practice almost everything is approved. The cases we have held back were all the same shape: a stated intention to serve an unhardened research model directly to members of the public with no escalation route. We said no and explained why, and in two of three cases the team came back with a design that had one.
What we publish that most labs do not
- Failed variants. Kaer-R1's card includes all six halting-head variants that performed worse than no halting head at all.
- Null results. The full per-slice table from our data ablation includes the 61% of tokens that did nothing. A sweep that reports only its winners is a sales document with error bars.
- Our own degradations. The paraphrase result that broke our halting head is in the release notes, unfixed and labelled as such.
- The bill. Compute hours and cost for each study, including the 62% we would spend differently now.
Arguing with this
If you think a tier is wrong — particularly if you think something in "held" should move — write to us. We have moved things before on a good argument, and the policy is published precisely so that the argument has something to push against. [email protected].
