Our staged access policy for weights, harnesses and Atlas data

Why Kaer releases are staged rather than open-by-default, what each tier includes, and the cases where we hold a release back — written down so it can be argued with.

Hands in gold armour writing with a stylus on a black tablet

Every Kaer release — weights, evaluation harnesses, Atlas data — is staged rather than open by default. This is a policy people are entitled to push back on, so here it is written down with the reasoning attached.

The tiers

TierWhat it includesWho gets it
OpenProtocols, rubrics, coding manuals, per-model published results, failed variants, null resultsEveryone, no request needed
On requestModel weights, halting heads, training recipes, evaluation harnessesResearch and evaluation use, two working days
SubscriberAtlas model identities, raw session traces, full domain tableInstitutions and publishing researchers, under terms
HeldScenario banks; weights for stated public-facing use without a fallback pathNobody, currently

Why the scenario banks stay closed

This is the one that gets the most objections and it has the simplest answer. A published benchmark is a training target. Publish our banks and they join the next crawl inside a month, after which the numbers mean nothing to anybody — including us.

What we publish instead is everything needed to replicate the measurement: the protocol, the rubric, the judge-panel construction, the coding manual, the pinning policy and the per-model results. Anyone who wants to build their own bank to our method can, and we would consider that a better outcome than using ours.

Why weights are on request rather than open

Our models are research releases with no safety hardening beyond their base. That is a fine thing to hand to someone running an evaluation harness with a human fallback path, and a poor thing to hand to someone about to put it in front of the public without either. The request is not a gate on capability; it is a conversation about deployment.

In practice almost everything is approved. The cases we have held back were all the same shape: a stated intention to serve an unhardened research model directly to members of the public with no escalation route. We said no and explained why, and in two of three cases the team came back with a design that had one.

What we publish that most labs do not

Arguing with this

If you think a tier is wrong — particularly if you think something in "held" should move — write to us. We have moved things before on a good argument, and the policy is published precisely so that the argument has something to push against. [email protected].

← All posts Kaer-R1 7B release notes →

Common
questions

Why not just release everything openly?

For the evaluation material, because a published benchmark is a training target and an open scenario bank stops measuring anything within a month. For the weights, because our research models have no safety hardening and we would rather know roughly where they are going. Neither reason is about commercial advantage, and we have tried to write the policy so that claim can be checked.

What gets held back, in practice?

Scenario banks, always. Model weights where the stated use is public-facing deployment without an evaluation harness or fallback path. Nothing has been withheld on competitive grounds and if that ever changes we will say so here.

How long does a request take?

Usually two working days for research and evaluation use. Longer if the intended use is public-facing, because that conversation is worth having properly.