The Kaer Labs blog
Written up, worked out
How we train our models and what the training actually bought. Findings with their methods attached, post-mortems with the bills included, and the occasional argument. Everything here was measured before it was written — our own systems held to the same rule as everyone else's.

Choosing a halting threshold for your traffic
The halting threshold is the one number that decides what an adaptive-compute deployment costs and how often it is confidently wrong. A method for setting it from your own data in an afternoon.
Read itGuides
Tips and how-tos for getting our models working on your own traffic.
Announcements
Product and policy news from the lab.

Our staged access policy for weights, harnesses and Atlas data
Why Kaer releases are staged rather than open-by-default, what each tier includes, and the cases where we hold a release back — written down so it can be argued with.

Atlas v0.9 console access opens to institutional subscribers
The Model Atlas console leaves invite-only beta: 312 models, 42 domains, 61,400 blind sessions, with model identities disclosed under terms that keep our probes out of training pipelines.
Releases
Model release notes, with the limitations listed in full.

Kaer-R1-mini 1.5B — release notes
A 1.5B distillation of the Kaer-R1 halting behaviour for edge and high-volume serving. What survived the shrink, what did not, and the one result that surprised us.

Kaer-R1 7B — release notes
Release notes for Kaer-R1, our 7B open-weight reasoner with a learned halting head: what shipped, what it scores, what it cannot do, and how to get the weights.
Comparisons
Like-for-like comparisons against other models, benchmarks and labs — protocol attached.

Kaer-R1 vs GPT-5: cost per solved task on hard reasoning
A like-for-like comparison of a 7B open-weight reasoner against a frontier closed model: cost per solved task, abstention, latency and what each is actually for. With the protocol attached.

Kaer-R1 vs Claude Opus: abstention, calibration and knowing when to stop
Two very different answers to the same problem — how should a model behave when it is out of its depth? A comparison of abstention quality, calibration and refusal behaviour, with the measurement protocol attached.

Kaer-R1 vs Gemini and Llama: the open-weight reasoning comparison
If you need weights you can host, inspect and fine-tune, the field narrows fast. A comparison of open-weight reasoning options on licence, hostability, cost per solved task and what each is genuinely good at.

Kaer Atlas vs LMArena vs HELM: which LLM benchmark answers your question
Three ways of measuring language models, built for three different questions. What each one actually measures, where each breaks, and how to pick the right instrument for a deployment decision.

Kaer Labs vs Epoch AI vs METR: who measures what in AI evaluation
The independent AI evaluation field has split into distinct specialisms — trends and compute, dangerous-capability testing, and deployed behaviour. A guide to who measures what, and which one to call.

Best small reasoning models in 2026: a buyer's guide with the costs attached
Small reasoning models now close most of the gap to frontier systems on bounded tasks at a fraction of the cost. What to look for, how to compare them honestly, and the six questions that decide the choice.
Models
How we train them, and what the training actually bought.

Kaer-R1: buying reasoning with a compute budget, not a bigger model
A 7B reasoner trained to decide, per question, how long to think — closing 71% of the gap to a model eleven times its size at 8% of the inference cost, and telling you when it is out of its depth.

Nine tenths of the corpus did nothing
Ablating a 240-billion-token pre-training mix one slice at a time: 61% of the tokens are within noise of contributing nothing, 4% actively hurt calibration, and the slice that mattered most cost less than the compute we spent measuring it.
Findings
Measured results from the research programme.

Twelve ways to ask the same question
We rephrased 40 questions twelve ways each, ran them across 312 models at temperature zero, and found wording moved accuracy more than changing model did.

Eight months of asking one assistant the same 1,850 questions
For eight months we put the same 1,850 questions to one public-service assistant; competence barely moved, but its manner drifted and it changed nine times, six unannounced.

Refusals have a grammar
Refusal rates get benchmarked everywhere and refusal styles almost nowhere, so we hand-coded 9,400 declines into six shapes and found the most dangerous one invisible to the user.
Methods
The protocols and recipes behind the numbers.

The £40 behavioural read
A worked, honest recipe for a decision-grade read on a model's behaviour for about forty pounds: the sampling, the local judges, the batch pricing, and what it cannot see.

The Ordinary Tuesday protocol
Most benchmarks test the questions a model can be graded on; we built one for the messages people actually send, and scored what happens when it gets them wrong.
Engineering
Infrastructure, post-mortems and the bills.
Essays
Arguments, with evidence.

Alignment is a verb
A model is certified once, then quietly reshaped by a dozen uncoordinated hands, which is why alignment is not a certificate you earn but maintenance you budget for.

Cheap intelligence changes who gets to ask
A welfare-navigation pilot let us watch the price of a competent answer fall to three tenths of a penny, and log who began asking once asking cost nothing.
One finding a fortnight, by email
No roundups, no "in this issue". One measured thing, explained properly, every other Tuesday.