The Kaer Labs blog

Written up, worked out

How we train our models and what the training actually bought. Findings with their methods attached, post-mortems with the bills included, and the occasional argument. Everything here was measured before it was written — our own systems held to the same rule as everyone else's.

Hands typing on a laptop on a gilded table
Guides · 06 Sep 2026

Choosing a halting threshold for your traffic

The halting threshold is the one number that decides what an adaptive-compute deployment costs and how often it is confidently wrong. A method for setting it from your own data in an afternoon.

Read it

Guides

Tips and how-tos for getting our models working on your own traffic.

Announcements

Product and policy news from the lab.

Releases

Model release notes, with the limitations listed in full.

Comparisons

Like-for-like comparisons against other models, benchmarks and labs — protocol attached.

Hands in gilded armour working at a modern keyboard

Kaer-R1 vs GPT-5: cost per solved task on hard reasoning

A like-for-like comparison of a 7B open-weight reasoner against a frontier closed model: cost per solved task, abstention, latency and what each is actually for. With the protocol attached.

Comparisons · 9 min02 Sep 2026
Woman in a blue turban writing on a tablet

Kaer-R1 vs Claude Opus: abstention, calibration and knowing when to stop

Two very different answers to the same problem — how should a model behave when it is out of its depth? A comparison of abstention quality, calibration and refusal behaviour, with the measurement protocol attached.

Comparisons · 8 min28 Aug 2026
A lone figure with a laptop under a vast painted sky

Kaer-R1 vs Gemini and Llama: the open-weight reasoning comparison

If you need weights you can host, inspect and fine-tune, the field narrows fast. A comparison of open-weight reasoning options on licence, hostability, cost per solved task and what each is genuinely good at.

Comparisons · 8 min25 Aug 2026
Renaissance still life with books, marble bust and a modern laptop

Kaer Atlas vs LMArena vs HELM: which LLM benchmark answers your question

Three ways of measuring language models, built for three different questions. What each one actually measures, where each breaks, and how to pick the right instrument for a deployment decision.

Comparisons · 9 min20 Aug 2026
Renaissance scholar peering into a brass microscope

Kaer Labs vs Epoch AI vs METR: who measures what in AI evaluation

The independent AI evaluation field has split into distinct specialisms — trends and compute, dangerous-capability testing, and deployed behaviour. A guide to who measures what, and which one to call.

Comparisons · 8 min16 Aug 2026
Laptop beside a stack of red leather books and a marble bust

Best small reasoning models in 2026: a buyer's guide with the costs attached

Small reasoning models now close most of the gap to frontier systems on bounded tasks at a fraction of the cost. What to look for, how to compare them honestly, and the six questions that decide the choice.

Comparisons · 9 min10 Aug 2026

Models

How we train them, and what the training actually bought.

Findings

Measured results from the research programme.

Methods

The protocols and recipes behind the numbers.

Engineering

Infrastructure, post-mortems and the bills.

Essays

Arguments, with evidence.