Independent AI & machine learning lab — London

The lab that builds the models and reads them just as hard.

We train machine learning systems that reason more reliably per unit of compute — then measure them until we can say exactly what they do. Models, methods and numbers, published where anyone can check them.

Renaissance painting of a woman in white silk robes working at a modern laptop
Field study n°—01 · The daily interface
  • HealthcareTriage advice at 3am
  • EducationA tutor for the kid nobody tutors
  • GovernanceCasework nobody should wait months for
  • ClimateForecasts that reach the farm in time
  • EnergyGrids balanced by the minute
  • TransportRouting that respects the school run
  • AgricultureSoil advice in the local dialect
  • JusticePaperwork read properly, for once
  • FinanceSmall print, actually explained
  • CitiesPlanning decisions people can question
  • WelfareBenefits found before they lapse
  • ResponseThe first hour of a bad day
  • LanguageThe other six thousand languages
  • ScienceHypotheses tested while you sleep
  • AccessibilityInterfaces that meet people halfway
  • CultureArchives that answer back

The models got smart.
Now they have to get good.

Read our mission

Four ways we
do the work

One lab, four programmes: we train the models, we take them apart, we watch them in the world, and we publish all three. Each exists because we kept asking a question nobody could answer with a leaderboard.

Renaissance youths gathered around a smartphone

Public evaluations

Model
Atlas

Model development

Kaer Reasoners

Small, auditable models trained for the thing that actually matters: getting a hard question right on the second read, not the fiftieth sample. Open weights wherever we can, always with the eval card attached.

Read the research

Private evaluations

Civic Systems

When an institution puts a model between itself and the public, something changes. We find out what — before the public has to.

Work with us
Renaissance scholar peering into a brass microscope

Interpretability research

Co — 
Intelligence

How we look
inside

Five instruments, one discipline: never claim more than you observed, never observe less than you claimed.

Behavioural probes

Structured scenarios, run thousands of times — against our own checkpoints before release, against everyone else's before deployment.

Interpretability audits

We trace an answer back to the mechanism that produced it. On models we trained we can go all the way down; on closed ones we say where the trail stops.

Deployment tracing

We follow live systems for months. Habits change slowly. Harms compound quietly.

Alignment stress-tests

We push until the incentives show through. Better us than someone with worse intentions.

Cost curves

A right answer nobody can afford is still a wrong answer. Every model we train is reported in accuracy per pound, not accuracy alone.

Benchmarks happen
in the lab

Behaviour happens
to people

The Atlas
of Model
Behaviour

Renaissance still life with books, marble bust and a modern laptop
300+models mapped since we started counting
42domains where behaviour gets measured
1.2Beveryday interactions in the trace corpus
18×cheaper to evaluate than a year ago — and falling
50+institutions who read us before they deploy
Deploying a model where it matters? Read us first. Open the Atlas
Hands in gold armour writing on a black tablet

Memory. Judgement. Doubt.
Taste. Care.

Everything that makes a human, human — that's the standard we align to. Not engagement. Not throughput.

How we hold that line
Classical cloudscape painting
Notes from
the field.

Model behaviour became public infrastructure. Nobody voted for it.

So we do two jobs at once. We build models we would be willing to defend line by line — and we keep the notebook on everyone's, ours very much included, where anyone can read it.

Watch carefully, write it down, publish it, and charge nothing for the truth. The training runs are ours to justify. The evidence belongs to whoever has to live with the result.

Anyone can generate.
Understanding is the scarce thing.

If you're about to put a model in front of people who never asked for one — talk to us first. It's cheaper than the alternative.

Start the conversation

Fair
questions

What is Kaer Labs, in one breath?

An independent AI and machine learning company. We train models — small, efficient, auditable ones, built for reliable reasoning rather than leaderboard position — and we study what models do once they're out in the world, in hospitals, classrooms, ministries, kitchens. Both halves get published, so that trusting a model stops being an act of faith.

Do you actually train models, or only evaluate them?

We train. The Kaer Reasoners programme is our own line of models: mid-size, densely supervised, tuned to spend compute on the questions that deserve it and stop early on the ones that don't. We are not trying to out-scale anyone — the frontier labs have that covered. We are trying to show how much of frontier reasoning is reachable at a hundredth of the budget, with weights you can inspect and an evaluation card you can argue with. Building and measuring are the same discipline here: we would not trust our own numbers if we had never had to hit them ourselves.

What do you mean by "democratising intelligence"?

Three curves, all bent the right way. Reliability high enough to put in front of a stranger. Cost low enough that the next question is effectively free. And evaluations open enough that you don't have to take our word for anything. Miss any one of the three and intelligence stays a luxury good.

Aren't there already a hundred benchmarks?

There are. Most of them measure what a model can do on its best day, with a well-posed question and a patient grader. We measure what it does on an ordinary Tuesday, for a tired person, with a badly-worded question and something real at stake. Those are different numbers. The second one is the one that matters.

Who reads your work?

Institutions deciding whether to deploy, builders deciding what to build on, and researchers who want behavioural evidence rather than leaderboard positions. The Atlas is public. The private studies belong to the partners who commissioned them — minus anything the public deserves to know, which we negotiate up front, in writing.

What does "aligned with human values" actually mean here?

We keep a working list — memory, judgement, doubt, taste, care, humour, patience — the unglamorous faculties that make a human, human. A system is aligned, in our books, when it strengthens those faculties in the people who use it rather than quietly renting them out. It's a high bar. That's rather the point.