AI research and software

We measure before we claim.

Hatteria Labs studies how language models work from the inside. Every claim has to pass a control before we rely on it, and we build products on the parts that hold up.

We measure

We study what happens inside language models: how they store information, how their internal parameters are organized, how they can be simplified, and how different kinds of models can be combined. Every finding also has to hold up against a control measurement that checks whether the effect is real or just chance.

We verify

Decision thresholds are set before the measurement, not after. When a later test overturns an earlier conclusion, the correction stays visible next to it.

We build and operate

Our own software, built and operated by the same people who run the experiments. That is how we find out how a result behaves once someone else depends on it.

Products

What we build

Our products are built and run on our own servers in the European Union. Torumata is live, AIDJ and Slidify are in closed testing. Running them keeps the research anchored in what works outside a benchmark.

More about the products →

torumata.com

An SEO audit for the era of AI search. See how language models actually read your site.

What it does

  • A technical SEO and AI readability score across five categories, so progress is measurable rather than felt
  • A language-model analysis, run page by page, of which questions the site can actually answer (AEO)
  • Measuring whether AI actually cites you (GEO): we ask, on your behalf, the questions your site should be able to answer, and find out who ChatGPT, Claude, Gemini and the Exa search index cite instead of you

For: Site owners and marketing teams working on SEO who see their traffic moving from search results to chatbot answers. And companies that need to demonstrate their website's accessibility under the European standard.

AIDJ

Closed testing

aidj.cloud

The DJ that runs your party on autopilot.

What it does

  • Guests request a track from their own phone after scanning a QR code, with nothing to install
  • A spoken DJ introduces the track in its own synthesised voice and crossfades into the next one
  • A message can be added to the event in progress without any moderator involvement, timed for a specific moment, or delivered immediately

For: Weddings, company parties, bars and clubs, birthdays and school events, anywhere the music matters and nobody should have to spend the evening managing a playlist.

Slidify

Closed testing

slidify.cloud

Your guests' photos, live on the screen.

What it does

  • Guests scan a QR code and upload from a mobile browser, no app, no account, nothing to explain
  • Photos reach the projector or television within seconds of being taken
  • The slideshow skips what it has just shown and favors pictures nobody has seen yet

For: Anyone hosting a wedding, a celebration or a company event who would rather collect the evening's photographs as they are taken than retrieve them afterwards.

How we work

How we know a result holds up

A measurement can easily show an effect that isn't really there. Our process is built to catch those cases in our own work.

Every result is checked against chance
Next to every result, we run the same measurement again on data that has been shuffled, destroying the structure the claim depends on. If that shuffled version comes out about the same, the claim fails. We record it either way.
Thresholds are set before the measurement
The criteria for judging a result a success or a failure are written down before the experiment runs, specific enough to be answered with a plain yes or no. A verdict written after looking at the data is not a verdict.
Corrections stay visible
When a later replication overturns an earlier conclusion, the correction stays visible next to the original rather than replacing it.
Easy to measure is not the same as important
When something is hard to measure directly, it is tempting to measure an easy substitute instead, for example how much smaller a model got rather than how well it still remembers afterwards. That substitute often measures only itself, not the thing that actually matters. We have hit this more than once, each time with a finished conclusion that would have passed review. When that happens, we retract the conclusion and measure the thing that actually matters. The correction stays visible next to the original claim.
The end-to-end effect is the only score
Many intermediate metrics along the way look convincing but may have nothing to do with what actually matters. What counts is only what the change does to the model's real, end-to-end behavior, checked on data the model never saw during development.

Are you working on something related?

Alongside our own products we accept a limited number of research and engineering engagements in the area of language models and the software around them.

Get in touch