Models learn from the people who teach them.So we start with the people.

Saolabs produces expert-labelled training and evaluation data for AI model companies and labs. We source and vet the specialists, then run the labelling work ourselves.

What we deliver

Saolabs delivers data that a model can learn from or be measured against. That covers labelled inputs, such as images, video, audio and sensor logs, and written supervision, such as expert answers, rankings and step-by-step reasoning.

The work is done by people who practise in the field the data describes. A dosing question is reviewed by someone with medical training. A contract clause is read by someone who has drafted contracts. A failing test is diagnosed by an engineer who writes code for a living.

  • Labelling: images, video, audio and sensor data, labelled precisely
  • Training data: expert-written answers, rankings and reasoning
  • Evaluation data: expert-written tests and scoring
  • Alignment data: preference judgements and reference answers

Why expertise changes the data

Most labelling errors that matter are not typos. They are plausible answers that a non-specialist would accept and a specialist would not. A model trained on those labels learns to sound right in exactly the places where being wrong is costly.

Take a simple case. A user asks a model: "I missed a dose. Can I take two next time?" The model says yes. A medicine expert flags the answer as unsafe and writes the better one: "Don't double up; take the next dose as usual and check with your doctor." That correction is the data.

Expert labellers also notice when a task is ambiguous, when a guideline does not cover a case, or when the honest answer is "it depends". Those notes are often as useful to a research team as the labels themselves.

Data types we work with

Saolabs works across the modalities that current models are trained and evaluated on. The common thread is that each type needs a human with domain knowledge to judge what correct looks like.

Formats follow your pipeline. We agree the schema up front, whether that is JSONL for SFT pairs, ranked completions for preference training, or frame-level annotations for video.

  • Language and code
  • Vision, video and voice
  • Agent trajectories
  • Simulation and robotics sensor logs
  • Physical-world sensor data

One team, sourcing through delivery

Saolabs sources pre-vetted experts and produces the data with them. The same team that recruits a cardiologist or a tax lawyer also writes the guidelines they work to and reviews what they produce.

Experts come in through our Hiring OS, outreach agents, recruiter network and jobs marketplace. They are vetted through AI voice interviews, real-world work tests in their field and skills assessments before they are matched to a project.

How an engagement typically works

We start by scoping: what the data is for, what a good label looks like, and where the hard cases are. From there we produce a small sample set so your team can check the output against your own judgement before anything scales.

Once the guidelines are settled, we match experts to the work and move into production, with expert review and quality checks along the way. Delivery comes with notes on known limits, such as categories with thin coverage or cases where experts disagreed, so you can decide how to weight the data.

Questions, answered.

What is expert-labelled AI training data?

Expert-labelled training data is data annotated or written by people with professional training in the relevant field, such as doctors, lawyers or engineers. It is used where correctness depends on domain knowledge that a general annotator would not have.

How is expert-labelled data different from crowdsourced labels?

Crowdsourced labelling suits tasks most people can judge, like spotting a car in an image. Expert labelling is for tasks where a plausible answer can still be wrong, such as a dosing instruction or a legal interpretation. Experts also flag ambiguous cases and gaps in guidelines, which improves the dataset beyond the labels themselves.

Which industries does Saolabs source experts from?

Saolabs sources experts in medicine, law, finance, engineering, software, science, mathematics and languages. Experts are vetted for their field before they work on a project.

What data formats can Saolabs deliver?

Saolabs agrees the delivery format with each client during scoping so the data fits the existing pipeline. Typical examples include JSONL for supervised fine-tuning pairs, ranked completions for preference training and structured annotations for vision, video or sensor data.

Does Saolabs build its own AI models?

No. Saolabs produces expert-labelled training and evaluation data that helps other companies make their AI models and agents safer.

Every safe modelhas an expert behind it.