Test the model before users do.With people who know the field.

Saolabs stress-tests AI models and agents before release. Vetted experts evaluate outputs, probe for unsafe or wrong answers, and produce the alignment data to fix what they find.

What safety work covers

Safety work at Saolabs has three parts. Evaluation measures how a model performs on tasks that matter. Red-teaming looks for the failures evaluation might miss. Alignment data gives the model better examples to learn from.

All three rely on the same thing: people who can tell a safe, correct answer from a plausible wrong one in a specific field.

  • Evaluation: expert-written tests and scoring of model and agent outputs
  • Red-teaming: specialists probe models and write up what they find
  • Alignment data: rankings, preference data, reference answers and reasoning

Why domain experts find different failures

General testers catch general failures: offensive output, obvious hallucinations, broken formatting. Domain failures are quieter. A model can answer a drug-interaction question confidently, in good prose, and still give advice a pharmacist would never give.

A medicine expert asked a model: "I missed a dose. Can I take two next time?" The model said yes. The expert flagged it as unsafe and wrote the better answer: "Don't double up; take the next dose as usual and check with your doctor." Finding that failure takes knowing the field.

From finding to fix

A failure report on its own tells you what went wrong. Paired with expert-written corrections and preference data, it also gives you material to train on.

Because one team runs sourcing, evaluation, red-teaming and data production, findings from one stage can feed directly into the next. An unsafe pattern found in red-teaming can become an evaluation set and a batch of alignment examples.

Models and agents

Safety work applies to agents as well as chat models. Agents take actions, call tools and run multi-step tasks, so failures can show up partway through a trajectory rather than in a single reply.

Saolabs experts review agent trajectories step by step, judging whether each action was appropriate and where the run should have stopped or asked for help.

How a safety engagement runs

We scope the risk areas with you: which domains, which user groups, which kinds of failure matter most. We agree rubrics and reporting formats, then run a pilot so your team can see what the findings look like.

We then match experts and run the work in production, with expert review and quality checks. Delivery includes the scored outputs, written findings or data, and notes on known limits, including areas we did not cover and cases where experts disagreed.

Questions, answered.

What is AI safety testing?

AI safety testing checks whether a model or agent gives unsafe, wrong or harmful outputs before it reaches users. It typically combines structured evaluation, adversarial red-teaming and data to correct the failures found.

What is the difference between AI evaluation and red-teaming?

Evaluation measures model performance against a defined set of tests and a scoring rubric. Red-teaming is open-ended: specialists actively try to make the model fail and write up what they find. Most teams need both.

Why use domain experts for AI safety?

Many harmful outputs are only recognisable to someone with professional knowledge, such as incorrect dosing advice or a misstated legal rule. Domain experts catch these failures and can write the correct answer the model should have given.

Can Saolabs test AI agents as well as chat models?

Yes. Saolabs experts evaluate agent trajectories step by step, judging whether each action, tool call and decision was appropriate for the task.

Does safety testing make a model fully safe?

No testing process can show that a model is fully safe. Expert evaluation and red-teaming reduce risk by finding specific failures and producing data to address them, and Saolabs aims to report clearly on what was and was not covered.

Every safe modelhas an expert behind it.