What expert evaluation is
Expert evaluation measures a model's outputs against the judgement of people who practise in the field. Saolabs specialists write test prompts, define what a correct answer requires, and score responses against a rubric.
It sits between automated benchmarks and live user feedback. Benchmarks are fast but often saturated or easy to overfit. User feedback is real but arrives after release. Expert evaluation tests the hard, specific cases before users see them.
Tests written for your use case
Public benchmarks measure what someone else cared about. Saolabs experts write tests around the tasks your model will actually face, in the domains where it will face them.
A legal evaluation might focus on jurisdiction-specific rules and when to refuse. A software evaluation might check whether generated code handles the error paths, not just the happy path. The test set reflects your risk, not a leaderboard.
- Domain question sets with expert reference answers
- Scenario-based tests for multi-turn conversations
- Agent tasks scored step by step, not only on the final result
- Targeted sets for known weak spots or past failures
Rubrics and scoring
A score is only useful if you know what it measures. We agree rubrics with your team before scoring starts: correctness, safety, completeness, appropriate uncertainty, or whatever criteria matter for the task.
Experts score against the rubric and add short written comments where a number would hide something important. Where several experts score the same item, the level of agreement shows which results are firm and which reflect genuinely hard judgement calls.
Evaluating agents
Agents fail differently from chat models. A run can reach the right final state through an unsafe step, or stall because one tool call went wrong early.
Saolabs experts review agent trajectories action by action. They judge whether each step was reasonable, where the agent should have stopped or asked, and whether the outcome would be acceptable to a professional in that field.
How an evaluation engagement runs
We scope the domains, tasks and decisions the evaluation needs to support, then draft test items and rubrics with you. A pilot round shows how the scoring works in practice and lets us tighten the rubric.
We then match vetted experts and run the full evaluation, with expert review and quality checks. Delivery includes item-level scores, written comments and notes on known limits, such as areas the test set does not cover.