Inside AI Assurance

Validation by People

Trained validators who work through your agents, workflows and processes by hand and tell you whether the system did the right thing. Not a script—engineers, domain specialists, and where the behaviour calls for it, people representative of your actual users, under a defined method.

Why a Machine Can't Score This

An automated evaluation compares output to an expectation. That works when the right answer is knowable in advance. Most of what matters about an agent isn't.

Was the escalation warranted? Was the tone appropriate to a distressed customer? Did it stop when it should have stopped? Would a reasonable person, reading the transcript, say the system behaved well? Those are judgements, and a judgement needs a judge. Ground truth for agent behaviour is ultimately a human opinion — the question is only whether it is collected carelessly or under method.

You can automate "did it return what we expected?" The harder question ("should it have?") still has to be anchored to human judgment, even when a model helps you ask it.

The Method

  1. A Rubric, Agreed Before Anyone Starts

    We write down what good behaviour means for your system — the dimensions, the scale, the worked examples at each point — and you approve it. Disagreement about the standard is resolved before it contaminates the results, not after.

  2. Scripted and Unscripted Scenarios

    Scripted runs cover the paths you know about. Unscripted runs are validators using the system as a real person would, which is where the interesting failures live.

  3. Independent Scoring

    Validators score without seeing each other's answers, and agreement between them is measured rather than assumed. Low agreement on a dimension is itself a finding: it usually means the rubric is ambiguous, or the system's behaviour genuinely is.

  4. Every Judgement Recorded

    Each score is traceable to a validator, a scenario and a transcript. The result is auditable rather than anecdotal — you can show it to a regulator, a board, or an engineer who disagrees with it.

  5. The Disagreements Are Part of the Deliverable

    Where validators split, you get the split. Averaging it away would hide the most useful thing in the dataset.

Who the Validators Are

QASource engineers and domain specialists, working under the same access controls and in the same delivery centres as the rest of the engagement. For regulated products, the people assessing your agent's behaviour have usually spent years validating the system it sits inside.

That matters more than it sounds. A validator who doesn't understand claims adjudication cannot tell you whether an agent's claims decision was defensible. Domain knowledge is the difference between a score and an opinion worth having.

We can staff this today with people who already work on systems like yours. The controls they work under →

Where It Fits

  • Before release. An assurance finding on whether the agent behaves acceptably across realistic use, with the human dimensions scored rather than asserted.
  • As a ground-truth set. Human judgments become the labeled data your automated evaluations are measured against. Without it, your evals are checking themselves.
  • After release, on a sample. Behavior drifts when models, prompts and data change. A recurring human review of sampled production traffic catches what a regression suite was never designed to see.
  • On workflows and processes, not only agents. The same method applies wherever a system makes decisions a person used to make.

Bring Us the Behavior You Can't Score

Tell us what your agent does and what "good" would mean, and we will tell you how we would measure it.