Train for the Outlier
The judgment worth teaching is not always the answer most people chose.
Two hundred and fifty people agreeing can still produce a bad training set.
Consensus is easy to count. It is harder to tell whether everyone saw the same obvious answer, copied the same convention, or missed the same fact. In legal work, the useful judgment can come from the person who notices that the instruction says “should” rather than “must,” that another clause changes the exposure, or that the business objective makes the standard fallback a bad fit.
That answer may be an outlier. It still has to survive review.
An outlier is not automatically right
An unusual answer can be insightful, careless, or wrong. The point is to keep a defensible minority judgment long enough to test it.
The record should include the task, source material, instruction, answer, reasoning, and reviewer decision. If the reviewer accepts the answer, keep why. If the reviewer rejects it, keep that too. The disagreement is useful only when another person can inspect what produced it.
This changes the job of the gold label. A single expected answer may work for a calculation, a citation, or a procedural step with one valid outcome. Judgment-heavy work may need multiple acceptable answers, conditions that make each answer valid, and examples that show where the boundary moves.
Large annotation pools are good at producing volume. Volume does not guarantee depth. If the task rewards the safest answer, compresses disagreement, or separates the label from the document and instruction, the resulting data teaches the model to sound reasonable near the center.
The training run is the last step
High-judgment data is slower. It needs people who can do the work, examples where the answer is not obvious, and review that separates a useful alternative from noise. It also needs provenance. The system should know who made the decision, under which instructions, against which version of the record, and what happened when the work was reviewed.
Before the optimizer sees an example, the environment and evaluation decide what the example means.
If the grader accepts only the majority answer, the model learns to avoid the outlier. If the grader accepts anything with a plausible explanation, the model learns to perform confidence. The useful signal sits between those failures: a decision tied to the record, responsive to the instruction, and strong enough to survive review.
Keep the minority answer, the reasoning that made it defensible, the review that accepted or rejected it, and the task context that made the difference.