Back to Blog

The Murky Middle

The obvious cases are easy to label. Legal judgment lives in the part that depends on context.

In our redline work, the obvious cases are rarely the interesting ones.

High-risk language can be easy to reject. Low-risk language can be easy to accept. The middle is where the answer starts moving with the context. A word in the instruction changes the weight. The person giving the instruction changes what it means. A commercial objective makes a fallback more attractive. Another clause changes the exposure.

That middle is difficult to label because it contains judgment rather than one fixed rule.

Medium risk is not one thing

A medium-risk issue is often treated like a point between high and low. That loses the reason it is difficult.

The right decision may depend on which party the lawyer represents, how the product works, whether the deal is strategic, what other protections exist, and which risks the business can carry. Two reviewers can reach different answers and both can be defensible. One may notice a fact the other missed. One may assign more weight to a business instruction. One may have seen the same failure before.

Collapsing those answers into one gold label throws away the disagreement that the model needs to understand.

A useful training example should include more than the final accept-or-reject label. It should preserve the document, instruction, relevant facts, proposed edit, explanation, and review. When experts disagree, the record should show why.

That gives the system something it can learn from. It can see which facts changed the outcome and which details were noise. It also gives the evaluator a way to distinguish a surprising answer from an unsupported one.

The grader has to hold the boundary

The grader has to allow defensible alternatives without becoming so loose that every polished answer passes. Mechanical checks can confirm that the agent read the right files, changed the right artifact, and cited real language. A rubric can test whether it addressed the required issues. Expert review still has to decide whether the judgment holds together.

An evaluation filled with obvious cases can make a model look better without showing whether it can do the harder work. The test set needs examples where small changes in context should move the decision and examples where irrelevant changes should not.

The comparison should not stop at an average score. Count how many additional decisions the changed model gets right. Look at where it changed its mind for the wrong reason. Keep the alternatives that survived review.

A useful system has to show what moved the decision and preserve enough of the record for someone else to challenge it.