Back to Blog

LiTiL ClauseCheck 1.5B: turn contract text into a provision map

A small contract model that checks whether a requested provision appears and helps build a reviewable provision map.

Most contract systems eventually need a plain answer to a plain question: does this clause contain the provision we are looking for?

LiTiL ClauseCheck 1.5B is built for that job. Give it one contract clause and one supported provision question. It returns Yes or No. The questions cover commonly reviewed terms such as assignment, renewal, termination, intellectual property, confidentiality, and limits on liability.

That short output is useful because it can become part of a durable contract index. Instead of asking a model to write a review memo for every clause, the application can record the document ID, clause location, provision name, answer, and model revision. Those records can drive search filters, portfolio views, and review queues.

Where it fits

ClauseCheck belongs after document parsing and clause segmentation. A parser finds the clause boundaries. The application chooses the provision questions that matter for the agreement type. ClauseCheck runs those questions against the relevant clauses and returns binary answers.

The resulting map can do several things:

  • mark agreements that contain an assignment restriction;
  • find clauses that may limit liability;
  • queue renewal language for an operations review;
  • open a review screen at the exact clause; and
  • send the tagged clause to an extractor or playbook model.

The handoff matters. ClauseCheck tells the system whether the requested provision appears. LiTiL Contract Extractor 1.7B can then capture the exact date, amount, party, or language. LiTiL Contract Playbook 3B can apply a supplied company rule to the tagged clause.

The interface is intentionally small

The model is a PEFT LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct. The user message contains the clause, the question, and an instruction to answer Yes or No. Greedy decoding and a four-token generation limit keep the response short. The application should accept only an exact Yes or No and reject anything else.

A request has this shape:

Clause:
{contract clause}

Question: {supported provision question}

Answer (Yes or No):

The binary answer does not remove the need to keep evidence. Store it beside the source clause. A reviewer should be able to click the tag and read the language that produced it.

What the measurements support

The public card reports results on 5,162 clause-and-question pairs. The matched Qwen base reached 78.86% accuracy. LiTiL ClauseCheck reached 92.27%.

The difference becomes clearer when you look at present provisions. The base had 98.22% precision when it answered Yes, but its recall for provisions that were actually present was 59.21%. In other words, it was conservative and missed many provisions. ClauseCheck raised present-provision recall to 89.18% while keeping precision at 95.21%.

On clause text absent from the training pool, the adapter reached 92.32% across 2,553 pairs. Its Brier score was 0.0559 versus 0.1769 for the base, with lower being better. These are measurements for the published provision-question interface, which is the interface an application should preserve when reproducing the result.

A practical implementation

Start with a small question set for one agreement type. For an NDA, that might mean confidentiality scope, term, return or destruction, assignment, governing law, and liability language. Run each question against likely clauses, validate the output, and retain the source location.

Then measure the behavior that affects the actual review queue. Count missed provisions, incorrect positive tags, parse failures, and the number of clauses a reviewer has to reopen. A high precision score can still be frustrating if recall is too low for the workflow. ClauseCheck's measured change is useful because it found substantially more present provisions while retaining a clean binary interface.

The model is small enough to use as a local classification layer. The 1.5B base is about 3 GB in BF16 weights, and the adapter adds about 17 MB. The public card suggests roughly 4 to 6 GB of accelerator memory for short BF16 prompts after runtime overhead.

The repository includes the adapter, exact loading code, tested base revision, parsing rule, and a saved example: LiTiL ClauseCheck 1.5B on Hugging Face.