Back to Blog

LiTiL Contract Extractor 1.7B: capture the term and keep the source

A contract extraction model that returns the language answering a specific question or a stable absence value.

A clause label tells you what kind of language you found. Most contract workflows still need the exact answer: the effective date, governing law, party name, renewal term, insurance requirement, or another value from the agreement.

LiTiL Contract Extractor 1.7B handles that second step. Give it contract text, a category, and one question. It returns the relevant language inside an <answer> element. If the term is absent, it returns <answer>NOT_PRESENT</answer>.

That output is easy to parse and easy to retain with its source. A contract system can write the extracted value to a structured record while preserving the original span and clause location for review.

Where it fits

The extractor belongs after document parsing and retrieval. For a long agreement, the application should first find the passage most likely to contain the answer. Then it sends one category-specific question with that text.

<context>
CONTRACT TEXT OR RETRIEVED PASSAGE
</context>

Category: Effective Date
Question: What is the effective date of this agreement?

The model returns either a supporting span or NOT_PRESENT. The parser should require one <answer> element. Preserve the returned text before applying date, currency, or entity normalization, then store the normalized value beside the original span.

This separation helps when the next system needs evidence. A renewal dashboard may display a normalized date, but a reviewer can still open the clause that produced it. A playbook model may use the extracted term as deal context, but the application retains the contract language behind that fact.

What the measurements support

The public evaluation contains 2,091 contract-question pairs from 51 contracts. The 459 training contracts and 51 test contracts are disjoint.

LiTiL Contract Extractor reached 74.46% normalized exact match, compared with 70.78% for the matched Qwen base. Token F1 increased from 0.7306 to 0.7690; character Jaccard increased from 0.7460 to 0.7927. The largest saved gains were on document name, agreement date, effective date, parties, and insurance questions.

The aggregate contains two different behaviors: copying the right span when a term is present and returning NOT_PRESENT when it is absent. A useful application test should report those behaviors separately. It should also measure the retrieval step because the extractor cannot return language it never receives.

A practical implementation

Start with the fields the business already tracks. For each field, define the contract category, canonical question, output parser, and normalization rule. Retrieve a focused passage, call the model once, and attach the result to the source location.

A basic pipeline can look like this:

  • Segment the agreement and keep page or character offsets.
  • Use a tagger, classifier, or retrieval query to find likely passages.
  • Ask one extraction question at a time.
  • Parse exactly one <answer> block.
  • Convert NOT_PRESENT to a structured null.
  • Normalize the value without discarding the returned span.
  • Route uncertain or conflicting fields to review.

Contract Extractor is a PEFT LoRA adapter for Qwen/Qwen3-1.7B. The adapter is about 266 MiB. The public card suggests 6 to 8 GB of accelerator memory as a practical starting point for short or retrieved passages with the BF16 base. Training used a 2,048-token envelope, so focused passages are the best match for the published setup.

The repository includes the adapter, prompt format, parser, loading example, and exact measured results: LiTiL Contract Extractor 1.7B on Hugging Face.