Back to Blog

LiTiL models for organizing contracts and extracting their terms

Clause Classifier and Contract Extractor work alongside our legal search models to turn contract passages into usable information.

An open intelligence layer has many parts. Our previous release of three models that help find relevant passages in text was only part of the equation. Finding the passages are great, but you also need to know what kind of provision it is, and whether it answers a specific question. To that end, we're releasing two specialist models to add to your stack: LiTiL Clause Classifier and LiTiL Contract Extractor 1.7B. These are useful when you want to organize an agreement, and populate contract records with supporting text. They can also support Retrieval Augmented Generation (RAG), where retrieved text helps a model write an answer.

These two releases add other jobs around that search. Together, they show how LiTiL's open intelligence layer can fit inside a contract system, with separate components for organizing text, finding it, and extracting information from it.

For LiTiL Classifier we fine-tuned ModernBERT-large on 60,000 public LEDGAR provisions through LexGLUE for classification. For extraction, we adapted Qwen3-1.7B using public CUAD contracts, questions, and annotated answers. Each model learns a different output: a category for the classifier, supporting contract language for the extractor.

What is a classifier?

A classifier takes an input and assigns it a category. An email classifier might label a message as spam. A contract classifier might label a provision as confidentiality or governing law. During training, we showed the model examples paired with their categories. It learns patterns in the text that help it choose a label for a new example, including when the wording differs from the examples it saw. With the LiTiL Clause Classifier it takes one provision, scores each of its 100 categories, and returns the highest-scoring label. Your application can store that label beside the original text or show the competing labels to a reviewer when the choice is close.

What does the extractor add?

An extractor takes text and a request for a particular piece of information. For a governing-law question, it's trained to return the relevant wording from text. Your application retains that wording beside a contract field and preserves the source reference so a reviewer can check it.

LiTiL Contract Extractor uses one CUAD category and question per request, across 41 supported categories. Its output is a text span inside an <answer> tag, or NOT_PRESENT. That absence result concerns the text it received; it does not prove the term is missing from an entire agreement when only an excerpt was supplied.

Classifier, embedder, retriever, extractor

These components produce different outputs, even when they work on the same text.

ComponentWhat it doesWhat your application gets
ClassifierAssigns text to a categoryA label, such as governing law, with category scores
Embedding modelConverts text into numerical vectors that can be comparedA representation of a question or passage for similarity search
RetrieverSearches a collection and selects relevant passagesRanked source text to show a person or pass to another model
ExtractorSelects language answering a specific question in the supplied textSupporting contract language, or an absence value

With dense embedding models such as LiTiL Embed 0.6B and Octen Law 8B, you encode document passages ahead of time and store their vectors in an index. When someone asks a question, you encode the question too. The search system compares vectors to find candidate passages, including passages that use different words from the question.

The retriever is the search component around that process. It uses the index, applies filters and permissions, and returns the selected text. A retriever can use embeddings, keyword matching, or a combination. An embedding model alone does not store your documents or run the entire search system.

LiTiL ColBERT 300M takes another approach to retrieval. It keeps vectors for individual tokens and compares the question and passage at that finer level. It needs a compatible search setup, such as PyLate, rather than the same single-vector index used by the dense models.

The classifier can add useful metadata to either setup. It does not need a search question to label a provision, and a retrieval system can work without classification.

How the pieces fit into a legal workflow

Let's say you need to find governing law clauses across your contracts, and then check them against your playbook....here's how you'd do it.

  1. Read and split the documents. The application extracts the text, separates provisions, and preserves their document and page locations.
  2. Label and index the provisions. Clause Classifier adds a category to each provision. An embedding model separately creates the vectors used for search. The application stores both beside the source text.
  3. Retrieve relevant passages. When a question comes in, the retriever searches the index. The application can use category labels as filters or ranking signals. Because labels can be wrong, strict filters can also hide relevant passages.
  4. Extract the requested term. Contract Extractor receives the selected text and a supported governing-law question. The application checks that the returned language occurs in the source and saves it with the document location. The classifier's LEDGAR labels and extractor's CUAD categories are different taxonomies; if classifier labels select the extraction category, the application must map between them rather than pass labels through unchanged.
  5. Use the result. The application can populate a contract record or send the source-backed result to a lawyer or playbook workflow. Deciding whether the term meets company policy is a separate review step.

For example, suppose the contrract says an it's governed by Oregon law. The classifier identifies it as a governing-law provision. The embedding model makes it searchable by meaning. The retriever brings it back for a relevant question. And then the extractor selects the wording identifying the governing law. Voila, you have a workflow.

Run it where your contracts live

Clause Classifier has been tested on CPU, so it does not require a dedicated GPU. Its full checkpoint is about 1.58 GB, plus runtime memory. Contract Extractor ships as a LoRA adapter: load it with the Qwen3-1.7B base model. The adapter's download size alone is not its memory requirement. See each model card for its loading instructions; this is not a measured memory budget for running the whole system together.

For a local contract tool, that means classification can happen inside your own environment. You choose where to store the text and which systems receive it afterward.

Start by splitting a contract into provisions and keeping each provision's source location. Classify each one, then attach the label to that record. Clause Classifier uses a 512-token input limit and gives each provision one label, so long or multi-topic sections need attention before classification.

One part of LiTiL's open intelligence layer

LiTiL's open intelligence layer is a set of specialist components that teams can arrange around their own documents and workflows. A clause library might use classification without a chatbot. A contract database might use extraction to fill fields. A legal assistant might combine both with retrieval to gather the source material for an answer.

The surrounding application handles document parsing, storage, permissions, category mapping, source checks, and the review interface. The models supply specific outputs at those steps. You can keep your existing database or search system and add the component you need, then test or replace that component separately.

Foundations and credits

LiTiL Clause Classifier's checkpoint is labeled Apache-2.0. It builds on ModernBERT from Answer.AI, LightOn, and collaborators, also released under Apache-2.0. We credit the LEDGAR authors and LexGLUE authors and maintainers for the public training and evaluation data, listed under CC BY 4.0. LiTiL fine-tuned the base model for this classification task.

Contract Extractor builds on Qwen3-1.7B and CUAD from The Atticus Project and its authors. Its repository currently lists its license as other; the classifier's Apache-2.0 label does not apply to the extractor. The model cards provide loading instructions and evaluation details. Keep applicable licenses and attribution notices with redistributed materials.