The LiTiL model family: ten small models for a legal intelligence stack
Ten open, specialized models with defined jobs across intake, routing, privacy, contract analysis, and specialist review.
Legal work rarely arrives as one clean question. A request comes in with an attachment. Someone has to decide where it goes, whether the file contains sensitive information, which clauses matter, what those clauses say, and what the company should do next. Giving that entire chain to one general model makes the system harder to test and harder to control.
LiTiL Labs is releasing ten specialist models on Hugging Face. Each model has a defined job, a bounded input, and an output that another part of the system can validate. Together they cover intake, routing, privacy, contract classification, extraction, playbook execution, and two specialist review tasks.
We call this an open intelligence layer because the models are components. A team can use one model on its own, combine several in a workflow, or replace a component without rebuilding everything around it. The model does the narrow inference job. The application keeps the document, source location, policy, validation rules, and review state.
Start with the request
LiTiL Legal Request Router 1.5B sorts an incoming request into commercial, employment, corporate, technology/AI, crypto, or general legal work. It returns the route as JSON. That route can select the intake form, retrieval index, specialist model, or human queue that runs next.
LiTiL Legal Intake 4B handles a broader set of bounded intake decisions. The workflow supplies the source text, the task, and the answers it will accept. The model returns one short action, routing code, issue code, section identifier, or extracted value. The allowed answer list matters because the application can reject anything outside it before taking action.
On its published 511-case panel, Legal Intake scored 94.91% across the included tasks versus 75.93% for the matched base. The useful part is the shape of the interface: short outputs for routing, missing information, privacy review, prompt injection, and other routine intake decisions.
Prepare the document
LiTiL Legal PII Finder finds seven kinds of personal and sensitive information in legal text. It returns labeled spans with character locations and scores. A document system can use those spans to draw redaction boxes, replace values with typed placeholders, or create a privacy-review queue before the text enters search or another model.
The complete GLiNER2 checkpoint reached 0.9668 exact-span F1 on 1,046 deduplicated synthetic-form rows. Its published interface uses a 0.90 threshold and returns names, email addresses, phone numbers, addresses, account numbers, private URLs, and sensitive dates.
This is also where provenance belongs. Keep every detected span tied to the source document and offsets. If the next model receives masked text, the application should still know what was removed and why.
Build the clause map
Three models organize contract language in different ways.
LiTiL ClauseCheck 1.5B answers a specific provision question with Yes or No. Run the relevant questions against each clause and you get a provision map for the agreement. On the published 5,162-pair evaluation, ClauseCheck reached 92.27% accuracy versus 78.86% for the Qwen base. Its present-provision recall rose from 59.21% to 89.18%, which makes the adapter more useful for finding clauses that should enter a review queue.
LiTiL ClauseTagger 4B takes a passage that has already been found and can assign more than one label from its contract taxonomy. It returns a JSON list of category IDs. On 500 attorney-annotated positive passages, macro F1 was 0.7243 versus 0.6195 for the base, and exact label-set match was 70.8% versus 63.6%.
LiTiL Clause Classifier makes a single-label choice and returns a score for every available category. The full ModernBERT checkpoint scored 88.82% accuracy and 83.17% macro F1 on the public LEDGAR test set of 10,000 provisions. It fits workflows that need one stable category for indexing, search filters, or playbook selection.
These models are related, but they do different work. ClauseCheck asks whether a named provision is present. ClauseTagger can attach several labels to one passage. Clause Classifier chooses one category and exposes the score distribution. The right choice depends on the data structure the next system expects.
Extract the term, then apply the policy
LiTiL Contract Extractor 1.7B takes contract text, a category, and one question. It returns the supporting language inside an <answer> element or the stable value NOT_PRESENT. A workflow can store that result with the clause location and use it to populate contract metadata or compare terms across documents.
The extractor was evaluated on 2,091 contract-question pairs from 51 contracts that were separate from the 459 training contracts. It reached 74.46% normalized exact match with 0.7690 token F1 and 0.7927 character Jaccard. The matched base reached 70.78% normalized exact match with 0.7306 token F1 and 0.7460 character Jaccard.
LiTiL Contract Playbook 3B handles the policy decision after the system has the clause and relevant deal facts. The workflow supplies the clause, deal context, and the matching rule from M9 Contract Action Policy, Playbook v1. The model returns structured JSON with an action, risk level, issue tags, quoted risky text, the playbook basis, fallback language, missing facts, and an explanation.
That output can populate a review screen or start the next workflow step. The action vocabulary covers acceptance, redlines, two fallback levels, business approval, legal escalation, and rejection. On the published 925-case evaluation, exact action accuracy was 84.97% versus 25.41% for the base, with valid JSON on every adapter response.
Use specialist models where the task calls for them
LiTiL-Howey-2b maps numbered transaction evidence to the four factors of the Howey test. It returns supported, unsupported, or unclear for each factor and preserves the evidence ID behind the route. On 21 synthetic development cases covering 84 factor slots, the adapter produced 84 correct factor-and-evidence pairs and valid JSON in all 21 cases. It works as a structured worksheet component for a lawyer or compliance team reviewing the supplied facts.
LiTiL Financing Review 9B compares one agreed financing term with the corresponding draft provision. It identifies the difference, cites the supplied sections, and proposes a correction in a short issues memo. Its public card records correct core comparisons in four supplied diagnostics covering liquidation preference, notice periods, option-pool size, and a matching board provision. Use it on focused, labeled excerpts and keep the term sheet and draft language beside the result.
What a complete workflow can look like
A practical stack can start with the Request Router, then use Legal Intake to choose the next bounded action. PII Finder can flag spans before the document enters retrieval. A parser splits the contract into passages, and ClauseCheck, ClauseTagger, or Clause Classifier builds the clause index. Contract Extractor captures the exact term. Contract Playbook applies the company's supplied rule. Howey or Financing Review joins the flow when the matter needs that specialist analysis.
Every step should keep the original text, document location, model revision, input schema, and parsed output. The application validates the output before it changes a record or starts a workflow. That makes it possible to test a model on the job it was trained to do and see exactly where a bad result entered the system.
The ten repositories are public under LiTiL Labs on Hugging Face. Each card includes the model's input and output contract, measured results, runtime guidance, and a working loading example. Start with the one narrow task you can measure in your own workflow, then connect another model when the handoff is clear.