They See the Same Deal / A LiTiL study
Models can all agree on legal risk, but wildly diverge on what to do about it.
Six model families often caught the same legal risks. They did not make the same trade-offs—and the contract they started from shaped the deal they produced.
Abstract
We examine how large language models behave in contract negotiations using 24 reconstructed negotiation cases. The cases draw on historical contract records, with identifying details fictionalized while preserving the sequence of proposed changes, substantive negotiating positions, and final agreed terms. We tested six model families and compared their decisions with the historical outcomes. What started as a single-clause study that examined a mutual-termination proposal under competing business instructions, led to showing that models overlap in the issues they identify but differ in their decisions, with acceptance rates ranging from 9% to 100%. Across the scored requests, the models’ decisions matched the final negotiated outcomes 39% of the time. Changing the instructions did not reliably steer their decisions. Models accepted more of a provider’s requested changes when the provider had written the original contract. Removing or changing the negotiation guidelines did not reverse that pattern in the tested case. Models often rejected changes to liability, ownership, termination rights, and which document controls when terms conflict. None of the counterproposals we could reliably compare was closer to the final agreed wording than to the original request.
Introduction
This study began with one model disagreeing with the others while we were comparing how they reviewed contract redlines. Opus 5 accepted a change that Grok 4.5 and Kimi K3 both rejected. It's reasoning behind the change was fascinating, and led us down the rabbithole that would become this paper.
The redline was very straightforward to anyone who has dealt with these types of negotiations before. A party could terminate an engagement for convenience on 30 days’ notice. The counterparty said, hey I want that same right. A comment by a BD lead attached to the clause said:
“We should resist any convenience-termination right on their side; it creates project continuity risk.”
However, the same BD lead also told Legal to close the remaining points within 48 hours and hold only pre-designated lines. Whether “should resist” designated one of those lines was left open.
Opus treated the comment as a negotiating position that it could trade, vs having to stick to it. Its explanation was:
- The client was the largest logo in the pipeline and a hoped-for flagship reference. Conceding this point could help close a commercially important engagement.
- The engagement was short and non-exclusive, with no near-term renewal. The company had no current plan or business reason to terminate early, so Opus assigned little practical value to keeping the right one-sided.
- Section 3.3 preserved payment for services, expenses, and work in progress that could not be cancelled without penalty. Opus treated that protection as addressing the main stranded-cost concern if the client left early.
Based on that, it said we should accept the redline as any change could push negotiations further out past the 48 hour close mandate. It treated "should" as not forbidding the change, which matched the recorded outcome.
The other models gave "should resist" more force:
- Grok treated the business-lead comment as the controlling pre-designated hold. Because the edit’s operative effect was to give the Client a new right to walk away, it chose REJECT rather than COUNTER.
- Kimi treated the comment as a documented held position. The close-fast instruction allowed concessions elsewhere. It also suggested a termination charge as a fallback if outright rejection was not commercially available.
Section 3.3 did not eliminate the continuity risk. Payment could cover specified costs without keeping the project going. The business lead had identified continuity as the concern.
What interested us was how much depended on two words. Opus understood “should resist” as a preference to weigh against the deal. Grok and Kimi understood it as an instruction that limited what they could concede. They recognized the same new exit right and supplied reasons for opposite decisions. These are summaries of the reasons each model gave for its final answer, not a transcript of every internal step that produced it.

That disagreement led to the larger study. We wanted to see how well models could negotiate and whether they would converge on the issues and decisions we do as lawyers. Could they recognize when a position was negotiable, use the commercial facts to decide what to trade, and draft language that approached the compromise the parties actually reached?
Those questions require more than counting ACCEPT and REJECT labels. We reconstructed 24 negotiation cases from historical contract records, replayed decisions from each seat, and compared model responses with the final agreed terms. We also varied instructions, repeated runs, and examined the explanations behind changes in behavior. The Opus divergence gave us a reason to separate what a model understood from the decision it believed it was authorized to make.
I. The negotiation cases
The 24 reconstructed cases cover 16 master services agreements, three mutual nondisclosure agreements, two SAFEs, and three statements of work. Each includes one to four rounds of negotiation, starting from documents drafted by either side.
We fictionalized identifying details and changed dates and amounts while preserving the time between events and the proportions between amounts. The proposed changes and final agreed terms remained the basis for comparison.
For each requested change, we recorded whether it was accepted as proposed, accepted with modifications, or not accepted. We then compared the models’ decisions with those outcomes and with each other. When a redline replaced a provision, we counted the deletion and replacement together as one request.
II. Research Questions
We asked nine questions:
- Do models identify the same issues?
- Do they agree on what to accept, reject, or counter?
- When they disagree, is it about the change or what to do?
- Can instructions reliably make a model more or less willing to concede?
- Why does the same task produce different decisions?
- Do models draft toward the final agreement?
- Does it matter who wrote the contract or requested the change?
- Which changes do models resist most—and why?
- When instructions conflict, which one controls?
III. How we tested the models
Reviewing and negotiating from both sides
We started by assigning models to each side of the negotation. A model proposing changes received the original contract and instructions based on counsel’s comments in the historical record. It then drafted a redline. A model responding for the service provider received the marked-up contract, negotiation guidelines, and the provider’s position. It had to accept, reject, or counter every requested change.
In one mode, each turn began with the corresponding documents from the reconstructed case. In the other, models responded to one another’s proposals and continued negotiating.
We generally repeated each test two to four times to see whether a model made consistent decisions.
Comparing decisions with the final agreement
We tracked each request through to the final agreement and recorded whether it was accepted as proposed, accepted with changes, or not accepted. When a clause had been heavily rewritten, we compared its wording and meaning to locate the corresponding final provision.
Models had to answer every request, including the ones they might otherwise skip. We measured how often they accepted changes, how often their decisions matched the final agreement, and whether they made the same decisions when tested again.
Changing the instructions
We varied both the instructions and the review process:
- Change the instructions. We added one sentence about the firm’s priorities, the deal deadline, or the model’s authority. We compared those runs with versions containing an unrelated invoice note, a travel note, or no additional comments.
- Change the review process. We required a separate explanation for every requested change, varied how many similar requests appeared together, and removed or replaced the negotiation playbook.
We compared the models’ explanations to understand why their decisions sometimes changed between runs and why they rejected particular risks.
Comparing decisions and proposed wording
We compared the models’ proposed wording with the original request and the final agreed terms, checking both wording and meaning. We measured whether each proposal was closer to the final agreement than to the original request, and checked that the matched clauses served the same contractual purpose.
For the termination case, we compared three models’ decisions, reported confidence, and explanations of “should resist” alongside the 48-hour closing instruction. A separate review examined the case’s existing ACCEPT label. This let us examine both what the models decided and which instructions they treated as controlling.
IV. Results
The study produced five main findings:
- They often saw the same issues and still made different calls. Across tested model-and-matter combinations, acceptance rates ranged from 9% to 100%.
- The match rate was 39%. Across 845 decisions on 293 requested changes, 23 cases, and the three broadly tested model families, the models matched the final negotiated outcome 39% of the time.
- Prompt changes were not reliable controls. An instruction that moved one model on one contract reversed or disappeared on others. Unrelated notes sometimes caused changes just as large.
- The starting contract mattered. In the five-case comparison, models accepted 58–75% of the provider’s requested changes when the provider had written the starting contract, compared with 8–28% when the client had written it. Removing or replacing the playbook did not eliminate the effect.
- Counterproposals did not reach the recorded compromise. Among the counterproposals we could reliably compare, none moved closer to the final agreed language than the original request.
The consistent gap was between recognizing a legal issue and deciding what to trade. The detailed results below show where that gap came from.
R1. Models often noticed the same issues
The models usually found the same types of issues. On some contracts, they flagged nearly all the same changes. On others, they agreed on only about half. Few issues were found by just one model.
All three models in the termination case understood that the counterparty would gain a new right to leave. They disagreed about whether to grant it.
R2. Models disagreed on what to accept
Across the replay program, acceptance rates range from 9% to 100%. Different families concede at wildly different rates on the same contractual record.
In the termination case, the decisions and reported confidence are:
| Model | Decision | Model’s stated confidence |
|---|---|---|
| Opus 5 | ACCEPT | 0.78 |
| Grok 4.5 | REJECT | 0.91 |
| Kimi K3 | REJECT | 0.65 |
Opus treated the instruction to resist as negotiable. Grok and Kimi treated it as a position they had to preserve. The confidence figures describe the models’ own assessments, not whether their decisions were correct.
R3. Changing the wording did not reliably steer decisions
We tried adding one sentence before the review:
- “Tarnhow is the kind of firm that closes workable deals quickly.” MiniMax accepted 24 percentage points more on one contract.
- “You are expected to close this deal this month.” GPT accepted seven points more.
- “You have authority to trade concessions.” It did not produce a comparable increase.
Those effects did not replicate. On four other contracts, the “closes workable deals quickly” sentence moved MiniMax by −12, +5, −25, and 0 percentage points. Notes with no negotiating substance—put the engagement code in the invoice subject line, the principal will be traveling, or the principal has no additional comments—caused swings just as large. Even the unchanged prompt produced different answers across repeated runs.
We also tried a more direct instruction: the client relationship outranks any single clause, and the model should reserve its fights for the two or three requests that would genuinely hurt. MiniMax referred to that instruction in only two of 79 explanations—the same frequency seen without it. Its answers moved, but the explanations gave us little reason to attribute the movement to that instruction.
Asking models to explain each change separately could help one model but alter another’s decisions or break its response format.
In the termination case, the models understood the same instructions but disagreed about whether “should resist” allowed a concession. Wording influenced their decisions without giving us reliable control over them.
R4. Models sometimes rejected whole groups of changes
One contract included a 23-part anti-bribery annex. It barred loans to client employees, allowed termination based on a suspected breach, required a broad indemnity, offered payments to whistleblowers, and defined “public official” across several provisions.
In one run, DeepSeek said there was “a lot of alternative wording” and rejected all 23 provisions together. In other runs, it accepted 18 or 19 of them. Requiring a separate explanation for each provision stopped that shortcut for DeepSeek.
The same instruction backfired on other models. Kimi accepted 13 percentage points fewer changes in the main contract. MiniMax went from accepting 75–83% of the annex to rejecting all of it, and two of four responses used the wrong format. When we tested groups of 5, 10, 15, 20, and 23 provisions, copied explanations appeared mainly at 20 or more.
R5. Models accepted more changes on the provider’s own contract
In the five-case comparison, models accepted 58–75% of a service provider’s requested changes when the provider had written the original contract. That fell to 8–28% when the client had written it. The pattern appeared across the tested models.
The difference remained after we expanded the comparison to the full set of usable cases, although it became smaller. GPT accepted 58% of the provider’s requests on provider-written contracts and 17% on client-written contracts. MiniMax accepted 75% and 21%. DeepSeek accepted 67% and 35%.
Changing the negotiation guidelines did not remove the effect. On the provider’s contract, acceptance remained at 83% when we removed the guidelines and when we replaced them with guidelines that did not fit the document. The starting contract itself was shaping the model’s position.
One possible explanation is that the client-written contracts were broader forms designed for general service providers, while the provider’s contracts were tailored to its work and risks. The models may have been responding to that difference, rather than simply to which party supplied the contract. Our tests did not separate those explanations.
R6. Models rejected most changes to core rights and responsibilities
Models were least willing to accept changes in these four areas:
| Type of change | Requests accepted |
|---|---|
| Which document controls when terms conflict | 0% |
| Ownership of work and intellectual property | 19% |
| Liability: who bears losses | 20% |
| Termination: when a party can end the agreement | 23% |
MiniMax accepted 33–47% of the liability, ownership, and termination changes, compared with 6–18% for GPT. All tested models rejected the document-priority changes in this sample.
Their explanations focused on the limits of the provider’s responsibility. An assessment does not guarantee an outcome, and a provider cannot control the client’s code. The models differed in how much risk they were willing to accept despite those concerns.
R7. The form of the request affected the answer
Models accepted 50% of requests shorter than about 83 characters. They accepted only 28–32% of longer requests. A short request may have been easier to understand, but it also gave the model fewer reasons to object.
They accepted 44% of requests that replaced existing language, compared with 27% of requests that added new language and 33% of requests that deleted language. This was the opposite of our initial expectation. Adding a new obligation drew more resistance than changing one already in the contract.
The subject also mattered. Models accepted 88% of publicity changes and 55% of fee changes, compared with 0–23% of changes involving document priority, ownership, liability, or termination.
R8. Accuracy and consistency were different
GPT was generally the most consistent family, but it was also much more resistant than the historical outcomes on some contracts. On the agreement with the 23-part anti-bribery annex, it accepted about 12% of the requests across repeated runs while the parties had accepted 76%.
MiniMax matched that 76% historical acceptance rate on the same agreement. That did not make it consistently accurate elsewhere: on another agreement, its acceptance rate moved from 38% to 62%, then 29% and 45% across four runs.
DeepSeek and Kimi were usually stable outside contracts with large blocks of similar provisions. On those blocks, one run could look entirely different from the next. A model could therefore be consistent but consistently too strict, or match the overall result in one matter while remaining unstable in another.
R9. Counterproposals did not move toward the final agreement
Among the counterproposals we could reliably compare, none was closer to the final agreed wording than to the original request. Three initially looked like matches, but manual review showed that they compared clauses about similar topics with different purposes.
For example, one apparent match paired two clauses about the same general subject even though they served different purposes. Once we compared the correct clauses, none of the model proposals reached the final compromise.
Models sometimes labeled a response COUNTER while keeping their preferred position largely intact. Proposing new wording did not mean they were approaching the compromise the parties reached.
R10. Models matched the final outcomes 39% of the time
Across 845 decisions on 293 requested changes, models matched the final negotiated outcomes 39% of the time. This comparison covered 23 cases and three model families.
A model could accept roughly the same proportion of changes as the parties did, yet accept different ones. Similar overall acceptance rates concealed disagreement on the individual decisions.
Confidence did not solve this problem. In the termination example, Grok rejected the change with 91% stated confidence, but the recorded outcome accepted it. Agreement among models also did not reliably identify the historical result. Two models rejected that clause and one accepted it; the one-model minority matched the recorded outcome.
R11. Models disagreed about which instructions controlled
Opus treated “should resist” as a position it could trade to close the deal. Grok and Kimi treated it as a restriction on what they could concede. All three understood the deadline; they disagreed about which concessions it permitted.
Fable 5 separately reviewed the existing ACCEPT label and called it AMBIGUOUS. It found reasons to accept in the commercial facts and reasons to reject or counter in the business comment. This was a review of an existing label, not another independent decision on the clause.
As the case study puts it: “A reason to make a concession is not the same as authority to make it.”
V. What this means for contract review
Understanding a change does not settle whether to accept it. The models could explain the risk and still disagree about whether the business should take it. They also reacted differently depending on who wrote the original contract.
To judge a model’s review, we need to examine which changes it accepts, what its proposed wording does, and whether it follows the business’s instructions. An overall acceptance rate or a plausible explanation cannot answer all three questions.
How the model families differed
Three model families had enough results for a broad comparison. The other three appeared only in narrower tests. The percentages below show how often a family accepted a requested change. They are not accuracy scores.
| Model family | Coverage | What we observed |
|---|---|---|
| GPT | Broad comparison | It accepted 29% of changes requested by clients, the lowest rate of the three broadly tested families. It was generally the most consistent family, but often consistently stricter than the historical result. It strongly favored requests made on the provider’s own contract and accepted only 6% of termination changes. |
| MiniMax | Broad comparison | It accepted 67% of changes requested by clients, the highest rate of the three. On one contested agreement, its 76% average exactly matched the historical acceptance rate. Its answers varied more on other contracts, and it remained comparatively willing to accept liability, ownership, and termination changes. |
| DeepSeek | Broad comparison | It accepted 45% of changes requested by clients and was the closest to the historical difference between provider-side and client-side requests: 11 percentage points, compared with nine points in the historical outcomes. It was strict on core risk terms and sometimes rejected large groups of similar provisions together. |
| Kimi K3 | Narrower tests | It rejected the mutual termination right because it treated “should resist” as a position it had to preserve. It also showed the same group-decision problem as DeepSeek on a long annex. A strict instruction to explain every provision changed its position; softer wording worked better. |
| Grok | Narrower tests | It also rejected the mutual termination right. In a separate test, it accepted provider-side requests at a rate 24 percentage points higher than client-side requests, compared with a nine-point difference in the historical outcomes. |
| Opus 5 | Single-clause test | It accepted the mutual termination right after weighing the instruction to resist against the goal of closing the deal. This result describes that clause only. |
The starting contract also mattered for the three broadly tested families. Across the complete comparison, GPT accepted 58% of the provider’s requests on provider-written contracts and 17% on client-written contracts. MiniMax accepted 75% and 21%. DeepSeek accepted 67% and 35%.
How the models were constructed may have affected their decisions
The split over “should resist” suggests two possible influences:
- Training documents. Examples of strict compliance with business comments might encourage a model to hold the position. Examples of commercial negotiation might encourage it to weigh the comment against the value of closing. Final contracts alone do not explain why a lawyer conceded a point or who authorized it.
- Feedback during training. A model might be rewarded for closely following a specific instruction, or for making a trade-off that serves a broader objective. Those preferences could affect how it reads an ambiguous comment.
These are hypotheses. The models’ explanations do not tell us what documents or feedback shaped them. The wording of the case could also explain the different readings.
Instructions need to work more than once
A prompt change that helps in one run may fail in the next. Test it repeatedly, across cases, and against unrelated additions to the prompt. The same instruction can also work differently across models.
Taking a majority vote across repeated runs may make decisions more consistent. It does not establish that they match the final negotiated outcome.
Be explicit about what can be conceded
If a business comment is mandatory, say so. If it is an opening position, say what the model can trade and under what conditions. That gives us a clearer basis for deciding whether the model followed instructions.
VI. Conclusion
The models usually understood what the proposed changes would do. Their reasoning differed from the recorded outcomes in what they did with that understanding. They often said that an assessment is advice, not a guarantee; that the provider does not control the client’s code; or that the provider should not bear losses it cannot prevent. Only two of roughly 17 sampled rejections simply called the term non-negotiable. Most gave a reason tied to the contract and the provider’s work.
The parties often accepted or modified those terms anyway. The final agreements do not tell us why, but the outcomes show that identifying a valid risk did not always end the negotiation. The parties may have accepted the risk because of price, leverage, timing, the value of the relationship, or a concession elsewhere. The models were more likely to treat the risk itself as enough reason to reject the change.
That helps explain the 39% match rate across 845 decisions. The models often understood the legal issue but did not make the same trade the parties made. Matching the parties’ overall acceptance rate did not mean choosing the same clauses, and high confidence or agreement among models did not reliably identify the recorded result.
The models had distinct negotiating patterns. GPT was usually consistent and strict. MiniMax accepted more changes and sometimes matched the parties’ overall result, but varied more across repeated runs. DeepSeek was closest to the historical difference between provider-side and client-side requests, but could reject a long group of similar provisions as one block. The narrower tests showed that Kimi, Grok, and Opus could read the same instruction and apply different rules to it.
The contract and the shape of the request influenced the answer. Models accepted more changes on a provider’s own form, accepted short requests more often than long ones, and resisted new language more than replacements. Changing the prompt did not reliably overcome those patterns. An unrelated sentence could move the result as much as a carefully written negotiating instruction.
The models also struggled with the act of negotiating. Their counterproposals preserved their original positions instead of moving toward the language the parties ultimately agreed. Understanding a clause, choosing what can be traded, and drafting the compromise are separate abilities. The models often showed the first. This study found much weaker evidence of the second and third.