AI Contract Review: I Tested ChatGPT Against a Contract Playbook

PRACTICAL AI EXPERIMENT

Vendor contract review can involve a surprising amount of repetitive comparison work. A reviewer reads a clause, checks the company’s contract playbook, decides whether the position is acceptable, and then moves on to the next clause. Multiply that across a long agreement and a steady stream of vendors, and a fairly straightforward task starts consuming a lot of time.

The comparison also gets more complicated once the playbook contains fallback positions. A company may prefer Net 45 payment terms while still accepting Net 30. Texas may be the preferred governing law while Delaware remains acceptable. Some rules depend on several conditions being satisfied at the same time, and occasionally the agreement simply does not contain enough information to reach a conclusion.

I wanted to see how well ChatGPT could handle that type of review using a simple internal contract playbook and a vendor agreement. I created synthetic documents, deliberately mixed acceptable and unacceptable positions, and ran two rounds without telling ChatGPT which clauses I expected it to flag.

I kept the setup similar to my accounts payable experiment. There was no contract-management platform, API integration or custom application involved. I uploaded an Excel playbook and a PDF agreement, provided the review instructions, and compared the output with answers I already knew.

Download the contract review experiment

The package contains the two synthetic contract playbooks, two vendor agreements and the exact prompts used for both rounds. If you want to reproduce the test while reading, download the files and start a fresh ChatGPT conversation for each round.


Download the AI Contract Review Experiment

The agreements and playbooks are fictional and provided for experimentation. This article is not legal advice and the example should not be treated as a substitute for qualified legal review.

Problem Statement

Many procurement and legal teams review vendor agreements against an internal set of contracting positions. Those rules might exist in a formal contract playbook, a spreadsheet, internal guidance or some combination of the three. The playbook usually describes what the company prefers, what it can accept without escalation and what should receive additional review.

That structure creates a good comparison problem for an AI experiment. The agreement provides the vendor’s language, while the playbook provides the business rules that should be applied to it. The useful output is a short review showing which clauses fit the approved position, which need attention and which cannot be evaluated from the information available.

The challenge comes from the details. A clause can differ from the preferred wording and still be perfectly acceptable. A rule can depend on a percentage, a notice period and a frequency limit all at once. A limitation-of-liability clause may contain the expected cap while the carve-outs materially change the exposure. A review process that ignores those details will create unnecessary exceptions or miss the ones that matter.

I therefore asked ChatGPT to use only three statuses: Match, Review and Cannot Verify. I also required evidence from both the playbook and agreement so that a human reviewer could see how the model reached each conclusion.

Plan of Approach

I ran the experiment in two rounds. Round 1 tested relatively straightforward contract positions along with one deliberately incomplete clause. This gave me a baseline for whether ChatGPT could match the correct clauses, identify obvious deviations and refrain from filling in missing information.

Round 2 introduced more realistic playbook logic. Several preferred positions had approved fallbacks, and other rules depended on multiple conditions being evaluated together. This round was designed to see whether ChatGPT would follow the actual business rules rather than flagging every difference from the preferred position.

Each round contained eight playbook topics, giving me sixteen controlled checks in total. Because I created both the playbooks and agreements, I had an expected result for every topic before uploading the documents.

I also told ChatGPT not to approve or reject the agreement and not to provide legal advice. The task was limited to preparing a structured business review that procurement or legal could evaluate.

Round 1: Applying a basic contract playbook

The first playbook covered payment terms, auto-renewal, termination for convenience, confidentiality, limitation of liability, data-breach notification, governing law and intellectual property.

The synthetic agreement contained a mixture of clauses that fit the playbook and clauses that clearly fell outside it. I also included a liability provision referring to undefined “excluded claims.” The base liability cap looked acceptable, but the supplied agreement did not explain which claims sat outside that cap.


Round 1 AI contract review setup with vendor agreement PDF and contract playbook spreadsheet uploaded to ChatGPT
Round 1 setup with the synthetic vendor agreement and internal contract playbook.
Prompt used for Round 1
I work in procurement and am reviewing a vendor agreement against our internal contract playbook.

Compare the vendor agreement with the contract playbook.

For every topic in the playbook, determine whether the agreement follows the preferred or acceptable position.

Use only these statuses:

- Match
- Review
- Cannot Verify

Do not approve or reject the contract.
Do not provide legal advice.
Do not invent terms that are not in the supplied documents.

Return a simple table with:

- Topic
- Agreement clause
- Status
- What needs attention
- Evidence from the playbook and agreement
- Suggested business change for legal or procurement review

If the playbook explicitly allows an alternative or fallback position, treat that as acceptable rather than automatically flagging it.

If the agreement does not contain enough information to verify a playbook requirement, say Cannot Verify.

Show me only what a human contract reviewer would need to know.

What happened in Round 1

The payment clause required payment within 15 days. The playbook preferred Net 45 and allowed Net 30, so ChatGPT correctly marked the clause for Review and suggested changing the term to Net 45 or at least Net 30.

The auto-renewal clause also received a Review. The twelve-month renewal period was within the allowed limit, but the agreement required 90 days of notice to prevent renewal while the playbook allowed a notice period of 30 days or less.

The termination-for-convenience and confidentiality clauses were accepted. The agreement provided the required 30-day customer termination right, while the five-year confidentiality period exceeded the playbook’s minimum requirement of three years.

The limitation-of-liability result was the one I was most curious about. The agreement contained the expected mutual cap based on twelve months of fees, followed by an exception for unspecified “excluded claims.” Since those claims were not defined in the supplied agreement, ChatGPT returned Cannot Verify and explained what information was missing.

The remaining results followed the expected playbook positions. ChatGPT flagged the ten-business-day security incident notification period against the 72-hour maximum, accepted Texas governing law and identified the intellectual-property provision for review because the vendor retained ownership of custom deliverables that the playbook assigned to the customer after payment.


Round 1 AI contract review results comparing a vendor agreement with an internal contract playbook
Round 1 output showing matches, review items and one clause that could not be verified from the supplied documents.
Topic Expected ChatGPT
Payment terms Review Correct
Auto-renewal Review Correct
Termination for convenience Match Correct
Confidentiality Match Correct
Limitation of liability Cannot Verify Correct
Data breach notification Review Correct
Governing law Match Correct
Intellectual property Review Correct

Round 1 finished with all eight topics classified as expected. The Cannot Verify result was particularly useful because the model had enough information to recognize the baseline liability cap while also recognizing that the undefined exclusions prevented a complete assessment.

Round 2: Testing fallback positions and multi-part rules

The second playbook was closer to how I would expect an actual contracting guide to behave. Several topics included a preferred position alongside one or more acceptable alternatives.

Payment terms provided a simple example. Net 45 was preferred, but Net 30 was explicitly acceptable. Governing law worked the same way, with Texas preferred and New York or Delaware permitted. An automated review should be able to distinguish an acceptable negotiated position from one that genuinely requires escalation.

I also included rules with several moving parts. Auto-renewal depended on both the renewal period and the notice window. Annual price increases had limits for percentage, frequency and advance notice. The limitation-of-liability rule required looking at the base cap and the way carve-outs applied to each party.


Round 2 AI contract review setup testing fallback positions and multi-condition contract playbook rules
Round 2 used a playbook with acceptable fallback positions and rules containing several conditions.
Prompt used for Round 2
I work in procurement and am reviewing a vendor agreement against our internal contract playbook.

Compare the vendor agreement with the contract playbook.

For every topic in the playbook, determine whether the agreement follows the preferred or acceptable position.

Use only these statuses:

- Match
- Review
- Cannot Verify

Do not approve or reject the contract.
Do not provide legal advice.
Do not invent terms that are not in the supplied documents.

Return a simple table with:

- Topic
- Agreement clause
- Status
- What needs attention
- Evidence from the playbook and agreement
- Suggested business change for legal or procurement review

Read the playbook carefully before deciding that a difference is an issue.

If the playbook allows a preferred position and an acceptable fallback position, treat either as acceptable.

If a playbook rule depends on more than one condition, evaluate all of those conditions together before assigning a status.

If the agreement uses different wording but satisfies the same business requirement, do not flag it merely because the wording is different.

If the agreement does not contain enough information to verify a requirement, say Cannot Verify.

Show me only what a human contract reviewer would need to know.

It recognized the acceptable fallback positions

The agreement used Net 30 payment terms. Since the playbook explicitly allowed Net 30, ChatGPT returned Match even though Net 45 was listed as the preferred position.

The governing-law provision produced the same behavior. The agreement selected Delaware, which appeared in the playbook as an acceptable alternative to the preferred Texas position. ChatGPT correctly left it alone.

This is an important part of the experiment because a review process that treats every departure from the preferred template as an exception can create a lot of unnecessary work. The approved fallback positions need to be part of the decision.

The rules with several conditions were handled well

The auto-renewal clause used one-year renewal periods and allowed either party to prevent renewal with 45 days of notice. Both conditions fit the playbook, so ChatGPT returned Match.

The annual price-increase provision required three checks. The agreement capped the increase at four percent, permitted an increase only once in a twelve-month period and required at least 90 days of written notice. The playbook allowed up to five percent, no more than one increase per year and at least 60 days of notice. ChatGPT evaluated all three pieces and accepted the clause.

The security incident provision also passed. Its wording differed slightly from the playbook, but both documents required notification without undue delay and within 72 hours after confirmation of an incident affecting customer data.

The intellectual-property clause followed the same pattern. The customer received ownership of paid-for custom deliverables while the vendor retained its pre-existing materials and general know-how, which was the allocation permitted by the playbook.

Two clauses still needed attention

The limitation-of-liability provision began with a mutual cap based on twelve months of fees, which matched the playbook baseline. The carve-outs changed the result because only the customer’s confidentiality, data-security and indemnification obligations were excluded from the cap. ChatGPT recognized the customer-only uncapped exposure and returned Review.

The data return and deletion clause also received a Review. The agreement made customer data available for export for 30 days after termination, but it allowed the vendor to retain that data for up to 180 days before deletion. The playbook required deletion within 30 days unless longer retention was tied to a documented legal obligation.

I thought this was a good test because both documents contained the phrase “30 days,” yet the obligation attached to those 30 days was different. One addressed availability for export and the other addressed actual deletion. ChatGPT picked up the distinction.


Round 2 AI contract review results showing acceptable fallbacks, multi-condition rules and clauses needing review
Round 2 output showing acceptable fallback positions alongside the liability and data-deletion issues that required review.

Results from both rounds

Across the two rounds I tested sixteen playbook topics. ChatGPT returned the classification I expected for all sixteen in this controlled synthetic experiment.

Measure Result
Playbook topics tested 16
Expected classifications returned 16
Missed planted review items 0
False-positive review classifications 0
Cannot Verify cases handled as expected 1
Experiment rounds 2

I would be careful about reading too much into a perfect score on sixteen checks. I designed these documents myself, the agreements were short, the playbook rules were explicit and I knew what conditions I wanted to test. The result tells me that the approach deserves further testing. It does not establish an accuracy rate for contract review in general.

Real agreements introduce far more complexity. Definitions can affect clauses many pages later, amendments can override earlier language, schedules may add new obligations, and different contract types may use different playbooks or escalation thresholds. Those are exactly the kinds of complications I would introduce in a later test.

How I interpret the result: ChatGPT handled this small set of explicit commercial rules very well, including approved fallback positions, multi-condition rules and one situation where the source documents were incomplete. That is enough for me to see potential in using AI to prepare a first-pass contract review, provided the output remains grounded in the company’s actual playbook and supporting contract language.

Where I think this could actually help

I can see this being useful as a first-pass procurement review before an agreement reaches someone who needs to spend time evaluating every clause manually. The output could separate positions that already fit the playbook from the clauses that genuinely deserve attention and the items where additional information is required.

A reviewer could then concentrate on the smaller set of issues that require judgment or negotiation. The evidence column is important here because it gives the reviewer the relevant playbook requirement and contract language together. Without that evidence, the person receiving the output would still need to reconstruct the comparison from the beginning.

I would also keep the three-status approach. Match, Review and Cannot Verify are simple enough that the output remains useful without pretending the model has made the final contracting decision.

This is the kind of bounded test I had in mind when I wrote about running an AI pilot with guardrails. You can define the expected outcome, deliberately introduce difficult cases, observe the mistakes and adjust the process before giving the system a larger role.

Would I automate the contract-review workflow?

I would test harder agreements before adding more automation. The current experiment answered the basic question I cared about: whether ChatGPT could apply an explicit playbook to a vendor agreement with enough discipline to distinguish approved positions, review items and missing information.

If later tests continued to perform well, an operational version could retrieve a vendor agreement, select the appropriate playbook, prepare the comparison and send the result to procurement or legal. That version would need playbook version control, reliable source citations, access controls, rules for amendments and conflicting documents, and an audit trail showing what the system evaluated.

Those are worthwhile problems, but they belong to a later stage. Someone reading this experiment can reproduce the current workflow in a few minutes and decide whether the idea is relevant to their own work before thinking about any of that infrastructure.

Before using this with real contracts

This experiment uses fictional agreements and a synthetic contract playbook. It is not legal advice, and the example workflow should not be used as a substitute for review by qualified legal or procurement professionals.

Contract playbooks vary between organizations and may depend on contract type, jurisdiction, transaction value, data sensitivity, regulatory requirements, insurance coverage, risk appetite and escalation authority. If you test an approach like this, use your own approved playbook and compare the results against agreements your team has already reviewed.

Pay attention to missed issues as well as false positives, require evidence for every conclusion, define what should happen when information is incomplete, and follow your organization’s policies for handling confidential agreements and vendor information.

Contract approval, legal judgment, negotiation decisions and formal acceptance should remain with the people authorized to make those decisions.

Key Learnings

  1. The playbook needs to distinguish preferred positions from acceptable fallbacks. Round 2 showed how easily a valid negotiated position could otherwise become an unnecessary exception.
  2. Rules with several conditions need to be evaluated as a whole. Auto-renewal, price increases and limitation of liability all depended on more than one part of the clause.
  3. Cannot Verify is a useful outcome. The incomplete liability clause in Round 1 showed why the system needs permission to stop when the documents do not support a conclusion.
  4. Evidence makes the review easier to trust and challenge. Showing the relevant playbook position beside the agreement language gives the reviewer a clear starting point.
  5. The quality of the playbook directly affects the quality of the review. Clear fallback positions, thresholds and escalation rules give the model something concrete to apply. Vague internal guidance would make consistent results much harder.
  6. Controlled experiments are useful before automation enters the picture. Knowing the expected result for every clause made it easy to see how the model handled the edge cases instead of being impressed by a polished-looking table.

If I were taking this experiment further, I would make the contracts harder before adding more technology around the process. I would test longer agreements, defined terms that affect clauses elsewhere in the document, amendments, schedules, conflicting provisions and playbook rules that change depending on the type or value of the contract. That would tell me much more about where this approach starts to struggle than immediately turning the current prompt into an automated workflow.

For now, I am comfortable with what this small experiment showed me. Across sixteen controlled checks, ChatGPT applied the playbook the way I expected, including acceptable fallback positions, rules with multiple conditions and one situation where the documents did not contain enough information to reach a conclusion. I would still want a much tougher test before using this around real contracts, but as a way to prepare a first-pass procurement review and point a human reviewer toward the clauses that deserve attention, this was considerably more useful than I expected.

Leave a Comment

Your email address will not be published. Required fields are marked *