Workflow · September 9, 2026
Draft a Copyright Litigation Risk Memo for AI Training Data Exposure
The task
You're in-house or outside counsel and your GC just forwarded another headline: The Seattle Times and Newsday are seeking an order requiring the destruction of copies of their copyrighted works, along with training datasets or AI models incorporating the material. Product and procurement want to know which of the company's AI vendors — and which internal training pipelines — sit in the blast radius. You need a privileged, structured risk memo fast.
Before AI
The manual version is a day-long slog: pull the vendor list from procurement, cross-reference each vendor's public disclosures on training data, dig up the MSA and DPA for each, map indemnity carve-outs, then draft a memo with a risk matrix and recommended redlines. Typical turnaround: 4-6 billable hours for a mid-sized vendor stack, longer if you're chasing signed contract copies.
The workflow
The prompts below run in sequence. Paste the first prompt plus your sample input into your firm-approved AI tool (one cleared for privileged work — see Gotchas). Feed each subsequent prompt into the same thread so it inherits context.
1. Triage the vendor list into exposure tiers
You are acting as a litigation risk analyst supporting outside counsel. This work is being performed at the direction of counsel in anticipation of litigation; treat the output as attorney work product and mark it accordingly. Below is a list of third-party AI vendors and internal AI use cases at a company. For EACH row, produce a triage table with these columns: 1. Vendor / Use case 2. Training data exposure tier (HIGH / MEDIUM / LOW) — HIGH if the vendor is known or credibly alleged to have trained on scraped web/news/book content; MEDIUM if training data provenance is opaque or mixed; LOW if the vendor uses only licensed, synthetic, or customer-provided data 3. Basis for tier (1 sentence — cite the public posture, e.g., "named defendant in NYT v. OpenAI," "publicly commits to licensed corpora only," "no public disclosure") 4. Downstream infringement channel (which of the company's products or workflows could surface allegedly infringing output) 5. Whether the company is likely a direct infringer, contributory infringer, or downstream user only After the table, list any vendor where you had to guess because the input was ambiguous. Do NOT invent facts about a vendor's training data — if unknown, say "provenance not disclosed" and tier as MEDIUM. Here is the input:
Company: Nordwell Financial Services (mid-size wealth manager, US) AI vendors and internal use cases under review: 1. OpenAI GPT-4o via Azure OpenAI — used in client-facing "Portfolio Insights" chat feature and internal research summarization. Contract: Microsoft EA + Azure OpenAI addendum, signed March 2025. 2. Anthropic Claude via direct API — used by legal ops team to summarize regulatory filings. Contract: Anthropic Commercial Terms, signed August 2025. 3. Perplexity Enterprise Pro — used by 40 analysts for market research. Contract: Perplexity MSA, signed January 2026, auto-renews Jan 2027. 4. Harvey AI — used by in-house legal for contract review. Contract: Harvey MSA, signed June 2025. 5. Internal model "NordGPT-1" — fine-tuned Llama 3.1 70B on a mix of (a) licensed Refinitiv news feed, (b) 12 years of internal research notes, (c) approximately 180GB of publicly scraped financial news articles collected in 2023 by a since-departed data science lead. No provenance log for the scraped set. 6. GitHub Copilot Business — used by 60 engineers. Contract: Microsoft Business Agreement, signed 2024. 7. Jasper AI — used by marketing for blog drafts. Contract: Jasper Business, signed February 2026. Context: The company was named in a subpoena last month in an unrelated matter and the GC wants a proactive risk read given the widening publisher copyright suits against foundation model providers.
2. Map indemnity coverage and gaps
Using the triage table you just produced, now build an INDEMNITY GAP MATRIX. For each vendor, produce a row with: 1. Vendor 2. Standard IP indemnity posture for this vendor as of your knowledge (e.g., "Microsoft Copilot Copyright Commitment covers eligible Azure OpenAI outputs subject to guardrail use," "Anthropic offers IP indemnity for paid API customers with standard carve-outs"). If you do not know the vendor's current public indemnity terms, write "REQUIRES CONFIRMATION — pull latest contract" and do not guess. 3. Key carve-outs that would likely defeat coverage in a publisher copyright suit (e.g., customer modification, disabling of safety filters, fine-tuning on customer data, use outside documented scope) 4. Cap on indemnity liability (state "unknown — confirm in signed agreement" if not public) 5. Whether the indemnity likely reaches the SPECIFIC theory in the current publisher wave: training-data infringement vs. output infringement vs. both 6. GAP SCORE: RED (no meaningful coverage), YELLOW (coverage exists but material carve-outs apply), GREEN (coverage likely reaches this risk) Then, for the internal fine-tuned model in the input, explain in 3-4 sentences why NO vendor indemnity applies and what first-party exposure the company holds directly.
3. Recommend protective clauses and next actions
Now draft the final memo. Format it as a privileged and confidential memorandum with these sections, in order: HEADER - To: General Counsel - From: [Counsel] - Re: AI Training Data Copyright Exposure — Risk Assessment and Recommended Actions - Privileged & Confidential — Attorney Work Product I. EXECUTIVE SUMMARY (5 bullets max, plain English, no hedging language a business reader can't act on) II. RISK LANDSCAPE (2 short paragraphs — reference the publisher litigation wave generally without citing specific case outcomes you cannot verify) III. EXPOSURE TABLE (reproduce the triage table from step 1) IV. INDEMNITY GAPS (reproduce the gap matrix from step 2) V. RECOMMENDED PROTECTIVE CLAUSES — for each RED or YELLOW vendor, propose specific redline language the company should push at renewal. Cover at minimum: - IP indemnity uncapped for third-party training-data claims, or carved out from the general liability cap - Vendor representation on training data provenance and legal basis - Duty to defend (not just indemnify) with counsel selection rights - Notice and cooperation obligations if vendor is sued over training data - Termination-for-convenience trigger if vendor is enjoined from using its current model - Data deletion / model unlearning cooperation if a court orders destruction of training sets VI. INTERNAL MODEL REMEDIATION — a numbered action list for the "NordGPT-1" style internal model: provenance audit, quarantine of unverified training data, decision memo on retraining vs. retirement, litigation hold considerations. VII. IMMEDIATE ACTIONS (next 30 days) — max 7 items, each with owner (Legal, Procurement, Engineering, or IT) and a specific deliverable. Keep the whole memo under 1,200 words. Use short paragraphs. Do not invent case citations. Where a legal conclusion depends on facts not in the input, flag it with "[CONFIRM]" so the reviewer catches it.
Gotchas
- Privilege hygiene. Only run this in an AI environment your firm or GC has cleared for privileged content — zero-retention API, no-training contractual terms, and access-controlled workspace. A consumer chatbot login will torch privilege claims.
- Model hallucinates indemnity terms. LLMs will confidently invent vendor contract language they've never seen. The prompt tells the model to say "REQUIRES CONFIRMATION" for anything it doesn't actually know — spot-check every GREEN rating against the signed agreement before it leaves your desk. Microsoft, Google, and Anthropic have all updated their IP indemnities more than once; treat any model-supplied version as a starting point, not an authority.
- Case law drift. The publisher wave is moving fast — new filings weekly, and cases like the Seattle Times and Newsday complaint keep shifting the threat model. Don't let the memo cite specific rulings without a human check.
- Internal fine-tunes are the real landmine. Vendor indemnities never reach models the company trained itself on scraped data. If your triage surfaces one, escalate that before you polish the redlines.
- Redline language is a starting draft. Every clause the model proposes needs a contracts partner's eye before it goes to the vendor. Treat it as a first pass, not final markup.
Time saved
A manual first-pass memo of this shape typically runs 4-6 billable hours: vendor inventory reconciliation, contract pulls, indemnity mapping, drafting, and formatting. This workflow compresses the mechanical work — triage, matrix building, and draft assembly — into roughly 30-45 minutes of prompt-and-review, leaving you the substantive hours for judgment calls and partner review. Conservative estimate: 3-4 hours saved per memo, more if you're repeating this across business units. The savings evaporate if you skip the verification pass on indemnity terms, so budget 20 minutes for that specifically.
What we got when we ran it
We executed every prompt above, in order, against the sample input (model: claude-sonnet-4-6). Outputs below are real and unedited apart from length truncation.
Step 1 output
# ATTORNEY WORK PRODUCT ## Privileged & Confidential — Prepared at Direction of Counsel in Anticipation of Litigation ### Nordwell Financial Services — AI Vendor Copyright Risk Triage Prepared for Outside Counsel | Work Product Protected Under FRE 502 & Attorney-Client Privilege --- ## I. TRIAGE TABLE | # | Vendor / Use Case | Training Data Exposure Tier | Basis for Tier | Downstream Infringement Channel (Nordwell-Specific) | Nordwell's Likely Liability Posture | |---|---|---|---|---|---| | 1 | OpenAI GPT-4o via Azure OpenAI — "Portfolio Insights" client chat; internal research summarization | HIGH | OpenAI is a named/primary defendant in The New York Times Co. v. Microsoft Corp. & OpenAI, No. 1:23-cv-11195 (S.D.N.Y.), and in multiple consolidated news publisher suits; publicly acknowledged training on broad internet corpora including Common Crawl and news content. | Client-facing Portfolio Insights chat could surface near-verbatim reproductions of news articles or analyst reports in client-visible output; internal research summarization could reproduce licensed or paywalled content in internal documents. | Downstream user / potential contributory infringer. Nordwell is not training the model but is deploying it in a commercial product. Contributory exposure depends on whether Nordwell had knowledge of infringement risk and continued to enable output distribution — the client-facing product is the highest-risk channel. Microsoft/OpenAI indemnity provisions in the Azure OpenAI addendum should be reviewed immediately for scope, carve-outs, and tender procedures. | | 2 | Anthropic Claude via Direct API — legal ops regulatory filing summarization | HIGH | Anthropic is a named defendant
…output truncated at 250 words.
Step 2 output
# ATTORNEY WORK PRODUCT ## Privileged & Confidential — Prepared at Direction of Counsel in Anticipation of Litigation ### Nordwell Financial Services — AI Vendor Indemnity Gap Matrix Prepared for Outside Counsel | Work Product Protected Under FRE 502 & Attorney-Client Privilege --- ## I. INDEMNITY GAP MATRIX | # | Vendor | Standard IP Indemnity Posture (as of knowledge cutoff) | Key Carve-Outs Likely to Defeat Coverage in Publisher Suit | Indemnity Liability Cap | Reaches Training-Data Theory? / Output Theory? | GAP SCORE | |---|---|---|---|---|---|---| | 1 | OpenAI GPT-4o via Azure OpenAI (Microsoft EA + Azure OpenAI Addendum) | Microsoft's Copilot Copyright Commitment (announced Sept. 2023) extends to Azure OpenAI Service customers: Microsoft will defend and indemnify eligible customers against third-party IP infringement claims arising from outputs of covered models, provided the customer uses the product as documented and has not disabled built-in safety features. The commitment is Microsoft's own undertaking, not OpenAI's — OpenAI has no direct privity with Nordwell. | (a) Customer must use the service within documented parameters — any customization of system prompts designed to suppress safety behavior or extract memorized content likely voids coverage; (b) Customer must not have "directed" the model to produce the infringing output (i.e., if a Nordwell system prompt specifically requests reproduction of a named publication's content, coverage likely fails); (c) Fine-tuning or model modification by customer (not applicable here, but confirm no AOAI fine-tuning features are in use); (d) The Commitment is widely understood to cover output
…output truncated at 250 words.
Step 3 output
# ATTORNEY WORK PRODUCT ## PRIVILEGED AND CONFIDENTIAL — ATTORNEY-CLIENT COMMUNICATION --- MEMORANDUM To: General Counsel, Nordwell Financial Services From: Outside Counsel Date: [DATE] Re: AI Training Data Copyright Exposure — Risk Assessment and Recommended Actions --- ## I. EXECUTIVE SUMMARY - NordGPT-1 is Nordwell's single largest legal exposure. Nordwell trained an internal model on approximately 180GB of scraped news content with no provenance log and no authorization documentation. There is no vendor indemnity. Nordwell is the infringer. Stop new deployments now. - Every third-party vendor indemnity in place has the same structural gap: vendor IP indemnities cover output-side claims but do not cover the training-data infringement theory — which is precisely the theory driving the current wave of publisher litigation. Do not rely on these indemnities as a shield against the cases being filed today. - Perplexity carries the highest third-party vendor risk. Its business model is the subject of active publisher suits, no indemnity is confirmed in the signed contract, and analyst outputs are flowing into client-facing materials. Use should be suspended or strictly quarantined pending contract review. - GitHub Copilot coverage is contingent on a single IT setting. If the duplication detection filter is not enabled for all 60 engineers, Microsoft's Copyright Commitment does not apply. This takes one day to verify and fix. - Four vendor contracts require immediate review. The Perplexity, Harvey, Jasper, and Anthropic agreements must be pulled and reviewed against the indemnity requirements in Section V. Renewal cycles for all must be flagged
…output truncated at 250 words.
---
This content is for informational purposes only and is not legal advice. Confirm confidentiality, privilege, and jurisdictional rules before using any AI tool with client matters.
Source: techcrunch.com
More for Legal professionals →
Get the next one in your inbox