Ethical AI Due Diligence: PE Checklist

published on 29 September 2026

If I cannot confirm data rights, human review, vendor control, and production monitoring, I should treat the AI system as a deal risk - not a feature.

For a PE buyer, the issue is simple: can the target keep using its AI after close without legal, security, or cost problems? This checklist says to answer that in 4 passes - Govern, Map, Measure, Manage - and tie every finding to a deal action: accept, fix, protect, or stop.

Before I get into the detail, here is the short version:

  • Start with inventory - if a tool is not listed, owned, and mapped to data and vendors, it is a risk item.
  • Test data rights first - possession of data is not the same as permission to use it for training, prompts, retrieval, or automated decisions.
  • Check third-party terms - AI value can break fast if a vendor can change model behavior, retain data, limit exports, or cut access.
  • Look at production proof, not demos - drift, error rates, bias checks, logging, and rollback matter more than slide decks.
  • Verify human override - for credit, hiring, insurance, medical, and fraud workflows, there should be a documented path to stop or review outputs.
  • Turn findings into deal terms - some issues fit post-close fixes, while others need repricing, indemnities, escrow, or a closing condition.

I would use 4 labels across each AI use case:

  • Acceptable - evidence supports use as-is
  • Acceptable with conditions - usable if stated controls or contract fixes are completed
  • Requires remediation - material gap, but fixable with owner, budget, and deadline
  • Transaction-critical risk - open issue can affect closing, value, liability, or day-1 use

The article’s main point is clear: AI diligence is not a policy review. It is a use-rights, control, and continuity review tied to EBITDA, liability, and exit value.

AI Due Diligence Framework for PE Buyers: Govern, Map, Measure, Manage

AI Due Diligence Framework for PE Buyers: Govern, Map, Measure, Manage

AI in M&A: Due diligence and the new reps and warranties | ITCC

Govern and map: inventory every AI system, owner, and data dependency

Start here. The Govern and Map pass is how you find hidden AI use before you test controls or output quality. If a system is not on the map, you cannot judge its risk.

AI inventory, governance policy, and oversight records

Ask for a current AI and automation register that covers production applications, pilots, proofs of concept, employee experiments, embedded AI features in SaaS products, foundation-model APIs, automated decision tools, and agentic systems. For each item, capture the system name and version, business purpose, deployment status, business unit, system owner, technical owner, vendor or model provider, users and affected individuals, input and output data, deployment location, applicable jurisdictions, risk tier, approval date, last review date, and decommissioning status.

Do not rely on the register alone. Reconcile at least 4 sources: the AI register, technical architecture diagrams and logs, vendor and procurement records, and user or engineer interviews. A common red flag is a register that includes only internally built models and leaves out undeclared tools such as Microsoft 365 Copilot, customer-service chatbots, résumé-screening software, or marketing-generation tools. Those may qualify as shadow AI. Any unmapped system is a transaction-risk item until ownership and data flow are verified.

Request the AI policy and related standards for procurement, development, validation, deployment, acceptable use, human oversight, privacy, security, intellectual-property review, records retention, incident response, vendor management, and system retirement. Then test whether the policy works in practice. Ask for AI steering-committee minutes, risk assessments, approval tickets, internal-audit reports, training records, exception registers, incident logs, model-change approvals, board or audit-committee reporting, and management certifications.

Pin down who can approve high-risk systems and who can suspend them. Useful evidence includes a governance charter naming the responsible executive, a RACI matrix, documented escalation timeframes, and reports that show open issues and overdue reviews. Red flags include a policy approved after systems went live, no named accountable owner, governance limited to legal review, repeated policy exceptions with no expiry dates, or board materials that discuss AI strategy but do not identify specific systems, risks, incidents, or remediation actions.

Sample high-risk systems and compare the recorded model version, prompts, retrieval sources, connected tools, user permissions, data flows, human approvals, and output destinations to the production setup. Pull samples from expense reports, browser or identity logs, cloud API bills, and software contracts, then match them to the formal register. If an unregistered system handles personal, confidential, regulated, or customer data, treat it as a control failure until ownership, purpose, permissions, and retention are documented.

Data rights, privacy, and chain of custody

After the systems are mapped, test whether the company has the right to use the underlying data.

Request a data-rights file for training, validation, fine-tuning, retrieval, prompts, and production inputs and outputs. It should identify the dataset or source, originating party, owner or licensor, collection method, applicable license, permitted AI uses, personal-data status, sensitive-data categories, geographic scope, retention period, deletion requirements, downstream sharing restrictions, and evidence of authorization.

Ask for data inventories, privacy notices, consent language, customer and employee agreements, data-processing agreements, vendor terms, data-use addenda, copyright or database licenses, records of deletion requests, and documented exceptions. Keep possession rights separate from use rights for training, fine-tuning, profiling, automated decisions, and vendor disclosure.

Trace 1 representative record from collection through deletion: transformation, labeling, training or retrieval, inference, logging, human review, storage, and deletion. Confirm whether prompts, uploaded files, embeddings, logs, feedback, and outputs are stored, whether they are used to train a shared vendor model, where they are processed, who can access them, and how they are deleted. Verify whether the target can isolate 1 customer’s data from another customer’s data and respond to deletion, access, correction, or contractual-use restrictions. If rights are unclear, the system is not yet usable for post-close operations.

High-priority red flags are plain. One is employees entering customer records, source code, health information, financial information, or trade secrets into consumer versions of AI tools without an approved enterprise agreement. Another is a vendor promise not to train on customer data while still allowing indefinite logging, human review, or use by subprocessors under separate terms.

The FTC has warned that AI providers may face liability when they fail to honor privacy and confidentiality commitments made to users or customers [3]. It has also warned that quietly expanding data use - for example, using consumer data for AI training through a retroactive terms-of-service change - may be unfair or deceptive [4].

Use a data-rights matrix to track each material dataset or data flow. Classify each as Cleared, Restricted, Uncertain, or Prohibited pending remediation, and assign an owner and deadline to every open item.

Dataset or flow Owner or rights holder Permitted AI use Personal data? Geographic scope Retention/deletion obligation Evidence reviewed Remediation status
Customer support tickets Customers; target as service provider or controller, depending on contract Service delivery; training only if expressly permitted Yes; may include sensitive content United States; verify any offshore processing Contractual deletion and legal-hold rules Customer agreement, privacy notice, DPA Open - training permission unclear
Licensed industry database Third-party licensor Search and internal analytics; no model training unless licensed Possibly United States and permitted territories License term and deletion at expiry License and vendor terms Cleared or escalate to counsel
Employee performance records Target or employer, subject to employment/privacy rules Limited HR administration; no unrestricted foundation-model upload Yes Relevant state and country restrictions Retention schedule and access controls HR notice, policy, access logs Open - purpose limitation review
Public web content Varies by source and license Review source terms and applicable rights before training May contain personal data Global source footprint Source-specific obligations Collection method, terms, takedown process Open - rights and provenance incomplete

A missing license, unclear consent, undocumented cross-border transfer, or inability to honor deletion requests needs an owner and a deadline. Do not label it as a mere data-quality issue.

Third-party AI tools and concentration risk

Apply the same rights and control review to third-party tools and embedded AI features.

Create a vendor and dependency register tied to each internal use case. Group the fields by category:

  • Contract terms: provider, model or feature, contract entity, indemnities, intellectual-property allocation
  • Hosting and security: hosting region, subprocessors, security certifications, breach-notification terms, audit rights
  • Data sharing: data categories shared, model version
  • Change and exit: change-notification process, exit notice period, termination assistance, export capability
  • Fallback: outage history and substitute options

Include AI features embedded in enterprise software even when the target did not buy them as separate products. Map retrieval databases, vector stores, orchestration frameworks, plug-ins, cloud-hosted models, labeling firms, managed-service providers, and consultants. NIST guidance recommends monitoring risks from third-party resources and documenting controls applied to them, and it notes that third-party integrations can increase intellectual-property, privacy, and information-security risks, particularly where external systems receive model inputs or outputs [2][1].

Concentration risk shows up when several material workflows depend on the same model provider, cloud region, identity platform, or data broker. Ask what happens if the provider changes model behavior, increases prices, has an outage, restricts a use case, loses a license, or terminates the account. A single-provider dependency without a fallback is a continuity risk, not a paperwork issue. If a critical AI process has no tested fallback, no exportable data, no replaceable provider, or no exit notice period, that is deal-impacting.

Only mapped, owned, and contractually usable systems should move into performance and control testing.

Measure: test model performance, drift, bias, and cyber controls

The point here is simple: you need proof that the AI performs in production and that the company can spot failure when it happens. Policies, model governance decks, and vendor certifications may help, but they do not answer the operating question. Operating evidence does.

Performance baselines, drift monitoring, and human override

Start with the record for each material AI system. Request the model card, or the closest equivalent, plus the intended-use statement, training and validation datasets, feature definitions, target labels, evaluation code, confusion matrices, calibration results, and production performance reports. You also want metrics tied to the decision the model is making: precision, recall, F1, error rates, AUC, calibration, latency, availability, and cost per transaction.

Do not overweight training-set accuracy. Out-of-sample testing and time-based testing matter more. If a claims model was trained on one period, test it on a later period. Then test it across product lines, geographies, and customer segments. A model that looks fine on its training data can still fail when the mix shifts in production. Compare live results to the approved baseline and investigate any drop in performance.

Drift is where many teams get caught flat-footed. Data drift changes the inputs. Concept drift changes the link between the inputs and the outcome. Either can weaken results without an obvious warning. Ask for monitoring dashboards, alert thresholds, alert history, retraining records, and incident tickets. Thresholds should trigger specific actions, not just sit in a dashboard. For example, a 10% relative drop may require investigation, a 20% drop may require suspension of automated decisions, and any material change to the data source, model, prompt, or business process may require revalidation.

Check the release controls too. Verify rollback to the last approved model or prompt version, shadow testing before release, approval records, and revalidation after any major product, market, legal, or data change.

For high-risk decisions - credit, employment, insurance, medical triage, and fraud escalation - confirm that a documented manual-review path exists and that the AI cannot act without human approval. A good diligence finding is not “accuracy is 92%.” It is more concrete than that: false negatives moved from 3% to 8% in a high-value customer segment, and there is no override or escalation record.

Baseline metric Current metric Tolerance Breach status Business impact Required action
False-negative rate 3.0% ≤5.0% Pass Acceptable missed-risk exposure Continue monitoring
False-negative rate in high-risk segment 8.0% ≤5.0% Breach Potential customer and loss exposure Suspend automation; conduct retrospective review
Selection-rate ratio for protected group 0.72 ≥0.80 screening indicator Breach Potential employment discrimination exposure Legal review; validate, redesign, or discontinue
Hallucination rate on approved test set 4.0% ≤2.0% Breach Incorrect customer or employee guidance Add retrieval controls and mandatory human review
Critical agent actions requiring approval 100% 100% Pass Limits unauthorized impact Maintain control
Critical agent actions logged 70% 100% Breach Inability to investigate or prove authorization Block deployment until logging is complete

Rollback should be tested, not assumed. Verify that the company can disable the model, revert to the last approved version, route work to humans, revoke agent credentials, and preserve logs.

Bias, explainability, and complaint handling

If model results differ by group, treat that as a legal and operating issue. It is not just a reporting footnote. Test selection or approval rates, false-positive and false-negative rates, calibration, ranking quality, and rejection or escalation rates by group. Where sample sizes allow, test intersectional groups as well. Report confidence intervals or statistical significance instead of leaning on a single percentage. In employment use cases, the four-fifths rule is a common screening indicator: a group's selection rate below 80% of the highest group's rate signals possible adverse impact, though sample size and statistical significance also matter.[5][6]

When you find a disparity, push past the headline and look for the cause. The problem may sit in biased labels, proxy variables, data imbalance, measurement error, workflow design, or threshold setting. The response may be better training data, model redesign, adjusted thresholds, added human review, a different tool, or stopping use of the tool altogether.

The risk gets worse when affected users cannot challenge the outcome or understand why it happened. Test whether people receive a clear, decision-relevant explanation rather than a technical feature-importance output. Then test whether they have a real appeal route. Sample complaints and appeals to measure response time, reversal rate, and whether repeat errors reach model owners. Also test reviewers with intentionally wrong outputs. That tells you whether staff are doing independent review or just clicking approve. A manual-review step is weak if people rubber-stamp the model.

Group Metric Threshold Observed disparity Impact Corrective action
Protected class A vs. reference group Selection rate ratio ≥0.80 0.72 Potential adverse impact under employment law Legal review; validate or redesign model
Non-English speakers False-negative rate ≤5.0% Above threshold Unequal service outcomes Retrain on representative data; add human review
Users with disabilities Screen-reader support Full support Partial support only Accessibility barrier Remediate UI; retest with assistive technology

Cybersecurity controls for AI systems and agents

AI security review is about unsafe actions, not just unsafe code. The board-level question is whether the company can stop the system before it creates liability. Request current architecture and data-flow diagrams, threat models, penetration-test reports, vulnerability scans, remediation evidence, access-control matrices, privileged-access logs, encryption settings, key-management procedures, and vendor security assessments. You want proof that controls operate over time, not just a certificate in a data room.

With agentic AI, the diligence lens changes. The issue is not only whether the model gives a correct answer. The issue is whether it can make an unauthorized state change, access too much data, or chain tools in ways no one expected. Run red-team tests for direct and indirect prompt injection, including malicious instructions hidden in web pages, documents, emails, database records, or retrieved content. NIST's adversarial-machine-learning taxonomy identifies prompt injection as malicious instructions intended to manipulate generative AI behavior.[7]

Map every tool and permission the agent has. Then test whether untrusted input can trigger email transmission, file deletion, financial transactions, code execution, privilege escalation, or data exfiltration. Controls should cover input and output filtering, allowlisted tools, scoped credentials, confirmation gates for consequential actions, sandboxing, rate limits, and an emergency shutdown mechanism.

For retrieval-augmented generation systems, treat retrieved documents as untrusted input. Controls should stop retrieved text from overriding system or developer instructions, and the company should preserve document provenance and version history.[8][9] Also confirm there is an AI-specific incident-response playbook that covers prompt injection, data leakage, poisoning, model theft, unsafe tool execution, harmful outputs, bias complaints, vendor outages, and emergency shutdown. Then verify that tabletop exercises or red-team tests have actually been run.

Record each failure for the Manage pass as a remediation item, a term, or a closing condition.

Manage: classify compliance gaps, deal risk, and post-close fixes

Start by sorting each finding into 1 of 3 buckets: current violation, control gap, or post-close fix. Then route each item using the buyer labels from the introduction - price, protection, or post-close remediation. That keeps the diligence record tied to the deal model instead of leaving AI issues as a side memo.

Compliance exposure, contractual risk, and insurance review

Use the use-case file to confirm rights, notices, logs, and obligations. For each material AI use case, map the laws and commitments that apply against what the company is doing in practice. That includes federal and state privacy laws, sector rules, consumer-protection duties, employment laws, copyright and trade-secret issues, export-control limits, and customer or supplier commitments. Also log any open regulator inquiries, litigation, employee complaints, customer claims, and internal escalations. Check whether statements made to customers, employees, regulators, and investors match actual system behavior. If a document is missing, do not assume there is a violation. Record it as a control gap and assign an owner, deadline, and evidence requirement.

The FTC states that it enforces federal competition and consumer-protection laws against anticompetitive, deceptive, and unfair business practices involving AI [12], and FTC materials identify AI-generated fake reviews and testimonials as prohibited conduct under the agency's finalized rule on fake reviews and testimonials [11].

After you map the legal exposure, test whether each vendor contract supports lawful use after closing. Run a linked-contract review for each material AI vendor. Confirm whether the target can lawfully use the vendor's inputs and outputs, and whether the vendor can change the model, training practices, hosting location, or subprocessors without meaningful notice. Focus first on the contract points that can hurt continuity or leave the buyer exposed:

  • narrow or excluded indemnities
  • liability caps that are too low
  • no audit rights or incident-notification rights
  • restrictions that block migration of data, models, prompts, logs, or customer configurations after closing

Put a U.S. dollar value on replacement or renegotiation costs. Then decide whether the relationship needs pre-close consent, a transition-services arrangement, or a post-close contract remediation project.

Next, test whether the insurance program covers the AI risks you found. Compare the company's actual AI activity against the wording, exclusions, sublimits, retentions, and notice requirements in its cyber, technology errors and omissions, employment practices liability, directors and officers, and media liability policies. Ask direct questions: Does coverage apply to biased or discriminatory automated decisions? Copyright or other intellectual-property claims? Errors in AI-generated advice or content? Acts by vendors or autonomous agents? Regulatory investigations or fines? Pull loss runs, open claims, reservation-of-rights letters, and any exclusions discussed during renewal. AI losses can sit in the gap between policy forms. If you find a material gap, reflect it in the purchase agreement, the escrow or indemnity analysis, and the post-close placement plan.

Deal treatment: price, reps and warranties, indemnities, and closing conditions

Match the protection to the exposure. The core issue is simple: can the system be used now, fixed after close, or is it too risky to buy? Skadden reports that buyers increasingly seek AI-specific representations when an AI feature is central to valuation or presents risks not adequately addressed by general intellectual-property, privacy, cybersecurity, and compliance representations [10][13]. Skadden also identifies licensed or otherwise authorized training data as a potential subject of a specific buyer representation [10].

The purchase agreement should spell out the covered systems, lookback period, materiality thresholds, notice procedures, survival period, access to records, cooperation duties, and whether remediation costs are part of the protection. Score each finding on likelihood (1-5), impact (1-5), detectability (1-5), and urgency (1-5). But don't let the scorecard hide obvious problems. Any confirmed unlawful use, safety risk, open regulator matter, or inability to operate lawfully should be treated as high risk no matter what the numbers say.

Finding Likelihood Impact Detectability Owner Deadline Deal treatment
Customer data used for model training without documented rights 4 5 4 General counsel and chief data officer Before closing or immediate freeze Specific indemnity; possible price adjustment and closing condition
No recurring bias testing for a material hiring model 3 5 3 Chief human resources officer and model owner Within 30 days after closing Warranty plus post-close covenant; pre-close use limitation
Vendor contract lacks deletion, audit, and incident-notice rights 4 4 4 Chief procurement officer and CIO Within 60 days Contract remediation covenant; escrow if replacement risk is material
Open regulator inquiry concerning automated consumer decisions 4 5 5 General counsel Before signing; ongoing through closing Specific indemnity, escrow, cooperation covenant, and possibly closing condition
Model drift dashboard is incomplete but current performance remains within tolerance 3 3 2 Chief technology officer Within 90 days Post-close remediation item

Use a closing condition when continued use could be unlawful or unsafe, when a required license or vendor consent is missing, or when the buyer cannot run the business on day 1 after close. Use a post-close covenant when the business can keep operating lawfully and the fix is measurable, estimable, and controllable. Write that distinction down in plain terms. Do not leave it buried under a generic label like AI risk.

30/60/90-day post-close remediation register

Build the remediation register before signing, not after. Every accepted finding should have 1 accountable owner, a funded budget range in U.S. dollars, a dependency, and a measurable acceptance criterion. Review the register weekly for the first 30 days, every 2 weeks through day 60, and again at the day-90 operating review. Use bottom-up estimates for each issue. Break out external counsel, data repermissioning, engineering, model validation, vendor renegotiation, security testing, and ongoing monitoring as separate cost lines.

Finding Risk rating Owner Cost estimate Deadline Deal remedy Acceptance criterion
High-risk use case lacks a lawful data basis Critical General counsel and chief data officer $50,000-$250,000 Day 10 Immediate freeze; specific indemnity if historical exposure exists Use suspended or approved data rights documented; affected decisions addressed
Material model not revalidated after major version change High Chief technology officer and model owner $75,000-$300,000 Day 30 Post-close covenant; price protection if performance is material Independent testing confirms agreed accuracy, drift, and human-override thresholds
Vendor contract lacks training, deletion, audit, and incident rights High Chief procurement officer and CIO $25,000-$150,000 Day 60 Covenant; escrow or termination right if vendor refuses Signed amendment or documented replacement plan with tested data migration and continuity controls
Privacy notices do not describe AI processing High Chief privacy officer $40,000-$175,000 Day 60 Warranty and remediation covenant Updated notices, consent or alternative legal basis, and retention schedule in place

Each line in the register should also include a residual-risk decision from the accountable owner confirming that the remaining exposure is understood, accepted, and monitored.

Conclusion: the PE checklist for deciding what is usable, fixable, or deal-breaking

The close decision should be auditable. It should not rest on a broad statement that the target "uses AI responsibly." The investment committee needs a clear record that ties each conclusion back to the underlying evidence - system inventory, data provenance, testing results, vendor and cyber controls, legal analysis, transaction treatment, and a named remediation owner with a deadline. That trail is what supports the final call.

Use NIST's AI RMF as the path for that decision: Govern, Map, Measure, Manage.

For each material AI system, the end state should fall into 1 of 3 labels:

Label What it means Typical deal treatment
Usable Rights documented; owner assigned; baseline testing passed; monitoring and incident controls operating Permit continued use with routine oversight
Fixable Risk is bounded; remediation is technically and legally feasible; owner, budget, deadline, and interim control exist Post-close covenant, integration workstream, or purchase-price protection
Deal-critical Unlawful or unverified data rights; material bias or security exposure; missing essential records; unacceptable vendor lock-in; no credible remediation path Closing condition, specific indemnity, escrow, repricing, or walk-away decision

Use these labels to build the portfolio view. One weak system can matter more than an average maturity score across the rest of the estate. Keep the categories separate: usable now, usable with controls, fix after close, pending evidence, and deal-breaking systems.

NIST gives the team a solid way to organize the work, but counsel still has to apply binding U.S. legal, regulatory, and contractual rules.

If the team needs outside diligence support, use a specialist shortlist. The Top Consulting Firms Directory can help shortlist advisers with PE diligence and AI experience. Keep the diligence record - test results, contracts, risk ratings, buyer decisions, and the 30/60/90-day register - as the evidence file.

FAQs

What counts as deal-critical AI risk?

Deal-critical AI risk is any AI failure that can hit compliance, disrupt operations, or damage the brand.

The main risk areas are algorithmic bias, data privacy breaches, model drift, weak data quality, and dependence on third-party tools that open security or compliance gaps. For operators and board-facing leaders, these are not edge cases. They can affect customer trust, audit exposure, service delivery, and, in some cases, enterprise value.

How do I verify AI data rights fast?

Require vendors to show, in writing, where data comes from, why they can use it, how long they keep it, and who they share it with. That means contracts, consent records, and terms that support dataset sourcing and sharing. Then verify those claims through an auditable trail - data lineage, retention records, and third-party sharing logs for each external party.

Use automated discovery and classification tools to map personal data across systems so you can see what sits where. Also confirm that each vendor has a DPA, or an equivalent agreement, that includes audit rights and supporting security evidence.

Which AI issues can wait until after closing?

Usually, items that call for continuous monitoring - not one-time validation - can wait until after closing.

That bucket often includes:

  • drift and bias dashboards or alerts
  • regular risk assessments
  • tabletop exercises and incident protocols

Pre-close, focus on the items that can change deal risk now: data rights, privacy basis, bias tests, cyber controls, third-party tools, and compliance gaps.

Related Blog Posts

Read more