What AI eDiscovery Actually Does

AI eDiscovery uses software-assisted methods to identify, organize, review, and produce potentially relevant legal documents. It does not replace the entire discovery process. Collection, preservation, chain of custody, searching, privilege review, redaction, and production still require decisions made under court rules, litigation holds, and the governing jurisdiction. AI is most useful in the middle of the process, where large volumes of documents must be classified, de-duplicated, clustered, or checked for responsiveness. As of September 24, 2026, the term “AI document review” covers several different technologies, including machine-learning ranking, natural-language search, optical character recognition, entity extraction, summarization, and generative AI. These systems are not interchangeable. A model that helps counsel search a contract library may not be suitable for reviewing millions of emails against a legal hold.

Also worth reading: How does AI eDiscovery verify document accuracy? · What is multi-agent litigation support software and how does it change eDiscovery and document drafting? · How to prepare for an eDiscovery technology review in 2026?

The core benefit is reducing the amount of human attention required to find a smaller set of potentially relevant material. It is not a guarantee that every relevant document will be found or that every privileged document will be protected. A defensible process therefore combines automated assistance with sampling, quality control, documented decisions, and human escalation. The Federal Rules of Civil Procedure still require parties to conduct discovery relevant to a claim or defense, and Rule 34(b)(2)(E) continues to describe an objection to producing documents that are neither relevant nor proportional. AI can help organize a review, but it cannot excuse a party from meeting that obligation. The safest description is assistive automation, not autonomous compliance.

How AI Differs from Traditional Review Methods

Traditional document review usually follows a linear sequence: collect a defined universe, de-duplicate documents, convert files into searchable text, apply searches, and have reviewers evaluate candidate documents. Technology-assisted review, or TAR, uses trained algorithms or active learning to rank or classify documents based on reviewer decisions. Generative AI adds a different layer by interacting with documents in natural language, drafting summaries, and explaining why a document may matter. TAR is a form of machine learning; generative AI is a broader class of models designed to produce content, and a product may use both. The distinction matters because a generative model can produce a plausible explanation that is not supported by the document it describes.

FeatureTraditional or TAR-assisted reviewGenerative AI-assisted review
Primary functionSearch, ranking, classification, and decision supportNatural-language queries, summaries, extraction, and drafting
Typical strengthRepeatable decisions over a defined document populationHandling varied language, questions, and document formats
Main riskInconsistent human coding or poor recallInvented text, missed context, and excessive reliance on fluent answers
Validation needRecall, precision, and reviewer samplingSource checking, testing, privilege controls, and output verification
Best fitLarge, well-defined review sets with known categoriesComplex questions, mixed document types, and legal research or drafting tasks
Human roleCoding, reviewing, escalating, and certifyingSupervising model use, checking outputs, and making legal judgments
The practical question is not whether generative AI is “better than TAR.” Each addresses different friction. A TAR system may classify an email as responsive more consistently than a general-purpose chatbot, while a generative model may help counsel summarize an unfamiliar contract faster than a coding interface. Many deployments combine both, but the organization must document which component made each decision. Reports describing AI for eDiscovery, including vendor and industry publications, frequently emphasize speed, but speed is valuable only if recall, privilege protection, and production defensibility remain acceptable.

How the Workflow Changes

A defensible AI eDiscovery workflow begins before any model is applied. The legal team defines the matter, custodians, date range, data sources, preservation obligations, and the meaning of responsiveness. It then collects data, preserves originals, records the chain of custody, and creates a defensible processing history. The team should test OCR quality, de-duplication, email threading, family grouping, and date handling before trusting automated output. These steps are not administrative decoration; faulty extraction can cause a model to miss a document whose image layer contains the relevant text. Federal Rule 37(e) addresses the loss of electronically stored information, so preservation and restoration procedures should be treated as distinct from review automation.

After processing, the team can use AI for search assistance, document ranking, clustering, entity extraction, issue coding, and review recommendations. Every generated response should remain linked to the underlying document, page, and metadata so a reviewer can reproduce the result. If a model says a document discusses an acquisition, the reviewer should be able to inspect the passage that supports the statement. Counsel should also decide whether sensitive documents may be sent to an external service and whether retention settings prevent later model training. In 2026, data governance is not limited to a vendor’s security page: it includes prompts, uploaded files, stored outputs, access logs, model settings, and administrator permissions.

The final stage is human quality assurance. Reviewers should sample both documents recommended for production and documents excluded by the system, with attention to privilege, confidentiality, personal information, and non-text content. Teams often set thresholds such as reviewing at least 5 percent of a coded population when beginning a new process, then adjusting after measuring error rates. Those percentages are operational suggestions, not universal legal requirements. The better practice is to establish an approved sampling plan, report precision and recall on a representative test set, and revisit the threshold when the review population changes.

Evidence, Accuracy, and Defensibility

AI-assisted review must be evaluated against a known standard rather than judged by how convincing its answers sound. A test set of adjudicated documents can measure whether the system correctly identifies responsive material, separates irrelevant material, and routes uncertain cases for human review. Precision measures how many documents selected by the system are correct; recall measures how many truly relevant documents the system successfully finds. A system with 99 percent precision can still miss many important documents if recall is poor, while a high-recall system may produce a large queue for human review. The team should report both measures, the error type, the sampling method, and the version of the model used. Generative outputs need an additional test for factual support, because a summary can be fluent yet incomplete or wrong.

Defensibility also depends on transparency. At minimum, the matter file should record the tool name and version, deployment date, configuration, data sources, reviewer instructions, sampling results, exceptions, and changes to the workflow. A party should be able to explain why a document was selected, why a redaction was made, and how a privilege decision was reached. Courts generally expect parties to understand the technology they are offering, although the exact level of explanation varies by jurisdiction and relief sought. Parties should not describe a tool as “self-learning” if they cannot explain what data was used, who approved changes, or how outputs were tested. Clear documentation can also help opposing counsel, experts, and the court evaluate proportional review methods.

Accuracy should be tested across the whole review universe. A model trained on contract disputes may perform poorly on instant messages, spreadsheets, handwritten notes, or messages in languages not represented in the test set. AI systems can also perform differently when a document is corrupted, heavily redacted, or stored as an image. The team should include those cases in testing and route them to conventional review. AI should not be used to infer a person’s legal position from tone, writing style, or demographic characteristics. Relevance and privilege are legal judgments tied to the case, and a model’s confidence score is not a substitute for a reasoned determination.

Cost, Pricing, and Business Case

AI eDiscovery costs depend on collection volume, data quality, hosting, processing, review volume, and the pricing model. Vendors commonly charge combinations of per-gigabyte processing, per-document review, per-user platform fees, implementation fees, and optional premium modules. A broad planning range may be roughly $5 to $50 per gigabyte for processing and $2 to $15 per document for some review services, but these figures are not universal quotes and can change with data complexity, volume, urgency, and service level. A generative assistant may add a separate subscription, per-seat fee, or usage charge. Organizations should request a written statement of units, minimums, overages, storage fees, and charges for extracting or exporting data.

The business case is strongest when AI reduces repetitive review while preserving quality. Management should calculate the fully loaded cost of current review labor, outside counsel spend, technology fees, privilege review, and rework. It should then compare those costs with a controlled pilot that measures time saved and error rates. A cheaper tool that increases privilege errors or misses a dispositive document can be more expensive after remediation, sanctions, or an extended dispute. AI may not be economical for a small, straightforward matter with only a few hundred documents and clear issues. It is more likely to help when the same coding decisions repeat across thousands of documents, or when counsel needs to answer complex questions across heterogeneous sources.

Pricing claims should be tested against actual data. Ask whether the quoted rate includes OCR, de-duplication, email threading, translation, audio processing, or privilege review. Confirm whether the provider can export audit logs and the underlying reviewer decisions, and whether the customer can leave with its data if the contract ends. The best result is not necessarily the lowest per-user price; it is a transparent system with predictable total cost, reliable validation, and a controlled migration path.

Common Mistakes in AI Document Review

The first mistake is treating every product marketed as “AI” as equivalent. Vendors may use AI for search, ranking, tagging, transcription, summarization, or full generative analysis. The buyer should ask what the system actually does and what happens when it is uncertain. A second mistake is skipping human review because a system produced a high confidence score. Confidence scores are useful triage signals, but they do not establish legal relevance or privilege. A third mistake is failing to control the data sent to an external model, particularly when the material contains trade secrets, personal information, or information subject to litigation holds.

Another common error is measuring productivity without measuring quality. A team may record how many documents were processed per hour while ignoring missed responsive material, incorrect privilege calls, or duplicate production. It is also risky to change prompts, models, or review criteria during the same coding phase without recording the change. Generative systems can behave differently after a configuration update, so version control matters. Organizations should not allow individual lawyers to use unapproved public tools for privileged matter files, and they should not assume that a vendor’s terms automatically satisfy the client’s confidentiality obligations. A written AI policy, role-based access, training, and incident procedures reduce these risks.

Finally, do not confuse legal research with discovery review. A generative research tool may help counsel locate a case or draft a clause, but it does not automatically know which custodial documents must be preserved or produced. Research outputs require verification against primary authority, while review outputs require verification against the source document. A platform that connects evidence to research and drafting may be useful, but the connection must preserve source provenance. The distinction between finding law and reviewing evidence should remain visible in permissions, workflows, and audit records.

When to Use AI and When Not to Use It

AI-assisted review is usually appropriate when the document population is large enough that manual coding is inefficient, the issues are reasonably defined, and the organization can test the tool on representative data. It can also be useful for early case assessment, where the team needs to identify custodians, entities, dates, and recurring issues before committing to a full review. In a smaller matter, ordinary search, keyword review, and targeted human analysis may be simpler and safer. If the legal questions are unstable, the data set is unusually diverse, or the risk of privilege waiver is high, a narrower assistive role may be preferable to broad automation.

The organization should begin with a pilot, not a production migration. Define the success criteria in advance, including minimum recall for responsive documents, acceptable privilege error rate, reviewer override procedure, and reporting requirements. A pilot often runs against a sample created with experienced lawyers and reviewed through a second pass. The team should test edge cases such as embedded spreadsheets, attachments, foreign-language text, duplicates, and documents with missing metadata. If results are unstable, increase human review or return to conventional methods. AI is not a reason to ignore proportional discovery; the Federal Rules permit proportional methods, but the court and the record govern the permissible process.

By September 24, 2026, organizations should also consider the EU AI Act’s staged application dates and its classification obligations. The Regulation entered into force on August 1, 2024, with provisions applying in phases beginning in 2025 and 2026, depending on the system category and applicable dates. An eDiscovery tool is not automatically a high-risk system merely because it uses AI, and legal interpretation should be confirmed for the specific deployment. NIST’s AI Risk Management Framework provides a useful governance structure even where an organization is not directly subject to a particular AI statute. The practical rule is to start controlled, document evidence, and escalate unresolved legal questions to qualified counsel.

How to Choose a Responsible Tool

A responsible tool should support the full matter workflow rather than only promising faster document review. Ask whether it can preserve original files and metadata, record every processing step, identify the model and version, export review decisions, and produce an audit report. The vendor should explain how it handles OCR failures, unsupported languages, attachments, email threads, privilege workflows, and access permissions. It should also state whether customer data is used to train shared models, how long information is retained, and what happens after contract termination. These are contract and governance questions, not merely product features.

The evaluation should include both the technical team and the lawyers responsible for discovery. Technical reviewers can test integrations, security, exports, and scalability. Legal reviewers should assess relevance definitions, privilege controls, explainability, and whether the system’s suggestions match the issues in the case. Procurement should review data processing terms, service levels, indemnity provisions, and the right to audit material subprocessors. A pilot should use a representative, controlled sample and compare the tool with the existing baseline. The selected product should be one that improves throughput without making errors harder to detect.

No single tool suits every organization, and vendor marketing should not be treated as independent evidence. A platform integrated with a recognized review environment may be convenient for one team, while another may prefer a lighter research or drafting assistant connected to an existing case system. The most defensible choice is the one that makes human decisions traceable. That means a reviewer can see the source document, understand the automated recommendation, correct the result, and export a record of the correction. AI can accelerate the mechanics of review, but legal accountability remains with the people responsible for the matter.