AI has moved from the margins of eDiscovery to its center. As of 2026, artificial intelligence touches nearly every stage of the electronic discovery lifecycle: identifying potentially relevant data sources, collecting and processing information, culling down document populations, reviewing documents for responsiveness and privilege, and even drafting production summaries and deposition outlines. Courts, litigants, and vendors have all adapted quickly. A 2025 survey found that roughly 61% of federal judges report using AI in some capacity, which is raising expectations for what courts expect from litigators who appear before them. If you are involved in litigation, an internal investigation, or a regulatory response, understanding how AI functions inside eDiscovery is no longer optional knowledge — it shapes budgets, timelines, sanctions risk, and case outcomes.

The Direct Answer: What AI Actually Does in eDiscovery

Also worth reading: How is AI legal education reform expected to evolve by 2027, and what does this mean for eDiscovery and document drafting? · How does the EU AI Act classify high-risk legal software and what certification is required for eDiscovery tools? · What are the definitive AI eDiscovery best practices for legal teams in 2026?

At its core, eDiscovery is the process of finding, preserving, collecting, reviewing, and producing electronically stored information (ESI) in legal matters. AI enters this workflow at multiple points. Traditional technology-assisted review (TAR), sometimes called predictive coding, uses supervised machine learning: a senior reviewer codes a training set of documents as responsive or not responsive, and the model then ranks or classifies the remaining population. Continuous active learning (CAL) variants retrain the model after every human coding decision, so the system improves as review proceeds.

Generative AI (GenAI) added a second wave starting around 2023–2024. Large language models can now summarize thousands of documents into digestible briefs, answer natural-language questions about a document corpus, identify key players and communication patterns, flag privilege issues with explanations rather than just scores, and draft first-pass work product such as privilege logs, chronologies, and deposition question outlines. In 2026, vendors have layered conversational assistants on top of these capabilities — for example, eDiscovery AI launched CaseBot, a conversational AI assistant that lets case teams query their own case data in plain English, and OpenText has promoted eDiscovery Aviator agents positioned as force multipliers for review teams. Casepoint was also selected by Elastic in 2026 to modernize eDiscovery and legal hold workflows, reflecting how search infrastructure and AI are converging.

The practical effect is that AI compresses the most expensive phase of litigation — document review, which historically consumed 60% to 80% of eDiscovery budgets — while shifting human effort toward validation, quality control, and strategy.

Why AI Became Necessary: Data Volume and Cost Pressure

The reason AI dominates modern eDiscovery is arithmetic. A single custodian can generate tens of gigabytes of email, chat messages, text messages, cloud documents, and collaboration data per year. A mid-sized commercial dispute involving 20 custodians can easily produce millions of documents after de-duplication. Manual linear review at typical contract attorney rates of $50 to $150 per hour would cost hundreds of thousands of dollars per million documents and take months.

AI changes this equation in three ways. First, machine learning classification can reduce the reviewable population by 70% to 95% by suppressing obviously irrelevant material before humans ever see it. Second, GenAI summarization lets attorneys review at the level of themes and threads rather than individual documents, which suits early case assessment. Third, AI-driven data mapping helps teams target collections more precisely, reducing over-collection that inflates processing and hosting costs. FTI Consulting's guidance on public sector disputes emphasizes the 'measure twice, cut once' principle: careful upfront scoping of what AI will do, on what data, with what validation, prevents expensive downstream errors.

There is also a defensive driver. Opposing parties increasingly use AI, so a team relying solely on keyword searches may miss hot documents that semantic models would surface. Courts have grown comfortable with TAR since the landmark Da Silva Moore decision in 2012, and judges now routinely approve negotiated ESI protocols that specify how predictive coding will be used, validated, and disclosed.

Stage-by-Stage: Where AI Fits in the EDRM Workflow

Understanding AI's role requires walking through the Electronic Discovery Reference Model (EDRM). During identification and preservation, AI-assisted data mapping tools scan enterprise systems to locate relevant repositories and automate legal hold notices, tracking acknowledgments and flagging non-compliant custodians. Casepoint's integration with Elastic search infrastructure reflects this trend toward faster, more scalable hold and collection pipelines.

During processing and early case assessment, AI performs near-duplicate detection, email threading (collapsing long chains so reviewers see only the most inclusive message), entity extraction, and language detection. Analytics dashboards show custodians, date ranges, and communication networks, letting attorneys decide within days — not weeks — whether a matter warrants settlement discussion.

During review, supervised TAR and CAL rank documents by predicted relevance. Reviewers start with the highest-ranked documents, where relevance rates are highest, and stop when statistical sampling shows they've reached the target recall level — commonly 75% to 85%. GenAI adds document-level summaries, issue tagging suggestions, and privilege rationale drafting. Arnold & Porter's analysis of 'the emerging framework on AI prompts, privilege, and discovery' highlights a critical wrinkle: when attorneys use GenAI during review, the prompts themselves and the outputs may become discoverable or bear on privilege analysis, a topic the National Law Review has described as train tracks now colliding between AI, eDiscovery, and privilege law.

During production, AI assists with redaction suggestion, privilege log generation, and quality control sampling. Post-production, the same corpus powers deposition preparation, motion drafting support, and trial exhibit organization.

Comparing Your Options: TAR, GenAI Review, Keywords, and Human-Only Review

Choosing a review approach is one of the most consequential decisions in any matter. The table below compares the four dominant options as they stand in 2026.

FeatureKeyword Search + Linear ReviewSupervised TAR / CALGenAI-Assisted ReviewHybrid (TAR + GenAI QC)
Typical cost per million docs$250,000–$500,000$80,000–$200,000$60,000–$180,000$100,000–$220,000
Speed to completionMonthsWeeksDays to weeksWeeks
Recall reliabilityUnmeasurable without samplingMeasurable via statistical validationImproving; requires validation protocolStrongest — dual-layer verification
Explainability to courtsLow — opaque term selectionHigh — established case law since 2012Moderate — evolving standardsHigh
Privilege riskHigh — human fatigue errorsModerateModerate — hallucination and prompt-discovery riskLowest
Best use caseSmall matters under 10,000 docsLarge structured email populationsEarly case assessment, investigationsHigh-stakes litigation with court scrutiny
Keyword-only review remains defensible for small matters but is increasingly criticized as unreliable for large corpora because keywords miss conceptual matches and generate massive false-positive rates — often 70% or more of hits being irrelevant. Pure GenAI review without human validation is risky: models can hallucinate content, misattribute statements, and their internal reasoning is not auditable in the way a TAR validation sample is. The hybrid approach — using TAR for measurable recall plus GenAI for prioritization, summarization, and quality control — is emerging as the default for sophisticated litigation teams, though it costs more than either method alone.

Practical Steps: Implementing AI in Your Next Matter

Start with scoping. Before selecting any tool, define the matter's custodian list, date ranges, data types, and target recall. Document these decisions in the ESI protocol or a defensible process memo. If you anticipate producing a TAR validation to opposing counsel or the court, agree on disclosure terms early — some protocols require sharing training statistics, others keep them privileged work product.

Second, run a pilot. Code a random sample of 500 to 1,000 documents yourself or with senior reviewers, then measure how the AI tool performs against that ground truth. Vendors' marketing claims about accuracy rarely survive contact with your specific data. A pilot costing a few thousand dollars can prevent a six-figure misfire.

Third, establish a validation protocol. Statistical sampling at defined intervals — typically every 2,500 to 5,000 coded documents under CAL — lets you estimate recall with confidence intervals. When the elusion rate (relevant documents remaining in the discard pile) drops below your agreed threshold, you can defend a stopping point.

Fourth, govern GenAI use explicitly. Adopt a written policy covering which models may touch client data, whether vendor models train on your inputs (they should not), how prompts are documented, and how outputs are verified. Several major providers, including Anthropic's expanded Claude offerings for law firms announced in 2026, now include contractual commitments that firm data is not used for model training — get that in writing regardless of vendor.

Fifth, train the team. Reviewers need to understand that AI rankings are probabilistic, that low-ranked documents still occasionally contain smoking guns, and that GenAI summaries must be spot-checked against source documents before being relied upon in filings.

Common Mistakes and How Courts Are Responding

The most frequent error is treating AI output as ground truth. GenAI models hallucinate — they fabricate quotes, invent dates, and confidently mischaracterize documents. Attorneys who cite AI-generated summaries without reading underlying documents risk Rule 11 sanctions, and courts have already sanctioned lawyers for filing briefs containing fabricated citations generated by AI tools. The same discipline applies to eDiscovery: an AI-flagged privilege call must be verified by a human before production.

A second mistake is poor prompt hygiene. Arnold & Porter's analysis notes that prompts typed into discovery-stage AI tools can themselves reveal case strategy, and if the tool is third-party hosted, those prompts may be discoverable or waive privilege over the analysis. Treat every prompt as if it could appear in front of a judge.

Third, teams often skip validation and then cannot defend their production. If opposing counsel demonstrates that a cheap AI cull missed clearly responsive documents, courts can compel additional review, shift costs, or in extreme cases draw adverse inference instructions. Defensibility lives or dies on documented methodology, not on the brand name of the software.

Fourth, organizations over-collect because AI makes processing seem cheap, then drown in hosting fees and review scope creep. Fifth, some firms ignore the EU AI Act framework adopted in 2024, which imposes obligations on high-risk AI systems that can reach AI tools used in judicial-adjacent contexts — cross-border matters require checking both the tool's compliance posture and your own governance documentation.

Costs, Pricing Models, and Budget Planning

eDiscovery AI pricing in 2026 generally follows one of four models. Per-gigabyte processing and hosting runs roughly $25 to $60 per GB per month for hosting, with processing one-time fees of $100 to $400 per GB depending on data complexity. Per-document review pricing through managed review providers ranges from $0.05 to $0.35 per document depending on complexity and language. Subscription platforms aimed at in-house teams charge $5,000 to $50,000 per year for self-service analytics and AI features. GenAI features are increasingly bundled but sometimes carry per-query or per-token surcharges.

Budget realistically: a 500,000-document commercial matter using a hybrid TAR-plus-GenAI approach might run $150,000 to $300,000 all-in, versus $600,000 to $1.2 million for equivalent manual review. However, savings evaporate if over-collection doubles your data volume, so invest in targeted collections first. Also budget for expert testimony if your AI methodology will be challenged — forensic consultants charge $300 to $700 per hour.

When to Act and What Comes Next

Act now if you are in active litigation or anticipating a dispute: preservation obligations attach immediately, and AI-assisted legal hold automation reduces spoliation risk from day one. If you are a law firm or legal department without an AI governance policy, draft one this quarter — courts, clients, and insurers are all asking for it. If you are a vendor buyer, demand validation studies, data-training opt-outs, and audit logs before signing.

Looking forward, expect agentic AI systems that execute multi-step review workflows autonomously under human supervision, tighter integration between eDiscovery platforms and general-purpose LLMs, and continued judicial refinement of disclosure standards for AI-assisted review. Law.com's coverage of generative AI realities in e-discovery makes clear that the gap between vendor promises and courtroom-tested performance remains real. The winning posture is neither blind adoption nor refusal — it is measured deployment with documented validation, human oversight at every consequential decision point, and honest accounting of what the technology does well (speed, scale, consistency) and poorly (judgment, context, accountability).

For teams handling AI-related disputes specifically, the same toolkit applies in reverse: AI-generated evidence, chat logs from AI assistants, and model outputs are themselves discoverable ESI, adding a new category of data source to preserve and review. The organizations that master AI in eDiscovery — both as a tool and as a subject of discovery — will hold a durable advantage in modern litigation.