What Legal Teams Mean by eDiscovery
Electronic discovery, or eDiscovery, is the controlled process of identifying, collecting, preserving, reviewing, and producing electronically stored information. For a legal team, it usually concerns emails, messages, spreadsheets, word-processing files, PDFs, database records, voicemail, and material shared through cloud applications. The process is not simply searching for documents; it involves proving that relevant information was preserved, locating it across several data sources, and presenting it in a defensible form. Federal civil discovery in the United States is governed principally by Federal Rules of Civil Procedure 26, 34, 37, and 45, although state rules and court orders may add requirements. A legal team typically relies on custodians, lawyers, and technical specialists to define what must be preserved and reviewed. As of September 2026, the basic workflow remains similar to earlier years, but AI-assisted classification, retrieval, summarization, and privilege analysis are increasingly common in commercial platforms.
Also worth reading: How Should a Legal Team Run an AI-Powered eDiscovery Document Review Workflow in 2026? · How Is AI Changing Legal eDiscovery in 2026, and How Should Firms Use It Safely? · How Should Outside Counsel AI Guidelines Govern Legal Research and eDiscovery in 2026?
The goal is not to collect everything without limit. Under Federal Rule of Civil Procedure 26(b)(1), the scope of discovery must be relevant to a claim or defense and proportional to the needs of the case. Proportionality considers the amount of information in dispute, the parties' access to it, the importance of the requested information, and the likely benefit compared with burden and expense. A team may reasonably collect emails and chat messages from 30 custodians but not every archived record from a company that has operated for 15 years. The defensible response is documented rather than maximal. A smaller, well-supported collection can be more defensible than an indiscriminate production containing large volumes of duplicate, irrelevant, or privileged material.
How the eDiscovery Process Works
A typical matter begins with a preservation notice. The legal team identifies relevant systems, data owners, date ranges, and litigation holds, then communicates those duties to custodians. In Microsoft 365 environments, this may involve applying retention labels, hold queries, or legal holds; in other environments, it may require disabling automatic deletion and preserving mailbox or application data. Collection can be targeted by custodian, date, mailbox, folder, keyword, or data type. Forensic tools may also be needed when deleted files, personal devices, or unusual applications could contain relevant evidence. The team should document each step because a later challenge often concerns not whether a document was produced, but whether the organization reasonably controlled the process that should have found it.
After collection, the data is processed, usually by deduplicating records, extracting text, indexing content, normalizing metadata, and establishing a reviewable family structure. Reviewers then evaluate responsiveness, privilege, confidentiality, and occasionally subject-matter restrictions. The team produces the selected documents in an agreed format, such as PDF, TIFF, or native format, and maintains an audit trail showing how each result was reached. Production may use a secure file-transfer platform, a negotiated production method, or a court-approved system. If the receiving side disputes the production, the producing party may prepare a privilege log, deposition testimony, or an affidavit from the person who directed the review. FRCP 37(e) addresses the loss of electronically stored information and can permit curative measures, but losing data is not excused merely because a vendor tool failed.
Where AI Fits and Where It Does Not
AI can help teams search large collections, rank documents by likely relevance, propose privilege classifications, summarize records, and generate first-pass chronologies or draft document descriptions. These features can reduce manual work, particularly when millions of records must be screened. A retrieval system that finds documents related to a defined issue may be more useful for targeted investigations than a general chatbot that cannot show why a result was selected. Vendors including Harvey, OpenText, DISCO, Reveal, and other eDiscovery providers have marketed AI-assisted review and research connections, while legal-industry commentary from sources such as JD Supra, LawNext, Thomson Reuters, and A&O Shearman describes rapid product development. The presence of these features does not mean that an AI system should operate without human supervision or evidence-based testing.
AI also introduces evidentiary and professional risks. A classification model may confidently mislabel a responsive email as irrelevant, or miss a privilege pattern because the training set did not represent the matter's language. A generated summary may omit a qualification that changes the legal meaning of a passage. Legal teams should therefore test systems on a representative sample, record the acceptance threshold, and retain human escalation for uncertain or high-sensitivity documents. For research and drafting, AI outputs should be checked against the source record and applicable court or ethical rules; fabricated citations, invented quotations, and unsupported factual statements remain unacceptable. The useful question is not whether an AI tool is accurate in the abstract, but whether its performance has been measured for the collection, language, issues, and risk level of the specific matter.
| Feature | Traditional review model | AI-assisted review model | Manual or specialist analysis |
|---|---|---|---|
| Initial document screening | Reviewer reads large volumes of records | Model ranks or classifies records, with reviewer oversight | Lawyer or specialist examines selected evidence for legal analysis |
| Typical strength | Straightforward accountability and familiar audit trails | Speed and consistency across very large collections | Contextual judgment, sensitivity, and difficult factual reasoning |
| Main weakness | Expensive and slow at high volume | Errors, bias, unexplained results, and vendor dependence | Costly per document and difficult to scale |
| Defensibility evidence | Reviewer testimony, protocols, logs | Validation sample, model documentation, reviewer decisions, audit logs | Affidavits, expert analysis, demonstrable legal reasoning |
| Suitable starting point | Smaller, focused matters | Large collections with tested and bounded tasks | Complex disputes involving novel evidence or high consequences |
First, the legal team should define the dispute before selecting technology. Identify claims, defenses, relevant custodians, likely data sources, and an initial date range. Then create a written collection plan naming systems, custodians, search methods, and preservation actions. During preservation, track acknowledgements and exceptions because a hold that is sent but not maintained can create its own problem. A reasonable pilot often includes 500 to 2,000 documents from a sample of custodians, or a smaller set if the issue is sharply defined. The team should test recall, precision, privilege detection, deduplication behavior, and search quality rather than treating the vendor's headline accuracy as a guarantee.
After validation, configure review in stages: culling, de-duplication, text search, machine-assisted ranking, human review, quality control, and privilege review. Set measurable checkpoints, such as reviewing 5% to 10% of the population or all records above a defined risk category. The exact percentage is not a legal safe harbor; it is an operational starting point that must be adjusted for matter complexity. Maintain a privilege log, document production specifications, chain-of-custody records, and correspondence with the other side. At the end of a phase, compare the result against the original issue definition and investigate unexplained gaps. A team that can explain why it selected 12,000 of 120,000 documents has a stronger position than one that merely says a platform produced the set.
Managed Services Versus Software and Self-Review
Legal teams can buy software, retain a service provider, or perform substantial work internally. These categories overlap: many platforms offer both licensed software and hosted review services, while a law firm may manage the legal decisions while a vendor handles collection and processing. The choice depends less on company size than on data volume, technical complexity, staffing, jurisdiction, and the sensitivity of the evidence. A small matter with three custodians and 20,000 documents may be manageable in-house. A regulatory investigation involving 30 terabytes across multiple cloud applications usually calls for specialists, documented forensic procedures, and a tested processing chain.
Cost pressure can encourage a team to accept the lowest bid, but cheap collection does not necessarily reduce total matter cost. Errors discovered after production can require reprocessing, supplemental productions, motions, depositions, and settlement discussions. A provider's proposal should distinguish per-gigabyte processing, hosting, review, export, data-transfer, training, and field-service charges. Ask whether pricing is hourly, per document, per gigabyte, or subscription-based, and whether minimum commitments apply. Contracts should address confidentiality, data location, subprocessors, model training, retention after termination, incident response, and the customer's right to export data. The vendor may offer strong automation, but the client still owns the decision about what is responsive, privileged, or confidential.
| Cost category | Illustrative pricing in 2026 | What determines the price |
|---|---|---|
| Collection and processing | Roughly $5 to $30 per gigabyte | Source complexity, forensic work, metadata, deduplication, and format conversion |
| Hosting and review platform | Roughly $3 to $15 per gigabyte per month, or a negotiated platform fee | Volume, retention, advanced analytics, user seats, and service level |
| Document review | Often $0.10 to $1.00 per reviewed document, with pricing varying widely | Language, issue complexity, reviewer expertise, and AI-assisted versus manual work |
| Data transfer or production | Variable fees, sometimes tied to size or number of files | Security requirements, file type, and delivery method |
| Legal services | Hourly or matter-based professional fees | Custodian interviews, strategy, privilege disputes, drafting, and negotiations |
Common Mistakes and Failure Points
One frequent mistake is beginning with a vendor instead of a legal theory. Search terms copied from a complaint can miss informal language, abbreviations, or documents that matter because of their relationships to other records. Another is collecting email while overlooking Teams chats, collaboration platforms, mobile messages, shared drives, or local archives. Teams should also avoid assuming that a keyword search is equivalent to a defensible review method, especially when the data is multilingual or heavily redacted. Destruction schedules should be suspended where appropriate, but over-preservation can create cost and privacy problems, so collection boundaries need regular reassessment.
Privilege review is another pressure point. AI may flag obvious legal communications while overlooking mixed business and legal content, or it may mark too many documents without improving the final privilege log. The team should define the governing jurisdiction and privilege standard rather than apply a generic checklist. Small sample quality-control reviews are useful, but they do not establish that every human decision was correct. In regulated matters, privacy obligations may conflict with broad discovery requests, and protective orders may be needed. Finally, a team should not allow generative AI to draft a production narrative, chronology, or legal argument without checking the underlying records; polished text can conceal an inaccurate premise.
When to Act and How to Measure Results
Act early when a dispute is reasonably foreseeable, a demand is received, a regulator begins an inquiry, or a complaint identifies events that may trigger a preservation duty. Waiting until a production deadline can eliminate meaningful options and make technical recovery harder. On the other hand, teams should not launch a full-scale collection merely because every business dispute could theoretically involve documents. A short legal and technical assessment, often completed within days, can identify sources, volume, risks, and the smallest defensible next step. A pilot can then test whether the proposed toolset works on real data before a large budget is committed.
Measure results with numbers that reflect legal work, not just software activity. Track days to collection, cost per processed gigabyte, percentage of documents deferred, recall in validated samples, privilege-review disagreement rates, production volume, and the number of supplemental productions required. Ask whether the tool shortened review time or merely created more records for lawyers to audit. Report unresolved data-quality problems, such as missing attachments, incorrect custodians, or inconsistent redactions. These measurements reveal whether the process is working and create a record that can be explained to opposing counsel or a court.
The most defensible eDiscovery process combines legal judgment, technical discipline, and continuous quality control. AI can make a large review faster, but it cannot decide the meaning of the case, waive privilege, or guarantee that every relevant record was preserved. In 2026, the practical advantage comes from selecting narrower questions, preserving data early, testing automation honestly, and preserving a clear audit trail. That approach is more demanding than uploading files to a dashboard, but it is better suited to disputes where authenticity, confidentiality, and legal responsibility matter.