What AI eDiscovery Software Actually Does

AI eDiscovery software is software that searches, classifies, extracts, summarizes, and routes documents collected during litigation, investigations, regulatory responses, and internal reviews. It does not replace the legal team’s judgment; it reduces the number of pages a person must read while producing traceable results for review, production, and possible courtroom use. Modern systems commonly combine keyword search, metadata filters, machine-learning ranking, generative summarization, and workflow controls in one platform. The 2026 product cycle includes more agentic features, meaning software can propose a sequence of review actions rather than waiting for every command, but vendor announcements should not be mistaken for proof of accuracy in a specific matter. The practical question is whether a system finds responsive material, explains why it did so, and gives counsel a defensible way to validate the result. A good answer therefore begins with the matter’s data, preservation duties, search terms, and review protocol rather than with a preferred vendor.

Also worth reading: What are the best practices for validating AI-assisted eDiscovery results before producing documents? · What is multi-agent litigation support software and how does it change eDiscovery and document drafting? · What Is the Actual Document Review Capacity of AI in Modern eDiscovery?

How the Review Process Works

A typical matter moves through eight stages: legal hold and collection, deduplication, normalization, text extraction, search or technology-assisted review, human review, quality control, and production. Some stages occur in different orders, especially when new files arrive or a client expands the document population. AI tools can identify duplicates, near-duplicates, email threads, attachments, privileged material, and responsive language while retaining the original document and its metadata. For generative functions, the software may create a short issue summary, extract dates and entities, or suggest a review coding, but the output should be treated as a proposal until a reviewer confirms it. Every decision should be reproducible through search terms, filters, saved models, version history, and an exportable audit trail. That audit trail matters more than a polished interface because opposing parties and courts may ask how a potentially privileged or responsive document was selected or excluded.

Core Capabilities to Look For

The most useful capabilities fall into four groups. First, ingestion and processing features must handle email, shared drives, mobile applications, databases, chat records, audio files, and common office formats without silently losing attachments or metadata. Second, investigation features include metadata search, field-level filters, date-range controls, near-duplicate detection, email threading, and configurable search-term lists. Third, review features include technology-assisted review, predictive coding, issue coding, redaction suggestions, privilege detection, translation, summarization, and reviewer queues. Fourth, governance features include role-based permissions, matter-level isolation, encryption, retention controls, audit logs, and configurable restrictions on sending matter data to an external AI service. AI eDiscovery and legal research products are increasingly connected, but a legal research tool is not automatically a full collection and review platform. CoCounsel Legal, for example, is positioned around legal research and drafting with Westlaw and Practical Law, while eDiscovery products are built around evidence processing, review, and production. Buyers should distinguish those categories before comparing features or pricing.

A Practical Evaluation and Adoption Plan

Start with a representative sample rather than a vendor demonstration using clean, familiar documents. Ask each candidate to process a fixed set containing known responsive documents, known nonresponsive documents, duplicates, privileged communications, encrypted files, and unusual formats. Measure whether the system found the planted documents, separated the correct families, preserved attachments, and produced usable logs. A common test is to provide 1,000 documents with 100 known responsive examples, then compare recall, precision, and reviewer corrections; a 95% headline accuracy claim is not meaningful unless the testing method and the definition of accuracy are disclosed. Run a pilot with at least 2 reviewers, record time per decision, and calculate the number of documents that changed classification after human correction. A second pilot should test privilege, confidentiality, and data-residency requirements, because retrieval quality is only one part of a defensible workflow. Finally, require written confirmation of data deletion, model training practices, subprocessors, incident response, and export formats before uploading client material.

Comparing Platforms and Alternatives

FeatureTraditional review platformGenerative or agentic assistantManaged review serviceManual or basic search tool
Best roleControlled in-house review workflowDrafting, summaries, coding suggestions, and guided tasksVendor-hosted reviewers and production supportSmall, simple, or low-volume matters
Main advantageMature filtering, review, auditing, and production toolsFaster assistance with language-heavy tasks and flexible workflowsAdds staffing, escalation, and operational capacityLowest technology complexity
Main limitationAI quality and usability vary by productMay produce unsupported statements or expose confidential dataHigher cost and less direct client controlReview time rises quickly with document volume
Validation burdenTest search, ranking, privilege, and exportsTest grounding, citations, permissions, and prompt controlsReview service-level commitments and chain of custodyDocument human decisions and search steps carefully
Typical buyerLegal departments and litigation firmsTeams wanting research, drafting, and review assistanceOrganizations needing end-to-end supportSolo practitioners handling limited collections
These categories can overlap. A platform may include generative AI while still requiring conventional technology-assisted review, and a managed service may use the same underlying software as a customer’s internal team. The relevant comparison is therefore contract scope, security controls, reviewer expertise, and total cost, not whether a product calls itself “agentic.” Traditional platforms often provide stronger evidence-chain and production controls, while generative assistants may reduce time spent summarizing a document or drafting a first-pass issue tag. Neither is automatically safer or more accurate. The table is a decision aid, not a ranking, and it should be adapted to the client’s privacy obligations and the forum’s rules.

Accuracy, Recall, Precision, and Useful Thresholds

Four measurements deserve attention. Recall measures how many known responsive documents the system retrieves; precision measures how many retrieved documents are actually responsive; coding accuracy measures agreement with human decisions; and throughput measures how many pages reviewers handle per hour. No single percentage should decide the purchase, because an imbalanced sample can make accuracy look excellent while missing important evidence. Teams should set minimum acceptance criteria before testing, such as 100% recovery of the known test documents, 100% preservation of attachments and metadata for accepted files, and zero unauthorized access events during the security exercise. For a larger evaluation, 5,000 to 20,000 documents can provide a more informative sample than a 50-document demonstration, although the right number depends on format variety and the number of review issues. Compare results across several recall targets rather than using one threshold, and require the vendor to explain changes in ranking, summaries, or coding when the underlying model is updated. The goal is measured performance under matter-like conditions, not a marketing score.

Common Mistakes in Buying and Using These Tools

One mistake is treating AI output as a legal conclusion. A generated summary can omit a qualification, misread a scanned page, or incorrectly suggest that a document supports a proposition. Another mistake is uploading every available file before agreeing on a search and review protocol, which increases cost without improving the result. Teams also err when they compare vendors using different collections, different search terms, or different reviewer instructions. A second error is failing to test privilege and confidentiality workflows; a system can be accurate on responsiveness while mishandling sensitive material. Procurement teams sometimes focus on generative features and ignore data deletion, model training, subprocessor location, and auditability, even though those controls may determine whether the product is permitted for the matter. Reviewers may also accept coding suggestions without recording corrections, making later quality checks impossible. Finally, assuming that a new agentic release is ready for unsupervised production is premature. As of September 2026, the market is moving quickly, but product announcements from vendors and industry coverage do not establish independent reliability in every jurisdiction or court.

Cost, Timing, and When to Act

Pricing is rarely a simple public per-page figure. Enterprise eDiscovery contracts may combine platform fees, user or reviewer charges, processing and hosting fees, optional modules, implementation services, and charges for advanced AI functions. A lower subscription can still produce a higher total cost if it requires more reviewers, more data processing, or additional security review. Ask for a written three-year cost model showing user counts, storage, processing volume, support, exports, and the price of each AI module; also state whether fees change when a model or feature becomes generally available. Implementation commonly takes weeks for a contained pilot and months for a complex multi-source matter, so timing depends more on collection quality and internal approvals than on model speed. Organizations should act when a matter has a substantial document population, a fixed preservation obligation, or a review deadline that makes manual reading impractical. Small matters may justify conventional search and manual review, while recurring programs with thousands of documents each month usually benefit from structured testing and workflow automation. The best buying window is before a crisis forces a rushed selection, not after a court deadline has already passed.