What an AI discovery pilot is—and what it should test

An AI discovery pilot is a limited, time-bound trial in which a legal team uses an AI-assisted technology to perform part of electronic discovery and measures the result against agreed operational and quality benchmarks. The trial might test technology-assisted review, document classification, search-term generation, early-case assessment, privilege analysis, or the retrieval and ranking of potentially relevant records. It should not be framed as permission to replace judgment with automated conclusions. As of September 27, 2026, the more useful question is not whether AI can rank documents, because many systems can do that, but whether it can reduce total review effort without increasing missed material, inconsistent decisions, security exposure, or later rework.

Also worth reading: What Should an Indian Law Firm’s AI Policy Cover for E-Discovery and Legal Drafting in 2026? · How Do Law Firms Build Secure Legal AI Governance for Research and E-Discovery? · What Are the Best Practices for Legal Discovery in 2026?

A defensible pilot normally examines a defined slice of a real matter or a legally representative sample rather than a demonstration assembled from unusually easy documents. The population should reflect the client’s data mix, including email, attachments, spreadsheets, images, PDFs, chat records, and mobile messages where those formats are material. Teams should preserve the existing chain of custody and require human authorization for changes that affect production. The objective is to produce evidence for a procurement decision, not merely favorable screenshots or an impressive demonstration. A pilot that omits exceptions and difficult records will usually overstate performance.

The evaluation should compare AI-assisted work with a documented baseline. For technology-assisted review, that baseline may be the current review rate, error rate, elapsed time, and per-million-document cost. For search, it may be recall under existing search terms and the number of documents requiring review. Results should be separated by document family and custodian group because an overall percentage can conceal poor performance on a small but important category. A 90% agreement rate on routine email may still be unacceptable if the system performs at 50% on a set of contracts containing liability terms. The pilot therefore measures both aggregate performance and weakness in the cases that matter most.

Establishing a practical evaluation method

Begin by writing down the decision the pilot must support. A useful threshold might require at least 95% recall on the gold-standard set, no worse-than-baseline privilege accuracy, and a reduction of at least 30% in documents sent to first-pass review. Those numbers are examples, not universal legal standards. The actual thresholds should reflect the sensitivity of the matter, the parties’ agreement, the governing procedural rules, and the organization’s risk tolerance. A criminal investigation, a commercial contract dispute, and a routine internal investigation do not have the same consequences if a responsive document is missed.

A sound method divides the evaluation set into training, validation, and blind test portions, even when the vendor provides its own model and says no customer-specific training is required. Reviewers should create relevance and privilege decisions under a written protocol, resolve disagreements through adjudication, and retain the resulting judgments as the test reference. The AI output should be evaluated on the blind set rather than on examples repeatedly shown during configuration. This prevents a vendor or legal team from tuning the trial until it reproduces familiar decisions while failing to generalize to new material.

The team should record several concrete measures: recall, precision, false-negative rate, false-positive rate, reviewer disagreement, override frequency, processing time, and total cost. Recall is especially important in discovery because a false negative can mean that a relevant document never reaches the attorney making the responsiveness decision. Precision measures how much irrelevant material enters review, while override frequency indicates where users reject the system but does not by itself prove that the AI was wrong. For generative search or summarization, reviewers should separately test whether cited passages actually support each response and whether the tool fabricated a document reference.

FeatureNarrow AI discovery pilotBroad enterprise AI trial
Typical scopeOne workflow, one data set, 4–12 weeksSeveral legal departments, workflows, and regions
Test dataRepresentative matter sample plus blind holdoutCurated benchmark or selected use cases
Main decisionContinue, revise, or stop one pilotArchitecture, governance, vendor, and investment direction
Cost visibilityDirect pilot and reviewer timeSoftware, integration, security, training, and change management
Main riskInconclusive sample or weak baselineHigh expense and organizational disruption before basic performance is known
Appropriate thresholdPre-agreed accuracy, recall, speed, and security criteriaRepeatable controls and proven business case across teams
## What the team should measure

The most informative dashboard combines quality, speed, cost, and governance rather than presenting a single accuracy score. For technology-assisted review, report recall and the reduction in review volume, but also show the number of documents used to calculate each result. A result based on 2,000 reviewed documents has a materially different evidentiary value from one based on 200,000. Where practical, publish confidence intervals so the legal team can see whether observed improvement exceeds sampling noise. A vendor claiming 92% agreement should not be allowed to obscure a small test population or substantial differences among custodians, document types, or languages.

Measure the full workflow, because deployment time can erase apparent efficiency. Record how long it takes to prepare data, configure the system, conduct quality control, remediate errors, export decisions, and deliver a production. A tool that classifies documents in seconds but requires two days of manual mapping may not help a small matter. Conversely, an AI feature that modestly reduces first-pass review can still be worthwhile if it operates on an existing review platform and does not create duplicate data exports. Time savings should be reported in both reviewer hours and elapsed calendar days because these are not interchangeable.

Cost should include more than subscription fees. During a 2026 pilot, total cost may include the license, data preparation, storage, exports or processing charges, security review, vendor assessments, attorney time, paralegal time, quality control, and remediation. Pricing varies too much for a reliable universal range: some tools are priced per user or month, some per document, some by volume tier, and others through negotiated enterprise agreements. Small teams should therefore request a written quote covering overage, minimum commitments, support, model usage, and termination. A pilot discount may reflect only the first 50,000 or 100,000 documents, not the cost at production scale.

Quality should also include the treatment of confidential information. The test should establish whether prompts, document content, embeddings, telemetry, or feedback are retained; whether any information is used to train a general or customer-specific model; and where the data is stored and processed. These questions are increasingly routine, particularly for international matters, but merely obtaining a security questionnaire is not enough. Contract language, user controls, and technical behavior must be checked for consistency, and exceptions should be approved by the responsible legal, privacy, records, and information-security functions.

Building the pilot around real legal work

A practical first phase selects one low-to-moderate risk workflow with a measurable baseline. Technology-assisted review on a settled internal investigation is often easier to evaluate than fully autonomous analysis of a novel trade-secret dispute, provided no privilege or professional-duty concerns are mishandled. The team should document the matter assumptions, custodians, date range, search sources, issue definitions, and exclusions before viewing the vendor’s results. It should then run a small configuration phase in which attorneys refine categories and workflows. That phase is distinct from the blind evaluation and should not be counted as proof of production performance.

The second phase sends the untouched evaluation set through the proposed workflow. Human reviewers should receive both AI recommendations and the ordinary case materials, but the protocol should prevent unconscious overreliance. In a blinded exercise, reviewers can see the machine result and independently assess the same documents, allowing the team to record adoption and override rates. A second group can review without the AI output to estimate anchoring. If reviewers accept nearly every recommendation without checking them, the apparent speed improvement may simply reflect automation bias rather than sound classification.

The final phase tests exception handling. This includes encrypted or corrupted files, duplicates, near-duplicates, foreign-language material, unsupported formats, handwritten records, image-only documents, long spreadsheets, and access-controlled sources. The system should also be tested against misleading text, such as an attachment name suggesting a contract when the file is a news article, or an email whose apparent subject differs from its content. Generative systems can produce plausible but false characterizations, so outputs should be traceable to source documents and reviewed as recommendations rather than factual findings.

For search, the pilot should compare the AI-generated terms or ranking with the team’s existing approach. Search quality cannot be judged by whether the output looks sophisticated; it must be tested against a known relevant-document set and reviewed for recall. A useful experiment may create two teams using the same issue definitions, one with conventional search and one with AI assistance, then compare documents found, time spent, and decisions overturned. Any privilege waiver concern should be reviewed under applicable law and orders, not treated as an ordinary model metric.

Comparing discovery AI, legal research, and drafting tools

Discovery tools, legal research systems, and drafting assistants solve different problems, even when they use similar underlying models. Discovery software is designed to find, classify, redact, and produce large collections. Legal research tools help locate authorities and explain how sources may apply to a legal question. Drafting tools generate or revise contracts, pleadings, memoranda, and clauses. A team that evaluates all three with one accuracy percentage is measuring the wrong things, because a citation error in research, a missed production document, and an incorrect contractual definition have different consequences.

FeatureAI eDiscoveryAI legal researchAI document drafting
Core outputSearch, ranking, tags, review queues, productionsAuthorities, citations, explanations, and research trailsText, clauses, comparisons, revisions, and first drafts
Main quality testRecall, responsiveness, privilege, traceabilityAuthority accuracy, currency, quotation accuracy, reasoningInstruction following, factual grounding, defined terms, and style
Typical human controlAttorneys review classifications and privilege callsResearcher verifies every material authority and pin citationLawyer validates facts, obligations, risk allocation, and final language
Major failure modeRelevant document missed or privilege mishandledHallucinated or outdated authority presented as realPlausible clause conflicts with instructions or governing law
Best pilot questionDoes it reduce review effort at stable quality?Does it improve verified research time and coverage?Does it reduce drafting effort without increasing review defects?
A single vendor may offer all three capabilities, but product modules should still be evaluated separately. The data permissions, audit trails, retention, and training terms may differ by module. Demonstrations often emphasize generative interfaces because they appear interactive, while less visible functions—such as field-level privilege analysis, deduplication, or production validation—may determine real operational value. Procurement should therefore compare the proposed workflow with the existing matter platform, not only the vendor’s newest user experience.

Alternative approaches include conventional search and analytics, human-only review, rules-based technology-assisted review, and established machine-learning platforms. Conventional analytics can outperform AI when the issue is narrow, the data is highly structured, or auditability is paramount. Human review remains necessary for legal judgment, especially privilege, waiver, confidentiality, and nuanced responsiveness. Rules may be better when conditions are explicit and stable, while AI may help with unstructured text and large-scale ranking. The best alternative is often a controlled combination, provided the division of responsibility is documented and the team measures the combined result.

Common mistakes that make results unreliable

The most frequent mistake is testing on data the vendor selected. A demonstration set may contain clean, already-classified documents while omitting the actual challenge: mixed relevance, privilege ambiguity, short messages, attachments, and inconsistent custodian behavior. Another common error is allowing the vendor to define the ground truth without independent review. If the same people configure the categories, adjudicate the gold set, and declare victory, the evaluation lacks separation between preparation and measurement.

Teams also treat precision, recall, accuracy, and agreement as interchangeable. In imbalanced datasets, a system can achieve very high accuracy by labeling almost everything irrelevant. Conversely, precision can decline when a recall-oriented system correctly surfaces more documents for human review. The evaluation should state which errors are acceptable, which require escalation, and how performance changes when the system operates at a stricter threshold. For a first-pass system, missing responsiveness may be more serious than adding irrelevant documents, but the legal team must make that decision rather than adopting a vendor default.

Another mistake is ignoring adoption. A technically capable tool may fail because reviewers do not trust its recommendations, the interface adds steps, or the results cannot be exported cleanly into the review platform. Conversely, enthusiastic early adoption can conceal a lack of independent testing. Measure how many recommendations are accepted, changed, or ignored, but investigate a suspiciously high acceptance rate as well as a low one. Legal professionals should remain accountable for final decisions, and a pilot should not make autonomous production decisions merely because a vendor advertises them as possible.

Finally, teams often treat security, privilege, and records obligations as items to resolve after the product works. That sequence is backwards for sensitive data. Before loading any material, confirm authorization, contractual restrictions, cross-border processing terms, data retention, subprocessors, incident notification, deletion, audit logging, and model-training controls. The team should test whether privileged material can be segregated and whether generated summaries retain references to the underlying record. These controls can materially change cost, usability, and even whether a particular pilot should proceed.

When to continue, expand, or stop

A pilot should continue when it shows a reproducible improvement, acceptable performance on the defined risk cases, and a workable path to production. By September 27, 2026, many legal teams have moved beyond proving basic machine-learning classification and are testing more agentic workflows, such as multi-step document analysis or AI-assisted investigation planning. That does not mean the technology is dependable without oversight. The reported experience of security operations organizations is instructive: one widely discussed forecast said 70% of SOCs would pilot AI agents, but only 15% would see results. Although the use case is not identical to discovery, the lesson is relevant: a pilot announcement is not evidence of production success.

Expansion is justified when the tool’s performance holds across more than one data set and the benefit survives full workflow costs. The team should set a decision date, often within 4–12 weeks, and define advance criteria such as at least a 30% reduction in review volume, no material decline in recall, resolved security findings, and a production price that remains acceptable at forecast volume. If those thresholds are missed, the team should not keep adding features indefinitely. It should identify the cause—training data, document quality, integration, interface design, or an unrealistic baseline—and decide whether a limited revision is credible.

Stop the pilot when errors threaten privilege or production obligations, vendor controls do not match the contract, reliable ground truth cannot be created, or the cost exceeds the value of the workflow. A lack of measurable benefit is not always a failure. A 10% speed improvement may be unattractive for a high-risk matter but useful for repetitive, low-value work; a 40% improvement may still be insufficient if the tool cannot preserve required metadata. The decision should be recorded for later audit and procurement review, including what was tested, what was not tested, and which assumptions require reevaluation.

Legal and regulatory developments also affect timing. In the United States, courts and counsel continue to address disclosure of AI use, competence, confidentiality, and the reliability of generated materials, but no single rule answers every discovery workflow. In the European Union, the AI Act adopted in 2024 introduced a risk-based framework whose obligations phase in over time; deployment, contract terms, documentation, and risk classification should be checked for the specific system and jurisdiction. Organizations should avoid assuming that an EU label such as “AI” creates a safe harbor, or that a private contract shifts professional responsibility away from the lawyer.

A defensible 2026 decision framework

The strongest answer is to run a narrow, measured pilot against a real baseline, with independent quality review, human accountability, and explicit security gates. Start with one workflow and one representative data set, then reserve a blind test that the vendor has not optimized against. Set thresholds before seeing results, report denominators and error types, and test exceptions rather than only clean examples. The final report should state not just whether the technology worked, but under what conditions, at what cost, and with what unresolved limitations.

This approach also keeps the broader site context in view. AI eDiscovery can reduce repetitive review, while AI legal research and drafting tools may change research and document-production workflows, but they should not be treated as interchangeable legal services. A useful pilot is therefore an operational experiment with documented controls. It produces better procurement evidence than a demonstration, protects the client and matter, and gives the legal team a defensible basis for deciding whether broader deployment is justified.