What an AI Discovery Pilot Is—and What It Should Prove
An AI discovery pilot is a time-boxed test of whether artificial intelligence can safely improve particular parts of legal document review, investigation, or research. It is not a decision to transfer a matter to an autonomous system, and “discovery” here should not be confused with general legal research. A well-designed pilot selects a narrow workflow, establishes a defensible answer key, compares human and machine performance, and produces an operational decision. The central question is whether the tool performs useful work at an acceptable quality, cost, and risk level under conditions resembling the real matter. By September 2026, legal teams have access to more capable document-analysis systems, legal research products, agentic tools, and generative drafting functions than they did in 2024. That wider availability increases the number of possible pilots, but it does not eliminate the need for evaluation.
Also worth reading: How Do AI Tools for PDF E-Discovery and Legal Research Work in 2026? · What Should an Indian Law Firm’s AI Policy Cover for E-Discovery and Legal Drafting in 2026? · What Are the Best Practices for Legal Discovery in 2026?
A useful pilot usually covers 5,000 to 50,000 documents, although the correct number depends on document volume, issue complexity, and the cost of generating reliable labels. The test should reserve at least 10% to 20% of the population for human-adjudicated validation, with a larger sample where decisions are inconsistent or the document set is heterogeneous. Teams should begin with one objective, such as responsiveness review, privilege identification, issue coding, first-pass chronology construction, or retrieval of authorities. They should not test five ambitious workflows at once, because that makes it difficult to determine which capability caused a gain or failure. Success should be expressed through agreed metrics rather than a vendor’s demonstration. Depending on the use case, those measures might include recall above 95%, precision above 90%, extraction accuracy above 98%, or a reduction in review time of at least 30%. These are possible pilot targets, not universal legal standards, and the final thresholds should reflect the consequences of each error.
Choosing a Use Case According to Legal Risk
The best first use case is usually bounded, repetitive, and easy for lawyers to check. High-volume email review can work because messages often contain dates, names, requests, and coded responsiveness conventions, but mixed attachment chains and encrypted archives can complicate the results. Privilege analysis presents a different risk profile: false negatives may disclose protected information, while false positives increase review burden without necessarily causing prejudice. A pilot involving requests for production is therefore likely to require a higher recall threshold and more attorney oversight than a pilot that merely organizes documents for research. Similarly, an AI chronology can save substantial time while still requiring counsel to confirm disputed events and source attribution.
The team should classify expected errors by severity before choosing metrics. A missed responsive email may affect production compliance; a malformed date in an internal chronology may merely require correction. In low-risk coding, 90% precision might be operationally acceptable if a lawyer must verify every output anyway. For privilege screening, a target below 99% may be inadequate for an unattended workflow, even if the aggregate score looks strong. Teams should also test performance across custodians, languages, date ranges, document formats, and opposing-counsel styles. Aggregate accuracy can conceal poor results for scanned records, spreadsheets, image-only PDFs, or unusually short messages. A vendor claiming “over 95% accuracy” is not providing enough information unless the pilot explains the task definition, sample composition, and treatment of partial matches.
Agentic legal tools introduce additional questions. A system that can search a database, call other services, and draft a response is more powerful than a passive classifier, but it also has more ways to fail through bad tool selection, stale data, or unauthorized action. A safe discovery pilot should initially restrict the system to read-only access to approved repositories, disable external transactions, and require a human approval gate. The evaluation should record not only answer quality but also whether the system cited correct evidence, respected data boundaries, and stopped when information was insufficient. A high answer score cannot compensate for a system that retrieves an out-of-scope document or acts without authorization.
Building the Test Data and Evaluation Method
A reliable pilot starts with a representative document population rather than a vendor-selected demo folder. The sample should preserve the natural proportion of emails, attachments, spreadsheets, PDFs, images, duplicates, near-duplicates, and encrypted or unreadable files found in the matter. Counsel should collect a stratified random sample, document the sampling method, and create an answer key through independent review. If budget is limited, reviewers can adjudicate disagreements between two experienced attorneys, after which a third reviewer resolves uncertain cases. The answer key should define categories precisely—for example, whether a document mentioning a topic is responsive, or whether privilege depends on the purpose of communication rather than a keyword such as “legal.”
The team should compare at least three baselines: the existing human process, keyword search or ordinary search tools, and the proposed AI-assisted process. A vendor should not be judged only against doing nothing. Automated search may already achieve high recall at low cost, while a trained reviewer may outperform a general-purpose model on a specialized issue. The pilot should also include a control sample that no system sees in advance, because repeated testing on the same labeled documents can encourage prompt tuning that does not generalize. A 60-day evaluation can use a 20-day setup period, 20 days of controlled testing, and 20 days of error analysis, although procurement, security review, and attorney time can extend that schedule to 90 or 120 days.
Results should be reported by category and error type, not only as one accuracy number. Precision measures how much of what the system selected was correct; recall measures how much of the material that should have been identified was found. For extraction, teams may report field-level accuracy and document-level all-fields accuracy. For legal research or drafting, reviewers should test whether citations support each proposition, whether quotations match the source, and whether the system distinguishes binding authority from commentary. Citations that look plausible but do not support the claim are especially dangerous. As a practical safeguard, any externally verifiable proposition should be checked against the primary source, and the pilot should fail if the tool fabricates authority or silently mixes current and historical law.
Security, Privilege, and Human Oversight
AI discovery should begin inside the same access controls used for the underlying legal matter. Data should be encrypted in transit and at rest, and vendors should explain whether customer documents are used to train shared models. Contract language should cover retention, deletion, subprocessors, breach notification, audit rights, data location, model changes, and the return of materials at termination. The procurement team should distinguish between a vendor’s standard product controls and any custom configuration required for a particular workspace. Claims about “enterprise security” do not by themselves answer whether privileged information is isolated from other customers or whether prompts and outputs are retained for support purposes.
Privilege review requires special care because uploading material to an external service can itself raise questions about confidentiality and waiver. Counsel should evaluate the governing jurisdiction, client agreements, court orders, and professional duties rather than rely on a universal rule. A pilot may proceed with synthetic, redacted, or lower-sensitivity data while security and legal questions are resolved. Access should be role-based, and the AI system should not receive credentials that permit it to email, file, transfer money, or modify evidence. Every proposed production set, privilege designation, or final research conclusion should remain subject to an identified human decision-maker.
The oversight design should specify what the reviewer sees. If the system merely returns a label, reviewers may accept it without examining the document. If it highlights passages, explains the basis, and links to source pages, verification may be faster—but an explanation can also be confidently wrong. Reviewers should receive training on likely failure modes and should be told that the system is fallible. The pilot should measure override rates, inter-reviewer agreement, time spent correcting output, and whether automation actually reduces total labor. A system that achieves high recall but sends 80% of the matter to lawyers for correction may not be commercially attractive. Conversely, a system that does not automate the final decision can still be valuable if it improves prioritization, search quality, and consistency.
Comparing the Main Implementation Options
Legal teams can buy a managed discovery platform, use a general-purpose AI service with approved legal workflows, or build an internal system around models and infrastructure. The choice should depend on matter volume, data sensitivity, technical capacity, and the need for specialized review—not on a general belief that AI is superior. Managed platforms commonly provide integrations, administration, analytics, and vendor support, which can shorten implementation time. Their weakness may be less flexibility, recurring per-user or per-document fees, or difficulty exporting data and audit logs. A legal-specific research or drafting product may offer better source integration and citation controls, but it may not replace document review features.
| Feature | Managed AI discovery platform | General-purpose AI service | Internal build |
|---|---|---|---|
| Setup | Usually fastest; vendor configures workflows | Moderate; security and prompt controls required | Slowest; requires engineering, legal, and security teams |
| Best fit | Large recurring review matters | Narrow experiments or research assistance | Specialized, high-control workflows with sustained demand |
| Cost profile | Subscription, processing, hosting, and support fees | Usage-based model, storage, integration, and review costs | Model, cloud, engineering, maintenance, and opportunity costs |
| Data control | Stronger if contracts and architecture are well designed | Depends heavily on contract, region, and product settings | Highest potential control, but operations remain the customer’s responsibility |
| Main risk | Lock-in and opaque scoring | Unapproved data handling and inconsistent controls | Internal capability gap and long-term maintenance burden |
| Typical pilot | 30–90 days | 2–8 weeks for a narrow test | 3–12 months for a production-grade system |
Cost, Pricing, and Expected Returns
There is no responsible single market price for an AI discovery pilot because pricing depends on volume, hosting, integrations, and whether labor is included. A small legal-team experiment may cost roughly $5,000 to $25,000 for tools, security review, sample creation, and attorney evaluation. A managed pilot involving 25,000 to 100,000 documents may range from $25,000 to $150,000 or more, particularly when review labels, data preparation, and human adjudication are included. Production deployments can be cheaper per document at high volume but still require licensing, hosting, implementation, and ongoing quality control. General model usage may be priced by tokens, requests, or document volume, but those figures can be unstable and should be confirmed directly with the provider.
The correct return calculation is based on avoided review time and improved quality, not merely the number of documents processed. If attorney review costs $250 per hour and a pilot saves 100 hours, the gross labor benefit is $25,000 before software, implementation, correction, and risk costs. If the system creates 40 hours of verification work and the project costs $60,000, the net benefit may be negative even when the model’s benchmark score is impressive. Teams should run the pilot with a preapproved budget ceiling, such as $50,000 for a 60-day test, and include a stop rule if projected review time exceeds the baseline by 20% or if a critical security control remains unresolved. Savings should also be discounted when the result depends on scarce senior-lawyer attention.
Legal research and drafting can produce different economics. A research tool that saves 30 minutes per memo may be worthwhile for five lawyers, while a discovery system that processes 500,000 documents may justify a larger implementation. The team should compare the product’s total workflow cost with the cost of improved consistency and faster turnaround. Public claims about productivity gains should be treated as hypotheses, not guarantees. The supplied 2026 research context includes examples of AI moving from pilots to operational payoff, but those examples do not establish that every legal team will achieve the same result. A credible business case should state assumptions, include human review, and show a conservative scenario as well as an optimistic one.
Common Pilot Mistakes and How to Avoid Them
The most frequent mistake is defining success as a dramatic demonstration rather than a repeatable workflow. A tool may perform well on 20 clean PDFs and fail on a custodian’s spreadsheet or a scanned image. Another error is beginning with production data before privacy, retention, and access rules are settled. Legal teams also tend to ask vendors to mark documents without supplying enough examples, leaving two reviewers to interpret the same category differently. The answer key should therefore be written as a decision protocol, with examples, exclusions, and escalation paths. Otherwise, disagreements may look like model errors when they are actually category ambiguity.
Teams should resist “AI-only” targets that remove the reviewer from the loop. A discovery tool is not the same as a lawyer, and its output may be used for production, privilege, or a client report. Set human approval thresholds based on risk, not novelty. Do not use legal research tools as substitutes for checking primary authority, and do not allow a system to summarize opposing counsel’s arguments without preserving source passages. Keep prompts, model versions, retrieval settings, and evaluation results so that a later decision can be reconstructed. If a vendor cannot explain how a result was produced or provide usable audit logs, that is a procurement concern even when the answer looks correct.
Another mistake is ignoring change management. Reviewers may distrust the tool, or senior lawyers may spend hours correcting outputs because the interface does not fit existing matter teams. Include a small reviewer group from the beginning, measure time per document, and provide feedback sessions during the pilot rather than after it ends. Set a clear decision date: continue, modify, pause, or terminate. “We will revisit later” allows an expensive experiment to become an unowned production dependency. The pilot report should identify what the system did well, where it failed, what remains untested, and what conditions would justify a larger deployment.
When to Act, Expand, or Stop
A legal team should act now when the workflow is sufficiently defined, representative test data is available, and the matter owner can assign accountable reviewers. Waiting is sensible when documents contain unusual formats, the relevant legal standard is unsettled, or the proposed deployment would send confidential material to an unapproved service. A narrow pilot on synthetic or redacted data can still test prompt design, interface usability, and review procedures, but it cannot establish performance on the real corpus. By September 2026, teams should at least have a documented inventory of potential tools, a security questionnaire, and a list of workflows that are not suitable for automation.
Expansion should follow evidence rather than enthusiasm. Continue beyond the pilot only if the tool meets the agreed recall and precision thresholds, produces correct supporting evidence, does not create material security problems, and saves enough total time to justify its cost. Require a second validation set before production use, and sample results after deployment. A tool that works for English email but not multilingual contracts should be restricted accordingly. If performance is weak, determine whether better extraction, search, or workflow design can correct it; not every failure requires a new model. Teams should also consider non-AI alternatives such as improved search, standard deduplication, better custodian questionnaires, or redesigned review protocols.
Stop the project when the system cannot reliably identify the relevant material, when its explanations are not trustworthy, when the organization cannot meet contractual and ethical duties, or when the business case depends on counting review time without counting correction time. A negative result can still be useful: it may prevent a production mistake, redirect spending toward conventional technology, or clarify a legal workflow that needed redesign. The most authoritative conclusion is that AI discovery is not a universal replacement for lawyers. It is a controlled engineering and legal judgment process. The organizations most likely to benefit will be those that define the task, test the tool against real-world evidence, protect sensitive information, and retain final human authority.