AI eDiscovery pilots should measure whether technology reduces routine review effort without increasing legal risk, missed documents, unstable results, or total cost. The most useful scorecard combines efficiency, quality, user experience, defensibility, and financial performance rather than treating “documents reviewed per hour” as the sole measure. For a properly controlled pilot, teams commonly compare AI-assisted review against a defensible human-review baseline, record precision and recall on a known answer set, and examine elapsed time from collection to production. As of September 29, 2026, there is no universal industry threshold that makes an AI eDiscovery pilot successful. A better standard is a documented target approved before testing, followed by evidence that the tool met or exceeded it under representative conditions.

What Makes a Useful AI eDiscovery Pilot?

Also worth reading: What Are the Best AI eDiscovery Metrics for Measuring Efficiency, Accuracy, and Risk in 2026? · What are defensible AI document review validation metrics for eDiscovery? · What are the accepted predictive coding validation standards in eDiscovery, and how do courts and practitioners actually measure whether TAR results are defensible?

An AI eDiscovery pilot is a limited production test in which a legal team applies assisted review, document classification, clustering, redaction, search, or prioritization to a controlled dataset. It should use representative material, including routine documents, potentially privileged records, duplicates, near-duplicates, emails with attachments, spreadsheets, and documents containing specialized terminology. The test must also retain ordinary review controls, such as authorized access, chain-of-custody records, version history, and attorney supervision. AI can accelerate sorting and first-pass analysis, but it does not replace counsel’s responsibility for relevance decisions, privilege calls, production scope, and final QC.

A strong pilot answers five operational questions. First, can the system process the actual data volume within the matter schedule? Second, does it identify responsive material with acceptable recall? Third, does its precision prevent reviewers from wasting time on a large volume of false positives? Fourth, can reviewers explain and defend material decisions? Fifth, does the saving in professional time exceed software, data preparation, hosting, security review, and oversight costs? The research supplied for this answer describes growing legal use of generative AI, including document review, but it does not establish a generally valid accuracy rate or financial return.

The pilot should run long enough to produce stable results. For a narrow classification test, several weeks may be sufficient, while testing multiple workflows across business units can require 6 to 12 weeks. Teams should avoid judging a tool from a demonstration containing 100 clean PDFs because enterprise repositories often contain scanned images, embedded objects, corrupted files, unusual metadata, and millions of records. Production-like conditions are more informative than a vendor-selected sample. If a test cannot be audited after the fact, it is not a defensible pilot even if its processing speed looks impressive.

Core Efficiency and Cost Metrics

The primary efficiency metric is net review time saved, not raw processing speed. Teams should record the hours attorneys, contract reviewers, paralegals, and vendor personnel spend on the same task with and without AI. AI may process documents quickly but still require substantial time to validate clusters, investigate false negatives, correct tags, and handle exceptions. A practical pilot target might be a 20% reduction in total human review hours over 8 weeks, but the appropriate threshold depends on matter value, document population, and baseline productivity. Savings are more defensible when the pilot measures time from first collection through final quality control rather than only model inference time.

Other useful measures include cost per document reviewed, cost per gigabyte processed, and cost per responsive document identified. These figures should include subscription fees, per-seat charges, usage fees, data ingestion, OCR, translation, hosting, security review, integration, administrator time, and attorney QA. Vendors may advertise a low price per seat while charging extra for OCR, advanced analytics, API use, or data export. Contracts should therefore state maximum data volumes, overage rates, minimum terms, termination rights, and whether the customer can retrieve its data and audit logs.

A simple financial threshold is to require expected annual savings to exceed total first-year costs by a margin set by the organization. A cautious internal rule might require at least a 1.5:1 benefit-to-cost ratio, although regulated or low-volume matters may justify a smaller direct return because the pilot can improve consistency or reduce future risk. The calculation should also account for opportunity cost: if reviewers spend less time on the pilot matter, can that capacity be used on other work? A tool that saves 40 hours but adds 20 hours of validation and administration has saved 20 hours, not 40.

FeatureNarrow document-classification pilotEnd-to-end review pilotTraditional hosted-review baseline
Typical duration2-4 weeks6-12 weeks4-10 weeks for comparable scope
Best dataset5,000-25,000 representative documents25,000+ documents across several repositoriesComparable production population
Main efficiency measurePrecision, recall, correction timeNet hours and cost through productionBaseline hours and cost per document
Main riskResults do not generalizeIntegration and workflow failures obscure tool valueHigher labor cost but familiar controls
Typical success targetAt least 95% agreement on the adjudicated set20%-40% net review-time reduction with no material recall declineBaseline against which AI performance is calculated
Important limitationSmall sample may overstate accuracyMore realistic, but more expensive to runDoes not test AI-specific benefits
## Accuracy, Recall, Precision, and Quality Control

Recall and precision are the most important quality metrics for predictive review. Recall measures the proportion of relevant responsive documents that the system successfully identified; precision measures the proportion of documents labeled responsive that are actually responsive. If a reviewed set contains 1,000 responsive documents and the model finds 950, recall is 95%. If the system marks 2,000 documents as responsive but only 1,000 are correct, precision is 50%. The model can therefore appear productive while producing too much work for reviewers, or appear cautious while missing important evidence.

Neither score should be accepted without the denominator. Teams should report results by document family, custodian, date range, file type, language, and source system. Overall accuracy of 98% can conceal poor performance on scanned invoices, low-quality images, or non-English records. A minimum review sample must also be large enough to support an error estimate. As a practical starting point, teams may adjudicate at least 1,000 documents or 5% of the pilot population, whichever is greater, and include all suspected exceptions. Statistical confidence depends on prevalence, clustering, and sample design, so the legal team should work with a qualified statistician when missed evidence could affect the matter.

Quality control should include privilege, confidentiality, and false-negative testing. Reviewers need to sample negative predictions, particularly records near the model’s decision boundary, rather than checking only documents the system marked as likely responsive. Search-term tests and recall testing remain useful because AI ranking cannot establish that a collection is complete. For high-risk matters, counsel may require dual independent review of a stated percentage, such as 5% to 10%, with escalation for disagreement. The pilot report should record each error, its cause, whether a human or system created it, and what corrective action followed.

Speed, Scale, and Workflow Performance

Processing throughput is worth measuring, but it must be separated from human productivity. Useful figures include pages or gigabytes processed per hour, documents classified per hour, time to ingest a repository, and time to export a production set. Benchmarks should distinguish OCR and conversion time from model processing time. A system may process 10,000 documents per hour while requiring three hours to resolve encoding failures or repair relationships, so end-to-end elapsed time gives management a clearer view.

The workflow test should include search, prioritization, review, coding, redaction, privilege analysis, and production generation where those features are proposed for use. A tool that performs classification well but forces reviewers to copy every decision into a second system may not improve throughput. Integration with matter calendars, document-management platforms, email archives, and case-management systems can affect both cost and adoption. Track clicks, keystrokes, context switches, correction frequency, and time required to open and complete a document.

User experience belongs in the pilot scorecard because claimed savings often fail if reviewers reject the interface or distrust its output. Ask participants to score ease of use, confidence, training time, and perceived workload on a five-point scale after the test. At least 5 to 10 representative users should evaluate the workflow where practical, and all reviewers should receive the same instructions. A speed gain accompanied by weak confidence or frequent overrides may indicate that the tool is optimized for demonstration performance rather than actual legal work.

Defensibility, Security, and Governance Controls

Defensibility means more than having a model accuracy statement from the vendor. Teams should know what data was processed, where it was hosted, which subprocessors received access, how long it was retained, and whether the output can be reproduced. An audit package should include the test protocol, data selection method, model and prompt version, configuration settings, reviewer instructions, adjudicated answer set, error log, and final approval. If the system uses retrieval-augmented generation or external services, counsel must confirm that client material was not used for unauthorized training or retained beyond the agreed period.

Security testing should cover role-based access, encryption in transit and at rest, deletion, export, incident response, and tenant separation. Existing legal hold and preservation duties continue while a pilot is running. Deleting a test dataset does not justify deleting a legal hold, and selecting a smaller dataset does not remove the duty to preserve relevant information. The pilot should not send privileged or export-controlled material to an unapproved service merely because the vendor describes the product as secure.

A model update can change results without a corresponding change in the matter. Contracts should identify notice periods for material model changes, provide version logs where available, and give customers control over deployment timing. A reasonable governance threshold is zero unauthorized disclosures and 100% traceability of production files, but organizations should define that threshold through privacy, records, security, and litigation-hold reviews. These controls may slow deployment, yet they are part of the real cost and performance of AI-assisted discovery.

Common Pilot Mistakes and Better Alternatives

The most common mistake is using vendor-created training data as the sole answer set. Vendors know which cases their system performs well on, while a customer’s records may contain unfamiliar abbreviations, scanned handwriting, foreign-language material, or legacy formats. A second error is measuring “time to first decision” instead of total time to defensible completion. Others include selecting only easy custodians, changing the review population mid-test, omitting exception handling, and failing to freeze a human baseline before deployment.

AI should also be compared with credible alternatives rather than with an intentionally weak process. Options include additional staffing, improved search terms, analytics or TAR, conventional hosted review, document clustering, or a hybrid model in which software ranks records and people review the results. The correct choice depends on volume, issue complexity, sensitivity, and existing contracts. A small matter with 8,000 documents may cost less to review conventionally than to configure and govern an AI platform. A repetitive, high-volume dispute may justify a longer pilot and deeper integration.

Decision factorBuy an AI review platformImprove conventional reviewRun a limited assisted-review pilot
Data volumeHigh and stableLow or variableModerate and representative
Review patternRepetitive with usable metadataHighly bespoke judgmentUnclear productivity potential
Cost structureHigh labor cost and scalable volumeLimited staff availabilityNeed evidence before commitment
Risk toleranceTested controls and approved deploymentLower technical dependencyControlled testing is acceptable
Expected resultLower cost per reviewed document at scaleBetter staffing or search designValidated quality, savings, and adoption
These alternatives are not mutually exclusive. A team may improve custodian questionnaires and search terms while testing AI. It may retain outside review counsel for quality control even if AI handles first-pass coding. The pilot should answer whether AI is the best incremental improvement under the organization’s actual constraints, not whether AI is fashionable.

When to Act and How to Launch the Pilot

Act now when the organization has a recurring, measurable review burden; access to representative data; security approval; and reviewers who can support a controlled test. High-volume matters with repetitive responsiveness, deadline, privilege, or contract coding may offer the clearest return. Avoid immediate enterprise deployment when documents cannot be lawfully exported, data quality is poor, the underlying issue is too unstable, or no one owns the workflow. A pilot is not a cure for unclear scope, incomplete collection, poor custodian selection, or unrealistic production expectations.

A practical launch takes 6 to 12 weeks for a meaningful end-to-end test. During weeks 1 and 2, define use cases, controls, baseline measures, and success thresholds. Weeks 3 and 4 should prepare a representative corpus, establish the adjudicated set, and complete security and privacy review. Weeks 5 through 9 can run the controlled workflow, collect errors, and measure reviewer behavior. Weeks 10 and 12 should support replication, sensitivity analysis, financial modeling, and a go, revise, or stop decision. Dates should be adjusted to data complexity, but compressing each phase usually weakens the evidence.

Before launch, approve numerical gates such as no decline from the human baseline in adjudicated recall, at least 95% precision on the test set, at least 20% net review-time savings, and 100% production traceability. Those figures are starting points, not universal legal standards. More demanding matters may require higher recall, dual review, or smaller error tolerances. The decision memo should distinguish results that were statistically stable from observations based on too few exceptions and identify unresolved limitations.

Interpreting Results and Expanding Carefully

The final pilot report should separate facts from claims. State the dataset size, date range, custodians, repositories, file types, languages, review population, duration, and number of participants. Report raw and weighted precision, recall, false positives, false negatives, user overrides, processing time, net labor hours, total cost, and security events. Include confidence intervals or uncertainty ranges where appropriate. A favorable 97% score from a narrowly selected sample does not prove 97% performance across one million future documents.

Expansion should occur only after identifying whether the benefit persists under larger volume and greater document diversity. Begin with the workflow that produced the clearest savings, preserve the human baseline, and retest before adding unrelated use cases. Review results after 30, 60, and 90 days of production use, including model drift, reviewer overrides, missed responsiveness, and cost changes. A 30-day check is appropriate for stable operations; quarterly governance may suffice once the system has passed production validation.

The definitive conclusion is that AI eDiscovery pilots should be judged by net human time saved, defensible recall and precision, reviewer trust, traceability, security, and total cost. A 30% classification-speed gain is not success if it produces unacceptable false negatives, and an AI-generated answer is not a production-ready privilege decision. The strongest pilot gives decision-makers reproducible evidence while keeping qualified legal personnel responsible for scope, judgment, and approval.