# What elusion rate acceptance threshold should we put in our ESI protocol?

legalpdf.io · August 22, 2026

> The Direct Answer: There Is No Universal Threshold, But 2–5% Is the Working Range The most common question litigators ask when drafting an ESI...

## The Direct Answer: There Is No Universal Threshold, But 2–5% Is the Working Range

The most common question litigators ask when drafting an ESI protocol is what elusion rate they should accept before a TAR (technology-assisted review) or predictive coding workflow is deemed adequate. The honest, defensible answer is that no court, rule, or regulatory body has ever fixed a mandatory elusion rate acceptance threshold. What exists instead is a body of case law, academic validation studies, and industry practice that clusters around a workable range. In TAR 1.0 workflows validated through the control-set method, parties and courts have generally treated elusion rates below roughly 2% as strong evidence of recall adequacy, rates between 2% and 5% as acceptable with disclosure or further sampling, and rates above 5% as presumptively problematic requiring remediation. These figures are not statutory; they are conventions drawn from cases like Da Silva Moore v. Publicis Groupe (S.D.N.Y. 2012), the first judicially approved predictive coding protocol, and from the Grossman-Cormack research on predictive coding published between 2011 and 2014.

**Also worth reading:** [How do you calculate a null set elusion rate in eDiscovery, and what does the number actually mean?](https://legalpdf.io/knowledge/how_do_you_calculate_a_null_set_elusion_rate_in_ediscovery_and_what_does_the_number_actually_mean.php) · [What is an AI privilege review validation protocol and how do legal teams implement one defensibly in 2026?](https://legalpdf.io/knowledge/what_is_an_ai_privilege_review_validation_protocol_and_how_do_legal_teams_implement_one_defensibly_in_2026.php) · [What should ESI protocol AI disclosure language look like in 2026, and do I have to tell opposing counsel we're using AI in eDiscovery?](https://legalpdf.io/knowledge/what_should_esi_protocol_ai_disclosure_language_look_like_in_2026_and_do_i_have_to_tell_opposing_counsel_were_using_ai_in_ediscovery.php)

The reason no hard threshold exists is that elution rate is a function of prevalence, sample size, and the statistical confidence level you select — not an intrinsic property of a review platform. A 2% elusion rate in a collection of 50,000 documents means roughly 1,000 missed responsive documents, while the same rate in a 500,000-document corpus means 10,000 missed documents. Courts evaluating TAR disputes have consistently focused on process transparency and statistical validation rather than any single magic number. When you draft your ESI protocol, you are not choosing a threshold because a rule requires it; you are negotiating a mutually acceptable statistical standard that both sides can defend if the production is later challenged under Rule 26(g) certification obligations.

## How Elusion Rate Actually Works in a Validated TAR Workflow

Elusion rate measures the proportion of responsive documents found in the discard pile — the documents the system classified as non-responsive after training was complete. It is calculated by taking a random sample of the documents not produced (the null set), having subject-matter experts review that sample for responsiveness, and dividing the number of responsive documents found by the total number sampled. If reviewers find 8 responsive documents in a 400-document random sample of the discard pile, the point-estimate elusion rate is 2%. That figure then gets projected across the entire discard population to estimate total missed responsive documents.

The mechanics matter because the elusion estimate carries a confidence interval that widens dramatically at low prevalence. At 95% confidence, a 400-document sample yielding zero responsive documents produces an upper-bound elusion estimate of roughly 0.75% (using the rule-of-three approximation: 3 divided by the sample size). To claim an upper bound of 2% at 95% confidence with zero hits, you need a sample of approximately 150 documents from the discard pile. If your sample finds some responsive documents, the required sample size grows substantially to keep the interval tight. This is why sophisticated ESI protocols specify not just a target elusion percentage but also the sample size and confidence level used to measure it — typically 95% confidence with samples ranging from 300 to 2,000 documents depending on corpus size and stakes.

Recall and elusion are two sides of the same coin. Recall equals one minus the elusion-adjusted miss rate relative to total responsive population. A workflow targeting 75% recall will tolerate roughly a 25% elusion-relative figure only when prevalence is low; at higher prevalence, the same recall target implies far more absolute missed documents. Drafting teams frequently conflate these metrics, which leads to protocols that promise numbers neither side can actually verify. Your protocol should define each metric operationally: who performs the elusion sampling review, how disagreements are resolved, whether the sampling is stratified by custodian or date range, and what happens when the measured rate exceeds the agreed threshold.

## Comparison of Common Threshold Approaches in ESI Protocols

Different negotiating postures produce different threshold structures. The table below compares the four approaches seen most often in negotiated ESI protocols and court-approved workflows as of 2025–2026 practice.

| Feature | Fixed Numeric Cap (e.g., ≤2%) | Confidence-Bound Approach (e.g., 95% CI upper bound ≤5%) | Disclosure-Only (no cap) | Iterative Remediation Trigger |
| --- | --- | --- | --- | --- |
| Typical proponent | Producing party seeking certainty | Sophisticated parties with eDiscovery counsel | Receiving party lacking statistical leverage | Either party wanting flexibility |
| Measurement burden | Moderate — single sample suffices | High — larger samples needed for tight bounds | Low — no formal validation required | High — repeated sampling rounds |
| Risk to producing party | Low if met; breach if exceeded | Predictable; bounded exposure | Higher — open to challenge | Controlled but potentially costly |
| Risk to receiving party | May accept inadequate recall at low prevalence | Best statistical protection | No assurance of completeness | Delay in receiving complete production |
| Court reception | Accepted where justified (Da Silva Moore line) | Favored by special masters | Viewed skeptically in high-stakes cases | Common in stipulated protocols |
| Cost profile | Lowest | Highest per-validation cost | Lowest upfront, highest dispute risk | Variable, often escalating |

The fixed numeric cap is popular because it is easy to draft and easy to test, but it has a structural flaw: it ignores prevalence. In a corpus where true responsiveness is 0.5%, even a perfect review cannot be distinguished statistically from a poor one using small elusion samples, and a rigid 2% cap may force wasteful retraining cycles that improve nothing. The confidence-bound approach solves this by tying acceptance to statistical proof rather than a raw percentage, which is why it increasingly appears in protocols drafted by AI-literate eDiscovery practitioners. The disclosure-only approach — where the producing party simply reports its elusion findings without committing to a cap — survives mainly in matters where the requesting party lacks bargaining power or the data volume makes validation economically irrational.

## Practical Steps for Negotiating and Drafting the Threshold Clause

Start by estimating prevalence before you negotiate anything. Run a quick judgmental or simple random sample of 500 to 1,000 documents early in the case. If prevalence comes back above 5%, a fixed low elusion cap is realistic and you can negotiate aggressively. If prevalence is below 1%, push for a confidence-bound formulation instead, because small-sample elusion estimates become statistically unstable and a fixed cap becomes either trivially easy or impossibly strict depending on sampling luck.

Second, specify the measurement methodology inside the protocol itself, not in a side letter. State the sample size (commonly 400–2,000 documents from the discard pile), the confidence level (95% is standard), who reviews the sample (senior reviewers or the subject-matter expert, not first-pass contract reviewers), and the reconciliation procedure for disputed calls. Third, build in a remediation ladder rather than a binary pass/fail. A well-drafted clause reads something like: if the measured elusion exceeds the threshold, the producing party conducts additional training rounds and re-samples within 14 days; if the second round still fails, the parties meet and confer within 7 days to consider alternative methods such as supplemental keyword screening of the discard pile or manual review of high-risk custodians.

Fourth, decide explicitly whether elusion testing applies to all custodians or only to high-volume or high-risk custodians. Applying full statistical validation to every custodian in a 200-custodian matter can add tens of thousands of dollars in review costs for marginal benefit. Many 2025-era protocols apply full elusion testing to custodians contributing more than 5,000 documents and use lighter-touch sampling elsewhere. Fifth, preserve the validation artifacts — sample lists, reviewer codes, reconciliation logs — for the duration of the litigation plus any appeal window, since these are the documents you will need if the production's adequacy is challenged at a discovery conference or in a motion to compel.

## Common Mistakes That Undermine Elusion Clauses

The most frequent drafting error is agreeing to a threshold without specifying how it will be measured. A protocol that says "elusion shall not exceed 3%" is unenforceable in practice if it does not state the sample size, confidence level, and review protocol, because either party can manufacture a favorable or unfavorable measurement through sampling choices. A 100-document sample finding 2 responsive documents yields a point estimate of 2% but a 95% upper bound over 6%; whether that "passes" depends entirely on unstated assumptions.

The second mistake is ignoring the relationship between elusion and precision. A producing party can drive elusion toward zero simply by over-including — lowering the cutoff score so nearly everything gets produced. This inflates the receiving party's review burden with thousands of non-responsive documents and shifts costs downstream. A balanced protocol addresses both metrics, often capping elusion while acknowledging that precision targets (or at least production volume expectations) protect the receiving side from drowning in noise. The third mistake is treating the threshold as a quality guarantee rather than a statistical estimate. Even a perfectly executed 2% elusion validation means the producing party likely missed responsive documents; the clause allocates risk, it does not certify perfection. Parties who litigate as though a passing elusion score proves completeness set themselves up for sanctions motions built on false premises.

A fourth mistake appears on the AI-assisted review side: assuming that modern continuous active learning (CAL/TAR 2.0) workflows eliminate the need for elusion testing. CAL reduces the need for a formal control set during training, but the leading judicial guidance and the Sedona Conference's TAR materials still support end-of-workflow validation sampling when the opposing party demands assurance. Skipping validation entirely because the tool is newer does not survive meet-and-confer scrutiny against a prepared adversary. Finally, teams routinely forget to account for family documents and near-duplicates in elusion sampling, which can distort results by several percentage points in either direction depending on how families were coded.

## When to Act: Timing the Threshold Decision in the Case Lifecycle

The right moment to lock the elusion threshold is during the Rule 26(f) conference or the equivalent pre-discovery conference, before substantial review spend occurs. Once a producing party has invested six figures in a review workflow, renegotiating the validation standard becomes a fight about sunk costs rather than statistics, and courts are less sympathetic to mid-stream changes. If you are the requesting party and the producing party resists any quantitative standard, the fallback position is to demand transparency: require disclosure of the review method, the training statistics, and a voluntary elusion sample report, even without a hard cap.

If you are already past the 26(f) stage without a threshold in place, act before production begins, not after. Post-production challenges to elusion performance almost always fail unless the requesting party can show the producing party's process was grossly deficient, because courts apply a reasonableness standard to Rule 26(g) certifications rather than a perfection standard. The practical window for meaningful negotiation closes once documents are in the receiving party's hands. For matters involving regulatory investigations or second requests under HSR review, build the validation framework into the document review plan at the outset, since agency staff increasingly ask for recall and elusion documentation when TAR is disclosed as the review method.

## Cost Considerations and the Economics of Tighter Thresholds

Tighter elusion thresholds cost money in three ways: larger validation samples, more training iterations, and more senior-reviewer time on the sampling reviews. A single 2,000-document elusion sample reviewed at blended senior-reviewer rates of $85–$150 per hour, with roughly 15–25 documents reviewable per hour for responsiveness calls, runs approximately $12,000–$20,000 per validation round. Add retraining and re-sampling cycles triggered by failures, and a contested validation process can add $30,000–$75,000 to a mid-sized matter. Against that, the cost of an inadequate production — motion practice, adverse inference arguments, redoing the review under court supervision — routinely exceeds $100,000 and can reach seven figures in large commercial litigation.

AI-assisted review platforms have compressed some of these costs. Continuous active learning tools typically reach stable recall with fewer training rounds than first-generation TAR, and several platforms now generate automated validation sampling reports that reduce the labor component of elusion testing. Still, the human review of the elusion sample itself remains the dominant cost driver and cannot be automated away, because the entire evidentiary value of the exercise rests on human responsiveness judgments. Budget realistically: for a matter with 250,000 documents and moderate stakes, plan $15,000–$40,000 for validation activities across the lifecycle, and treat that as insurance against the far larger cost of a challenged production.

## Bottom Line for Drafting in 2026

Anchor your ESI protocol's elusion clause in three numbers: a target elusion rate (2% is the conservative market standard, 5% the outer defensible bound), a 95% confidence level, and a sample size scaled to corpus size (minimum 400 documents, ideally 1,000–2,000 for large productions). Pair the cap with a defined remediation path, disclose the measurement methodology in the protocol text itself, and preserve all validation records. Treat the threshold as a negotiated risk-allocation device grounded in the Grossman-Cormack research tradition and the Da Silva Moore judicial lineage — not as a regulatory mandate, because none exists. Parties who understand the statistics negotiate better clauses; parties who do not end up with numbers that look rigorous on paper but collapse under the first serious challenge.

## Quick answers

### Is there a court-mandated elusion rate threshold?

No. No court or rule sets a mandatory elusion rate. Figures like 2% or 5% come from case law such as Da Silva Moore v. Publicis Groupe, academic validation studies, and negotiated industry practice. Courts evaluate the reasonableness of the overall process, not compliance with a fixed number.

### How large should my elusion sample be?

For a 95% confidence level, a minimum of about 400 documents from the discard pile is common, with 1,000–2,000 preferred for large productions. With zero responsive documents found, a 400-document sample supports an upper-bound elusion estimate near 0.75%; smaller samples produce intervals too wide to defend.

### Does TAR 2.0 / continuous active learning remove the need for elusion testing?

No. CAL reduces the need for a formal control set during training, but end-of-workflow validation sampling is still recommended when the requesting party wants assurance of completeness. Most negotiated protocols retain elusion testing regardless of the training method used.

### What happens if the elusion rate exceeds the agreed threshold?

Well-drafted protocols include a remediation ladder: additional training rounds and re-sampling within a set period (often 14 days), followed by meet-and-confer if the second attempt fails. Options at that stage include keyword screening of the discard pile or manual review of high-risk custodians.

### Can a producing party game the elusion rate?

Yes, by over-including documents and lowering the responsiveness cutoff, which drives elusion toward zero while flooding the receiving party with non-responsive material. Balanced protocols address this by monitoring precision or production volume alongside the elusion cap.

Canonical: https://legalpdf.io/knowledge/what_elusion_rate_acceptance_threshold_should_we_put_in_our_esi_protocol.php
Markdown: https://legalpdf.io/knowledge/what_elusion_rate_acceptance_threshold_should_we_put_in_our_esi_protocol.php/index.md
