What Outside Counsel AI Guidelines Should Actually Do

Outside counsel AI guidelines should define who may use generative AI, what data may be processed, which workflows require review, and how work performed with AI can be verified. They are not merely technology policies; they allocate professional responsibility between the in-house legal team and outside lawyers. A sound policy treats legal research, document drafting, and eDiscovery as distinct activities because each has different accuracy risks, confidentiality concerns, and evidentiary requirements. As of September 24, 2026, many organizations still rely on general confidentiality clauses that were written before public AI chatbots could retain prompts, search case law, or analyze thousands of documents. The appropriate response is a permission-based framework with named owners, documented tool reviews, and measurable approval steps. The goal is not to prohibit AI, but to prevent unapproved systems from receiving privileged, personal, regulated, or material nonpublic information.

Also worth reading: What should ESI protocol AI disclosure language look like in 2026, and do I have to tell opposing counsel we're using AI in eDiscovery? · What Are the Essential Enterprise Legal AI Compliance Protocols Required for Document Drafting and eDiscovery in 2026? · How Do Legal Teams Measure ROI Metrics for Artificial Intelligence eDiscovery Tools in 2026?

A useful rule separates four levels of use: no AI, public-information research, internal confidential work, and restricted litigation or investigation data. Public legal research may be permitted when no client information appears in a prompt, while opposing-party documents generally should not be submitted to a public service. A contract attachment is more reliable than a verbal assurance from a lawyer that a particular tool is secure. It should identify approved tools, prohibit human review as the sole quality control, require disclosure of material AI assistance, and preserve the lawyer’s duty to check citations, quotations, calculations, and factual assertions. Organizations that apply those controls have a clearer record of care than those that merely announce that AI is allowed or forbidden.

Why Traditional Outside Counsel Instructions No Longer Cover the Risks

Older billing and confidentiality provisions focused on who accessed a document, how it was stored, and whether it could be reused. Generative AI changes the question because information may be processed through a third-party service, retained in logs, used to improve a model, or exposed through a prompt embedded in an exported file. Confidentiality can be lost through screenshots, summaries, metadata, or an attachment copied into a model-generated response. Ordinary litigation holds do not automatically tell a vendor how machine-generated analysis must be handled. The lawyer remains responsible for the work even when a tool identifies authorities, proposes a draft, or ranks documents within a large review set.

The risks vary by task. A wrong case citation is frustrating but usually visible during legal research; a fabricated quotation, invented holding, or distorted chronology can be carried directly into a brief. Drafting errors may survive because language is fluent and the reviewing attorney assumes the model performed the verification. In eDiscovery, an apparently precise prediction may encode a mistaken assumption about responsiveness, privilege, or family relationships. AI can also transform volume into a problem: a model may review 500,000 documents in less time than a person reviews 500, but speed does not establish defensibility. That is why guidelines should be proportional to the decision being supported rather than to the novelty of the software.

Law.com’s discussion of Ohio’s AI ethics guidance, an Epstein Becker Green roundtable on AI and outside counsel guidelines, and Thomson Reuters materials on AI-assisted legal drafting all point toward a governance problem rather than a purely technical one. The July 16 New York roundtable illustrates that legal teams are actively comparing rules, but publication of principles does not by itself create an enforceable standard. Each organization must translate those ideas into contract language, matter protocols, and review records suited to its own practice areas. The key shift is from controlling tools to controlling workflows and accountability.

Core Provisions to Include in Outside Counsel AI Guidelines

The first provision should establish that counsel remains fully responsible for client work and professional judgment. It should prohibit delegating legal judgment to a system, treating an AI response as an authority, or using unreviewed output in a filing, advice memorandum, transaction document, or discovery response. Counsel must validate every quotation, citation, date, party name, numerical calculation, and material factual statement. For research, the instruction should require checking the primary authority in a reliable database or official reporter rather than accepting the model’s synthesis. For drafting, it should require comparison against the governing record and applicable law. These are stronger controls than asking lawyers to “use professional judgment,” because they identify the review actions that decision makers can later audit.

The second provision should classify information and restrict access accordingly. Public statutes, regulations, and opinions can normally enter an approved research tool, although the service’s retention terms still matter. Client strategies, employee data, health information, financial information, and privileged communications require a separately approved environment. Opposing-party material, expert reports, investigation files, and sealed court records often require the highest level of protection. Organizations should document whether prompts, outputs, embeddings, and training data are retained, who can access them, where processing occurs, and whether the information can be used to train a general model. A binary label such as “confidential” is not enough; access should depend on role, matter, and purpose.

The third provision should govern approved and unapproved tools. An approval record can be short: service name, intended uses, data categories, security review, contract terms, and the person who approved it. A business-oriented tool used only for public legal research is not automatically equivalent to a private analysis platform connected to a client data room. Shadow AI is a workflow failure when employees cannot find a safe option or do not know what they may enter. Therefore, counsel should not evade the rules by selecting an unapproved service when the policy is inconvenient. Reported violations should be assessed for client harm, privilege exposure, and remediation rather than treated as a mere etiquette issue.

Rules for AI-Assisted Legal Research and Document Drafting

Legal research guidelines should permit AI to generate search concepts, issue trees, and candidate authorities only when the lawyer confirms the underlying sources. A model may be useful for identifying terminology across jurisdictions, but it may also blend rules from different states, treat a dicta statement as a holding, or invent a decision. A reasonable workflow is to use the model for orientation, search in an authoritative database, open each cited case, and record the proposition for which it is being used. The final work product should not cite a source that merely appeared in the model response. If a proposition cannot be found in a retrievable authority, it should be removed even if the draft reads well.

Drafting rules need different controls because the likely error is not always an invented authority. A document may have a valid legal citation attached to the wrong clause, a defined term that appears inconsistently, or a representation that the source record does not support. Contracts, pleadings, policies, and board materials should therefore receive a source-by-source check of material facts and obligations. Companies can establish a review threshold, such as any AI-generated section containing 100 or more words, but a numerical trigger does not replace a risk assessment. A short description of settlement terms may carry more risk than a long background section. Guidance should prioritize business effect, filing deadlines, and the cost of correction rather than document length alone.

Disclosure rules must be drafted carefully. One organization may require notice in an internal matter log, while another may require disclosure to opposing counsel when AI affected a filed document or discovery communication. The policy should identify the circumstances, recipient, timing, and record rather than impose an unlimited warning whenever any tool was used. A blanket disclaimer can obscure a meaningful use without informing anyone about the actual process. Conversely, a lawyer’s private use of AI as an ordinary research aid may not require a public disclosure in every jurisdiction. A qualified lawyer should evaluate the applicable professional rules and court orders for the specific matter.

Specialized Controls for AI in eDiscovery

Ediscovery presents higher data-volume risks and therefore deserves a separate section. Counsel should state whether AI may assist with search-term development, document classification, privilege analysis, near-duplicate detection, translation, summarization, and quality control. It should prohibit the system from making final privilege calls or irreversible privilege waivers without attorney review. Any model used in review should be evaluated for its recall, precision, and performance across languages, document types, and known error patterns. The acceptable threshold cannot responsibly be a universal percentage; organizations should determine it from the matter’s risks and governing orders.

A defensible evaluation can use a test set of at least 1,000 documents, divided among clear positives, clear negatives, and a boundary group that experienced reviewers dispute. If the tool misses 10 percent of known privilege documents, that may be unacceptable in a sensitive investigation; the same rate may be irrelevant to a routine mailing search with independent review. The resulting measurement should inform whether the system is used for ranking, first-pass review, issue coding, or merely a search aid. A reported 90 percent accuracy rate is not interpretable without knowing which class was measured, how errors were defined, and how many documents were excluded. Vendors may express performance as a single metric, but legal teams need operational detail.

Human sampling should continue after deployment. A 5 percent quality-control sample of a 100,000-document population is 5,000 documents, which could be expensive and may still not be statistically sufficient for every category. Counsel should therefore distinguish random sampling from targeted sampling of low-confidence, high-value, and potentially privileged material. The review record should identify the model version, configuration, review thresholds, and changes made during processing. Reproduction matters because a new model or changed prompt can alter results even when the underlying collection stays the same. AI-assisted eDiscovery should be treated as a documented review method, not as a substitute for the attorney’s continuing responsibility for the produced record.

Comparing Governance Approaches and Tool Categories

Organizations generally have three policy choices, and each carries different costs and control levels. The best option depends on client data, matter sensitivity, and how much verification the organization can perform. Cost figures below are illustrative planning ranges rather than quoted market prices, because vendors price by users, documents, storage, or transaction volume. Before purchase, counsel should obtain current pricing in writing and confirm whether minimum commitments, data egress, premium support, and audit rights are included.

FeaturePublic generative AIApproved enterprise assistantPrivate or matter-specific systemHuman-only method
Suitable inputPublic laws, general questions, fictional examplesApproved client information within configured limitsHighly sensitive, regulated, or sealed materialAny authorized material under ordinary controls
Main advantageFast access, broad coverage, low entry costShared workflows, access controls, vendor supportGreater control over data and configurationHighest traditional control, slowest high-volume processing
Verification burdenHigh for citations, quotations, and hallucinationsMedium to high, with review and loggingMedium to high, because model errors still occurLow for generation, high for labor and consistency
Illustrative costOften $0 to $200 per user monthlyRoughly $100 to $500 per user monthlyCustom; often $50,000 to $500,000 or more per yearDriven mainly by lawyer and review-team hours
Main weaknessUnapproved retention and weak confidentiality controlConfiguration may be misunderstoodHigh cost, implementation burden, and limited transparencyPoor scalability and routine administrative work
No category makes review unnecessary. A private system can produce a confident error, while a public model can be used safely for questions that contain no client information. The decision should be made task by task rather than by awarding a contract to the vendor with the most impressive demonstration. A second alternative is to prohibit external AI entirely and obtain added human research, drafting, and review support, but that approach can increase fees and delivery time without eliminating hallucinations or missed documents. The practical middle ground is a limited public research use, an approved platform for internal matters, and human review before any consequential use.

A Practical Implementation Process for Legal Teams

Begin with the matters carrying the greatest exposure, including active litigation, internal investigations, regulatory inquiries, and transactions involving sensitive intellectual property. A cross-functional group can include legal operations, security, privacy, records management, procurement, and at least one practicing attorney. The group should map the actual tools being used and interview both lawyers and support personnel, since workers may not distinguish a public chatbot from a managed research product. A 30-day discovery phase can capture tool names, data categories, decision points, and failure reports without pretending that the organization already knows its risk profile. That phase should produce a policy that reflects real behavior rather than the systems administrators expect employees to use.

Next, create a short approval workflow with service tiers and named decision makers. A request should identify the model, intended task, information entered, deployment level, and users who will review the output. Security and privacy personnel can evaluate contract and data handling, while a lawyer determines the professional-review requirement. A decision should be completed within 10 business days for ordinary tools and a shorter period for urgent litigation needs. Tools used without approval should be reported promptly so the legal team can assess exposure. After 90 days, analyze adoption, incidents, review defects, and time savings rather than counting the number of AI users.

Training should use real, approved scenarios and include failure analysis. A 60-minute session that only describes features is unlikely to change work habits; a two-hour exercise showing a fabricated citation, an unsupported chronology, and a privilege error can. The trainer should ask participants to identify what went wrong, what evidence would correct it, and at what point a workflow should stop. Counsel should then revise templates, research checklists, and matter instructions so compliant behavior is convenient. Adoption within the first year should be assessed through documented review quality, not a target such as having 80 percent of lawyers use AI. A high usage number combined with weak verification is a warning sign, not a success metric.

Common Mistakes That Make Policies Weaker on Paper

The most common mistake is issuing a general ban while leaving employees to infer that informal experimentation is acceptable. This creates shadow AI without collecting the information needed to improve the rules. Another error is treating confidentiality as the only issue. A tool may handle data appropriately while still generating inaccurate research, an unsupported summary, or an inconsistent contract clause. A third mistake is assuming vendor assurances replace an attorney’s verification, particularly when terms change after procurement. Policies should include an annual review and a trigger for reassessment after a material product or contract change.

Organizations also fail when they measure success by time saved alone. A review that is 50 percent faster but introduces a 3 percent privilege error may be worse than a slower process, especially if the error reaches production. Conversely, a 2 percent error rate in a low-risk issue-coding task may be acceptable if the output receives sampling and the workflow is reversible. The correct metric depends on consequence. A reasonable dashboard can track verified citation corrections, privilege sampling results, unauthorized-data incidents, review time, and the percentage of outputs receiving documented human approval. A dashboard with five measures is more useful than a policy promising 10,000 productivity hours without a baseline.

Finally, the policy should not apply one workflow to every lawyer or jurisdiction. Support staff may need different permissions from partners, and public-sector or court rules may impose requirements that private practice does not. Assigned in-house counsel can adapt the guidelines for a matter while preserving the core controls. If outside lawyers are asked to follow client rules that conflict with their own ethical or contractual obligations, they should be permitted to raise the conflict before transmitting information. That preserves trust and gives the client a chance to choose a safer tool or a more supervised process.

When to Act, How to Budget, and What to Measure

An organization should act before the next high-sensitivity matter begins, even if it is not buying software immediately. Within 30 days, identify high-risk workflows and stop the transmission of client data to unapproved services. Within 60 days, publish interim rules, designate a decision owner, and require human verification for research citations and discovery conclusions. Within 90 days, complete a pilot on one matter or document class, then document the results. A 180-day review can determine whether the policy improved quality, reduced processing time, and produced a usable audit record. Waiting for a public enforcement action or a major client incident gives the organization little control over the evidence created during the interim.

Budgeting should include more than license fees. A $200 monthly subscription per lawyer can become a substantial annual cost when multiplied across 100 users, and a private platform can carry a six-figure implementation and support commitment. A small pilot may cost less than adding several contract lawyers to conduct repetitive review, but that comparison is invalid if the tool’s error rate is not measured. Include legal review, security assessment, procurement, training, evaluation data, and ongoing monitoring in the calculation. Ask for pricing that distinguishes standard review, advanced reasoning, storage, integrations, retention, and deletion; the same headline “per user” price may conceal different usage limits.

After six months, reasonable targets include 100 percent of material AI use being assigned to an approved category, zero known incidents involving unapproved confidential data, and a documented check of every AI-supported legal authority or production decision. Numerical quality targets should be set by task rather than copied from a vendor. For example, a research workflow might require 100 percent citation verification, while an eDiscovery pilot might establish a statistically justified precision threshold and a recall threshold approved by the responsible attorney. The decisive question is whether leadership can explain what happened, who checked it, and what changed when the tool was wrong. If it cannot, greater adoption should wait.