Direct Answer: What a Legal AI Security Review Should Decide

A legal AI security review is the documented process of deciding whether a generative-AI system may be used for eDiscovery, legal research, document drafting, or another legal workflow. The review should examine the vendor, model, data flows, user permissions, retention practices, output controls, audit evidence, and contractual remedies rather than treating “AI security” as a single product feature. Its central question is not simply whether the tool is accurate; it is whether the organization can explain what information was sent, who could access it, how outputs were produced and checked, and what happens if the system fails. For legal teams, confidentiality, privilege, work-product protection, court obligations, professional duties, and cross-border data restrictions usually matter more than a vendor’s generic claim that it uses “enterprise security.”

Also worth reading: What are the most important controls for maintaining data integrity and security in legal discovery? · What are the core requirements and security compliance standards for enterprise legal AI in eDiscovery and research? · How Do You Conduct a Realistic Legal Tech Cost-Benefit Analysis for AI Systems in 2026?

A defensible review should produce a written decision for each approved use case, identify the responsible owner, and state the conditions that require reassessment. As of 28 September 2026, that review should also account for agentic features: an AI system may now retrieve internal documents, call research or document tools, write files, or initiate software actions rather than merely answer a prompt. Ordinary accuracy testing is therefore insufficient. A 99% answer-accuracy figure, even if genuine for a measured benchmark, does not establish that the tool will reliably preserve privilege in 100% of matters. The correct output is a controlled approval, restricted pilot, remediation plan, or prohibition, supported by evidence that an accountable person can inspect later.

How to Assess Confidentiality, Privilege, and Data Handling

Start by mapping the data before evaluating the interface. Teams should record which matters use the system, what documents are uploaded, whether names, health information, financial data, trade secrets, or information subject to preservation orders enter prompts, and whether the vendor retains inputs for training or improvement. The review must distinguish data merely displayed in a browser from data transmitted to a model provider, stored in a tenant, logged by monitoring software, used to create embeddings, or available to subprocessors. For eDiscovery, the data set may contain millions of pages and an entire custodian population, so sampling limitations become important. For legal research, the input may be narrower but still reveal a client’s theory, settlement position, or unpublished strategy.

Privilege analysis should be use-specific rather than categorical. Uploading a document does not automatically destroy privilege, but a vendor’s ability to retain or reuse that document outside the engagement may create disclosure, contractual, or waiver arguments. The legal team should inspect the terms of service, data processing agreement, confidentiality clauses, deletion settings, backup schedule, support-access policy, and any provisions governing model training. It should also ask whether human support personnel can view customer content and under what approval. Encryption in transit and at rest is a useful baseline, but it does not answer who can decrypt the material or whether a subprocessor can access it under a separate policy.

A sound threshold is no production legal data until the contract and technical configuration agree. If the provider states that customer content is not used to train shared models, that promise should be documented in a binding agreement rather than accepted only from a sales page. Matters involving litigation holds, regulatory investigations, or government requests may require stricter isolation than ordinary corporate research. If a vendor cannot identify every subprocessor or cannot produce deletion and access logs, that is a reason to limit use or select another service, not a reason to assume encryption solves the problem.

How to Evaluate Security Architecture and Agent Permissions

The next step is to examine the system from authentication through execution. Determine whether customers use SAML or OIDC single sign-on, multifactor authentication, role-based access control, tenant isolation, customer-managed keys, regional hosting, and configurable retention. Ask whether integrations use scoped OAuth credentials, whether those credentials can be limited to read-only access, and whether the service can require approval before an AI agent sends email, modifies a document repository, exports a file, or executes code. A legal AI product may be secure in its native environment and unsafe when connected carelessly to a case-management, email, cloud-storage, or document-review system.

For eDiscovery and drafting workflows, permission should be narrower than the permissions held by a lawyer or paralegal. An agent authorized to search Westlaw or Practical Law does not necessarily need permission to read every internal matter. Likewise, access to a document repository should not automatically permit deletion, forwarding, or retention changes. Technical controls should include separate service identities, time-limited credentials, restricted folders, allowlisted destinations, and logs recording prompts, retrieved sources, tool calls, outputs, and administrative actions. Where feasible, high-impact actions should require human confirmation, while confidential searches should be filtered by ethical walls or ethical walls created for the engagement.

Security claims should be tested against an architecture diagram and current independent assurance reports. ISO 27001 certification, SOC 2 Type II coverage, penetration testing, and vulnerability-disclosure processes can provide useful evidence, but they are not interchangeable. A SOC 2 report may cover availability and logical access without proving that an AI system resists prompt injection or that customer data is excluded from model training. Request the report’s scope, period, exceptions, and bridge letter, then map the relevant controls to the proposed legal use. The review should also establish an incident-notification period; a practical contractual target for a legal vendor might be 24 to 72 hours for a confirmed security event, subject to the vendor’s ability to investigate without creating unreasonable delay.

How to Test Accuracy, Hallucinations, Citations, and Human Oversight

Accuracy testing must reflect the actual job the tool will perform. For legal research, evaluators should use a representative set of questions, jurisdictions, document types, date cutoffs, and negative cases in which no reliable answer exists. Reviewers should record whether each answer identifies the relevant authority, distinguishes binding law from secondary material, states the effective date, and admits uncertainty. Citation checking should be mechanical first, followed by human review of whether the cited source actually supports the proposition. A correct URL or quotation can still be placed in a misleading context, so citation existence is not equivalent to legal soundness.

For document drafting, testing should include missing facts, conflicting definitions, unusual jurisdictions, inconsistent party names, and prompts that invite unsupported assumptions. The team should measure the rate at which the system fabricates citations, invents contract clauses, changes negotiated positions, or omits required qualifiers. A reasonable pilot may begin with 50 to 100 carefully logged tasks, but the sample must be reviewed by people familiar with the relevant law. The results should be reported as rates rather than impressions: for example, “12 of 100 outputs contained an unsupported citation” is more useful than “the model sometimes hallucinated.” Any vendor benchmark should be reproduced or independently checked before it becomes part of the approval record.

Human oversight should be designed into the workflow, not added as a disclaimer after deployment. A lawyer should approve externally filed documents, negotiated language, filing deadlines, dispositive factual statements, and advice communicated to clients. For research, the reviewer should inspect the source passages and verify the current law independently. For eDiscovery, technology-assisted review should preserve defensibility through sampling, quality control, issue tracking, and a record of exceptions. Because an AI output can be fluent and wrong, review time should be budgeted in advance. If checking every answer costs more than the drafting time saved, the use case may not be economically sensible even when the system is highly capable.

Practical Steps for Conducting the Review

The first phase is a small, documented inventory of proposed tools and use cases. The review team should include legal, information security, privacy, records management, procurement, and the business unit requesting the tool. It should record the vendor’s legal entity, hosting regions, subprocessors, model providers, integrations, intended users, data categories, and expected volume. A useful approval threshold is to prohibit public or consumer AI accounts for confidential client or company data unless a formal exception and compensating controls are approved. Each use case should be classified by sensitivity and impact: public research may receive a lower threshold, while unreleased litigation analysis or regulated personal data should receive the highest.

The second phase is technical and contractual testing. Security personnel should verify single sign-on, multifactor authentication, access logs, retention controls, deletion behavior, backup deletion, encryption, tenant separation, and incident procedures. Procurement should compare the vendor’s actual commitments with the organization’s requirements, including breach notice, audit rights, confidentiality, data ownership, training restrictions, service availability, indemnity, limitation of liability, business continuity, and termination assistance. Legal should confirm that the agreement distinguishes the provider’s platform from third-party research content, integrations, and model suppliers. The team should preserve screenshots and test results because product interfaces and model versions can change without advance notice.

The third phase is a time-limited pilot with trained users and approved examples. Establish a baseline for time, spending, citation errors, rework, confidentiality incidents, and user corrections. Stop the pilot if the system sends data to an unapproved region, exposes another matter’s content, fabricates material citations, or takes an unauthorized action. After the pilot, the accountable legal owner should issue a written decision and schedule reassessment at least annually, after a major model or integration change, after a security incident, or when the use expands from research to autonomous action. A review that ends with an account being enabled is incomplete.

Comparison of Security Review Approaches and Alternatives

Organizations can compare several ways to obtain assurance. The best approach is rarely a choice between “secure AI” and “no AI”; it is a choice among vendor diligence, internal testing, contractual controls, and workflow restriction. The table below illustrates the trade-offs rather than assigning an unverified security ranking.

FeatureFull enterprise legal-AI reviewPilot with public or synthetic dataDirect use of a general AI account
Confidential legal dataPermitted only after approval and contractual reviewGenerally excludedNot appropriate
EvaluationArchitecture, contracts, logs, accuracy, and workflow testingLimited user tests and data-flow checksMinimal evidence
Typical costHigh internal effort; vendor and assessment fees varyModerate internal effort; often free to low-cost pilot toolsLow upfront cost, potentially high remediation cost
Main benefitBest basis for defensible production approvalFast, low-risk learningConvenience without a defensible control model
Main weaknessSlow and resource-intensiveDoes not establish production readinessPrivacy, privilege, and accuracy exposure
Alternatives include restricting AI to general legal questions without client facts, using on-premises or private-cloud models where feasible, purchasing from vendors with strong enterprise controls, or using conventional search and human review. A general research tool may be safer for a public statute than a document platform holding a client’s entire matter, even if both use sophisticated models. Similarly, a retrieval system connected only to an approved case-specific corpus can reduce exposure compared with a shared agent that can access every repository. These alternatives are not automatically cheaper or better; they change the amount of data, functionality, and operational burden involved.

Common Mistakes and When to Act or Stop

One common mistake is accepting security questionnaires without testing whether the answers apply to the legal product. A questionnaire may describe the vendor’s corporate environment while the purchased service uses a different cloud region, model, connector, or subprocessor. Another mistake is assuming that a confidentiality clause covers data sent to an AI model; the model provider and the legal application provider may have separate roles. Teams also fail by treating an accuracy demo as a security review, by allowing employees to paste material into consumer tools, and by giving an agent broad administrative permissions to make a pilot convenient.

The team should pause immediately after a suspected exposure, an unauthorized account creation, a model change that changes data handling, or evidence that citations cannot be reproduced. It should preserve logs, preserve the relevant agreement versions, contain access, and notify privacy or incident-response personnel. If privilege may have been disclosed, the matter owner and ethics or professional-responsibility counsel should assess notification, remediation, and waiver risk. If a regulatory deadline is approaching, the team may use the tool only for lower-risk preparation while relying on verified sources for the filing itself. A deadline is a reason to control risk, not a reason to dispense with review.

Cost is a separate consideration. Consumer subscriptions may be free or cost tens to hundreds of dollars per month, while enterprise legal platforms commonly quote prices based on users, matters, storage, search volume, or negotiated services; exact 2026 prices cannot be inferred from the research context. Internal review costs include staff time, security testing, contract negotiation, training, and ongoing monitoring. Compare those costs with the time saved and error reduction expected from the workflow. A tool that saves 20 hours but adds 30 hours of verification is not an effective business case. Obtain a written quote, define data limits, and include renewal, overage, integration, and support charges before approval.

The Minimum Approval Record and Final Recommendation

A defensible legal AI security review should leave behind a record that another authorized reviewer could understand. It should name the tool, version, vendor, intended use, approved users, data classes, hosting and integration boundaries, permitted model providers, retention period, access controls, human-review requirements, tested limitations, incident contacts, contract references, and the date of the decision. It should also record rejected features and future conditions. For example, the approval may permit research against a designated legal database while prohibiting autonomous filing, unrestricted email access, training on client content, and export to personal accounts.

The recommendation should be proportionate. Start with a narrow, reversible use case and synthetic or low-sensitivity data; expand only after the organization has evidence of accuracy, access control, deletion, and incident response. Treat agents as privileged software operators, not ordinary chat interfaces, because their ability to retrieve, write, and act creates risks that a standalone answer does not. Revisit the review when the model version, vendor ownership, data location, connector permissions, or legal function changes. A strong review does not guarantee zero risk, and it cannot turn every model output into reliable legal advice. It does something more practical: it makes the risk visible, assigns responsibility, and creates evidence that the organization made a reasoned decision before placing important legal information into an AI system.

For legal eDiscovery, the same approach should begin with collection, preservation, defensible search, responsiveness analysis, quality control, and production controls. For legal research and drafting, it should begin with approved sources, jurisdiction limits, citation verification, confidentiality settings, and lawyer approval. Those workflows have different failure modes, so a single security score cannot govern both. The final standard is whether the legal team can explain and evidence the system’s behavior, not whether a vendor can attach a favorable security badge.