Direct Answer for Indian Legal Teams
Indian legal teams should treat AI used for eDiscovery, legal research, document drafting, and matter management as a governed information-processing system, not as an autonomous lawyer or ordinary software feature. The minimum control set is role-based access, approved tools, data classification, India-specific data-protection compliance, human verification, audit logs, retention limits, vendor due diligence, and a documented incident response process. For eDiscovery, controls must additionally cover source-system integrity, defensible collection, chain of custody, privilege review, and reproducible document selection. For legal research and drafting, every material proposition, quotation, citation, calculation, and deadline must be checked by a qualified lawyer before external reliance.
Also worth reading: What Are the Best Legal AI Controls for eDiscovery and Legal Work in 2026? · What are the most important controls for maintaining data integrity and security in legal discovery? · What Are Legal AI Audit Controls, and How Should Law Firms Implement Them in 2026?
India does not yet have a single horizontal law that regulates legal AI in the same way the EU AI Act regulates certain AI systems. Instead, several laws apply together, including the Digital Personal Data Protection Act, 2023 and its 2025 rules, the Bharatiya Nyaya Sanhita, 2023, the Information Technology Act, 2000 where relevant, professional confidentiality duties, court procedures, and sector-specific regulation. The Bharatiya Sakshya Adhiniyam, 2023 also affects electronic evidence. As of 29 September 2026, therefore, “compliance” should not be reduced to locating one AI statute; it requires a documented assessment under every law that governs the underlying data and activity.
A practical rule is to assign risk according to three variables: what data enters the system, what the AI is allowed to do, and how consequential its output may be. Processing publicly available case law presents a different risk from uploading a client’s medical, employment, financial, or litigation files. Generating a search query is different from recommending a filing strategy or communicating a legal conclusion to a court. The controls should become stricter as data sensitivity, autonomy, and decision impact increase.
The Indian Legal and Ethical Control Framework
The Digital Personal Data Protection Act, 2023 establishes obligations for specified personal-data processing, including notice, lawful and legitimate use, purpose limitation, security safeguards, data-subject rights, and accountability. Its implementation timeline matters: the Act received assent on 11 August 2023, while the Digital Personal Data Protection Rules, 2025 were notified on 14 November 2025, with staged commencement provisions. Legal teams must verify the provisions in force on the processing date rather than assuming that every future obligation is already operative. Court exemptions and state-instrumentalities questions also require matter-specific analysis.
Legal confidentiality creates a separate control layer. A vendor’s promise that it does not train a public model does not automatically resolve issues involving privileged documents, client identity, trade secrets, work product, personal data, cross-border transfers, or onward access by subcontractors. A law firm should identify the data controller or fiduciary-like entity role, define processing purposes, restrict secondary use, and record retention and deletion periods. It should also assess whether a foreign server or support team means that personal information is transferred outside India.
Professional responsibility remains human-led. AI can reproduce an inaccurate authority, fabricate a citation, omit a limitation, expose confidential information through prompts, or make a recommendation based on incomplete instructions. Courts and regulators should not be asked to treat a generated answer as the work of an independent legal professional without checking the person who supplied it. The prudent standard is informed human judgment: a lawyer must be able to explain the purpose of the tool, the data submitted, the verification performed, and the reasons for accepting or rejecting its output.
Norms are moving faster than statutes. Indian Express reporting on “rogue” AI agents and law-firm controls, together with the White & Case India regulatory tracker, illustrates the need to monitor both formal regulation and enforcement expectations. These developments are not themselves legislation, but they help legal teams anticipate procurement, client, and professional-conduct questions. Regulatory trackers should be treated as monitoring aids rather than substitutes for current statutes and official guidance.
Data Governance for AI eDiscovery
E-discovery begins before an AI platform reviews any document. Controls must preserve the evidentiary record: collect from approved sources, preserve metadata, document search terms and filters, record exceptions, and maintain chain of custody. The Indian Evidence Act was replaced by the Bharatiya Sakshya Adhiniyam, 2023, and parties must comply with applicable electronic-record requirements and court directions. An AI-generated ranking does not excuse a party from producing relevant material, explaining its methodology, or meeting a disclosure deadline.
A defensible AI-assisted review process normally has five gates. First, the matter owner defines the issue, custodians, date range, document families, and legal criteria. Second, an independent validation set measures whether the platform is retrieving the expected population. Third, privileged, confidential, and restricted records are isolated under stated permissions. Fourth, reviewers examine the technology-assisted results rather than accepting automated selections. Fifth, the production log records the software version, settings, exceptions, reviewer decisions, and final output.
The software should support, rather than silently replace, legal judgment. “Continuous active learning” can improve prioritization, but it can also alter later selections or affect reproducibility if prior labels change the model. Teams should decide whether such learning is enabled, freeze the production environment, export decision records, and rerun a controlled sample before relying on a changed result. Quantitative measures such as recall, precision, deduplication rate, reviewer agreement, and error severity should be reported without reducing quality to a single percentage.
Privilege remains a high-risk area. A system that flags documents for likely privilege is not the same as a system that decides privilege. Training or testing on client documents can be unacceptable even where ordinary processing would be allowed, and privilege rules may differ between jurisdictions. Indian litigation teams should involve responsible lawyers early, test false-negative and false-positive behavior, and document why a suggestion was accepted or rejected.
Controls for Legal Research and Document Drafting
Legal-research systems should be restricted to sources the organization has independently approved. The tool should display the full authority, court, date, paragraph number, and current status, and it should distinguish primary sources from commentary. A response that includes a plausible-looking citation is not evidence that the source exists. Before use, the lawyer must open the authority in an official or trusted repository and confirm that the proposition actually appears at the cited location.
Prompt design is a security control, not merely a convenience feature. Prompts should prohibit requests to reveal system instructions, hidden training material, credentials, or other clients’ information. Inputs should be inspected for personal data, litigation strategy, trade secrets, and malicious documents. A hostile PDF can contain instructions aimed at an AI system, so document ingestion should separate text extraction from any permission to follow commands found inside the document. The model must never treat content inside evidence as an instruction from the user or legal representative.
Drafting tools create additional duties around authorship and reliability. The output may be used for an internal outline, a first draft, a due-diligence issue list, or a filing intended for court; these uses should have different approval requirements. A court-ready document needs substantive review for jurisdiction, limitation periods, local rules, citation accuracy, party names, numerical consistency, confidentiality, and compliance with applicable advocate rules. Teams should not use aggregate time saved as the only benefit metric, because a faster but incorrect filing can be worse than no automation.
Use approved templates and version control to reduce variation. A template library can encode approved clauses, defined terms, house style, and escalation rules, but the legal content still requires professional review. Generative drafting should be tested on known matters before production use, with a benchmark set containing deliberately difficult cases. A 95% agreement rate on routine clauses is useful but does not establish reliability on a novel, high-value issue; severity-weighted testing is more informative than an average score.
A Practical Control Model for Firms
The strongest approach is a tiered control model. Basic systems can handle public statutes, approved research platforms, and fictional exercises. Restricted systems may process one matter’s non-public documents under named-user access. High-impact systems, such as those supporting bulk evidence review, medical-record analysis, employee decisions, or court filings, require enhanced testing, legal approval, and incident reporting. This model prevents a small firm from buying the same expensive controls for every use case while ensuring that consequential deployments receive stronger protection.
| Feature | Basic legal AI use | High-risk eDiscovery or drafting use |
|---|---|---|
| Permitted data | Public law, fictional matters, internal non-sensitive notes | Privileged records, client files, personal or regulated data |
| Access | Named users and approved devices | Matter-based groups, least privilege, multifactor authentication, periodic recertification |
| Human review | Spot-check results and citations | Review by qualified reviewers with defined quality thresholds and exception escalation |
| Auditability | Usage log and version record | Full input, retrieval, model, prompt, citation, reviewer, and export history |
| Vendor terms | No training on customer inputs; defined retention | No secondary use, India-transfer assessment, subprocessors, deletion evidence, audit rights |
| Performance standard | Test a representative benchmark | Matter-specific validation, recall and precision testing, privilege testing, documented retesting |
| Incident response | Correct output and report internally | Contain data, suspend processing, preserve logs, notify relevant parties, and perform root-cause review |
The policy should include acceptable and prohibited uses. Examples of prohibited uses may include uploading privileged communications to unapproved public chatbots, using consumer accounts for client data, asking a model to predict a judge’s decision, or sending confidential material through personal accounts. Exception requests should identify the business purpose, data, proposed controls, expiry date, and person accepting residual risk. A blanket exception without an end date is not a control.
Vendor Review, Procurement, and Contractual Protections
Vendor assessment should occur before trial data is uploaded. A sales demonstration is not permission to import production material. The security questionnaire should ask where data is stored, whether it is used for model training, who can access it, which subprocessors are involved, how long it is retained, how deletion is verified, whether logs can be exported, and what happens after termination. Organizations should also examine encryption in transit and at rest, tenant separation, multifactor authentication, vulnerability management, backup practices, and breach-notification timing.
Contract language should convert general assurances into enforceable duties. Relevant terms include purpose limitation, prohibition on training on customer data without written consent, defined retention and deletion, confidentiality, security standards, subprocessor controls, India-specific data-protection responsibilities, assistance with data-subject requests, incident cooperation, audit evidence, suspension rights, and transition assistance. A firm should know whether the provider can support a legal hold, not merely an ordinary operational backup. If the service becomes unavailable, the organization must still retrieve its matter data and access its audit history.
Pricing varies by architecture and scale, so a universal figure would mislead. Public research or drafting subscriptions may cost roughly US$20 to US$200 per user per month, while enterprise legal-research seats can range from several hundred to more than US$1,000 per user per month. eDiscovery services priced per gigabyte can range from cents to several dollars per GB, with review, hosting, export, technology, and forensic services often charged separately. Indian enterprise implementations may also require one-time integration, identity, security, and governance costs. The evaluation should compare total cost over at least 24 to 36 months, not only the quoted licence fee.
No vendor should be accepted solely because its output looks sophisticated. References, independent penetration testing, financial stability, service history, and data portability deserve review. The contract should require notification of a material model change that could alter prior outputs. It should also make the customer responsible for legal instructions while making the supplier responsible for agreed security, availability, and processing duties.
Common Mistakes and Weak Controls
A common mistake is confusing access control with deletion. Restricting access does not prevent an authorized user from copying or exporting data, while deleting an account does not necessarily remove backups or vendor-side artifacts. Effective controls combine least privilege, download restrictions where proportionate, watermarking, export approval, retention schedules, and verified deletion. They also address screenshots, prompt copies, support tickets, and test environments, which are often overlooked.
Another error is treating an accuracy score as a guarantee. Benchmarks may use easy questions, familiar English, or materials unlike Indian court records. They may not test multilingual documents, scanned records, handwritten notes, conflicting authorities, or unusual metadata. A claim of “90% accuracy” without a defined task, dataset, and error cost cannot answer whether a legal team can rely on the system. Testing should report false positives, false negatives, abstentions, and failures by document type.
Teams also make the mistake of deploying before defining ownership. If no one is accountable for updates, prompt changes, and incidents, users may route around the approved service. Conversely, a central approval process that blocks every minor use can drive work into ungoverned consumer tools. A low-risk lane and a high-risk lane allow firms to apply proportionate review. Users should receive examples of acceptable prompts and examples of prohibited submissions.
Finally, organizations frequently treat a “no training” clause as the entire privacy analysis. A vendor can process data without training it and still retain logs, grant support access, use subcontractors, or transfer information across borders. Indian data-protection obligations, confidentiality, legal privilege, records requirements, and court directions must each be considered. AI governance is therefore not a substitute for ordinary information governance; it adds model-specific risks to established duties.
When to Act and How to Respond to Failure
A team should act before a tool enters production. The trigger is not simply an announcement that generative AI exists; it is the proposed upload of non-public data, connection to a document repository, ranking of evidence, creation of legal work product, or use in a decision affecting a person. By that point, sensitive information may already be exposed. A short pause for vendor, data, and legal review is less costly than attempting to reconstruct what was submitted after an incident.
A staged rollout reduces risk. Begin with a fictional or redacted test containing at least 50 to 100 representative items, then expand to 500 or more documents when the matter is complex. Establish thresholds before testing: for example, zero tolerance for known confidentiality breaches, mandatory correction of every fabricated authority, and a matter-specific recall target approved by counsel. Thresholds should reflect consequences, not industry averages. A system that fails to retrieve a small percentage of decisive documents may be unacceptable in a case where those documents change liability.
If an incident occurs, disconnect the system without destroying evidence, preserve logs and configuration, identify affected data and recipients, and stop further processing. The response team should determine whether notification is required under contract, data-protection law, professional rules, client instructions, or a court order. Counsel should decide on privilege, regulatory, client, and business notifications. Root-cause analysis should examine permissions, prompts, source data, model changes, human review, and vendor events. A corrected prompt is not enough if the underlying control failed.
The board or firm leadership should receive meaningful reporting. Useful figures include the number of users, approved tools, documents processed, incidents, blocked uploads, citation errors, privilege errors, retention exceptions, and hours of review saved. Savings should be net of procurement, testing, training, and review time. If the numbers cannot be measured, claims of efficiency should be treated cautiously.
Overall Control Standard
The defensible standard for Indian legal AI in 2026 is controlled, traceable, and human-verified use. For eDiscovery, that means preserving evidence integrity and reproducibility. For legal research, it means verifying every material authority against an authoritative source. For drafting, it means qualified review of substance, form, deadlines, and confidentiality. Across all uses, the organization must know what data was processed, which version and settings were used, who approved the result, and how the system will be stopped or corrected.
This approach is not a claim that AI is unreliable in every setting. It recognizes that a statistically useful output can still be legally wrong, procedurally unusable, or commercially damaging. Conversely, restrictive controls should not freeze adoption. Clear rules, approved environments, and staged testing can permit useful automation while limiting exposure. Indian legal teams should revisit the framework whenever legislation, court practice, model capability, vendor terms, or the sensitivity of the data changes.
The best immediate decision is therefore simple: before using an AI platform for actual client work, identify the data, classify the use, approve the vendor, define the human reviewer, set measurable failure thresholds, and preserve an audit trail. If those answers cannot be supplied, the system is not ready for production.