AI in eDiscovery
Review triage with the reasoning attached, and a defensible record of how you got there.
eDiscovery has used machine learning longer than almost any legal application. Predictive coding is two decades old and courts have accepted it for over ten years. What language models change is not whether AI is used but what the output can tell you.
What changes from predictive coding
Predictive coding classifies: it learns from a seed set and ranks documents by likely responsiveness. It works well, it is court-accepted, and it cannot tell you why a document was ranked where it was.
A language model can classify and explain, citing the passage that drove the call. For privilege review in particular that is a material difference, because a privilege log entry needs a basis and a reviewer needs to check it.
Defensibility is the constraint
Any review methodology has to survive a challenge from the other side. That means documented process, measurable recall, and a sampling protocol that demonstrates the approach worked, not merely that it was used.
A model that produces good results with no measurement is not defensible, and a defensible process with mediocre recall is still better than an indefensible one with excellent recall. Build the validation protocol before the review starts, agree it where you can, and log everything.
- Documented methodology agreed in advance where possible
- Statistically valid sampling of the null set, not just the responsive set
- Measured recall and precision, reported rather than asserted
- A full audit trail of every model decision and its stated basis
Where predictive coding still wins
On very large, homogeneous populations where the classification is simple and cost per document dominates, predictive coding remains cheaper and entirely adequate. Language models cost more per document and earn it on the harder calls.
The sensible design uses both: predictive coding to cut the population, language models on the residue where explanation and nuance actually matter.
The confidentiality dimension
A production set is the most sensitive material a firm handles, frequently including a client's privileged communications and an opponent's confidential business information under a protective order. Sending it to a third-party model is a disclosure that a protective order may not contemplate.
Read the protective order before choosing an architecture. This is a case where the deployment decision is dictated by a document already in the matter file.
Related
- AI in Legal, the full picture for this sector
- AI for Lawyers
- Legal Research
- Document Processing
- Contract Review
- Local & On-Prem LLM Deployment
- Model Fine-Tuning & Integration
- Custom AI Agents & Automation
Common questions
Will courts accept language-model review?
Courts have accepted technology-assisted review for over a decade, and the reasoning has focused on process and validation rather than the specific algorithm. A documented, measured, sampled methodology is what carries the argument, not the model choice.
Can it do privilege review?
It can triage and propose a basis for each call, which is a substantial improvement on ranking alone. Privilege determinations still get attorney review, because the consequence of a wrong call is waiver.
Does a protective order prevent using AI?
It might, depending on its terms about disclosure to third parties and where data may be processed. This is worth reading before scoping, and it is one of the more common reasons a matter runs on-premise.
How much can this cut review cost?
The population reduction is where the money is, and it varies enormously with the matter. We would rather size it against a sample of your actual corpus than quote a percentage that turns out to be wrong for your case.
Find out what this looks like for your organisation
A 30-minute call. We will tell you plainly whether AI is the right tool for the problem you have, including when it is not.
Book a free consultation