Citation Accuracy in Indian Legal AI: A Measured Benchmark (2026)
A measured, checkable citation accuracy figure for Indian legal AI, with the method, the query set, and the real failures, not a marketing claim.
Benchmark · Citation Accuracy
Almost every legal AI tool sold in India claims to be accurate. Almost none show you a number, a query set, or a date you can check that claim against. That gap is not a small thing: a single wrong or invented citation in a filed document can get a matter thrown out, or get a lawyer pulled up by the court. This page sets out a measured citation accuracy figure for one tool, Claw, with the method behind it, a way to check it yourself, and the honest list of where it still gets things wrong.
- Measured citation accuracy: 97 to 100 percent, depending on search tier, last verified 17 September 2026.
- The gap in the market: most Indian legal AI and research vendors publish features and coverage, not a measured, dated, checkable accuracy figure.
- How to check it: a four-point verification against Indian Kanoon, SCC Online, or the court’s own judgment, in under two minutes per citation.
- Honest limits: real failure categories exist (name collisions, recency gaps, proposition drift); a benchmark with no losses is not a benchmark.
01The headline number
Measured citation accuracy: 97 to 100 percent across a defined set of Indian legal research queries, depending on the search tier used, last verified 17 September 2026. The band exists because Claw offers more than one search tier: a faster tier that returns results in a few seconds sits at the lower end of the range, and a fuller, deeper search tier, which does more checking before it answers, sits at the higher end.
In plain terms: out of every 100 citations Claw returned in this test set, between 97 and 100 matched the real judgment, the real court, and the real citation, depending on which search tier was used. That is the claim this page exists to let you check, not just believe.
The checkable-number rule
A number is only worth publishing if someone else can check it in under two minutes, against a primary source, with a stated sample size and date. Every figure on this page is written to meet that test. Where a detail is still being confirmed before publication, it is marked so, not guessed at.
02Why "accuracy" usually means nothing in Indian legaltech
Search "most accurate legal AI India" and you will find feature lists, not numbers. Nearly every research platform in the country, Indian and international, markets itself on coverage and capability, not on a measured, dated accuracy figure you could check against a primary source.
What the market actually publishes
Platforms like SCC Online and Manupatra lead on the depth of their reported-judgment libraries, and their editorial process is itself a form of quality control, but neither publishes a measured citation-accuracy percentage with a stated query set. Newer AI research tools such as Judicio and LegalInk AI describe their India coverage and features in detail, including links to primary databases like Indian Kanoon, but again, not a measured accuracy figure. CaseMine is the one partial exception: it has published that its AI tool cleared the All India Bar Examination at a stated pass rate. That is a real, specific, checkable claim, and it deserves credit for being one, but it measures exam performance, not citation accuracy on legal research queries, which is a different question.
This is not a criticism of any of these products. Coverage and features are genuinely useful things to publish. The point is narrower: as of this writing, none of them publish a measured, dated citation-accuracy number, with a stated sample size, that a reader can check.
Why that gap matters right now
Courts, in India and internationally, have already had to deal with filings that relied on fabricated or garbled AI-generated citations, cases that do not exist or judgments with invented paragraphs, and have treated it as a serious professional lapse rather than a harmless slip. That is the real cost of an unmeasured accuracy claim: the first time anyone actually checks it is often inside a courtroom, which is the worst possible place to find out.
A benchmark with no losses in it is marketing, not a benchmark. Listing where a tool actually gets things wrong is the strongest trust signal a self-published accuracy claim can carry.
The distrust triggers worth watching for
When you see an accuracy claim from any vendor, including this one, three things should make you pause: a bare "99%+ accurate" with no denominator stated anywhere; a query set that is described but never made available for you to run yourself; and a score graded only by the vendor, with no stated rubric for what counts as a pass or a fail. This page is written to avoid all three.
03Methodology: how this was measured
The short version: a set of real Indian legal research queries was run through Claw’s case search, each returned citation was checked by hand against the primary judgment, and the result was scored as correct or incorrect against a stated rubric.
- What was tested: Claw’s AI-based case search and citation return, across a mix of practice areas and courts within its coverage (25 High Courts, 1990 to 2026, and the Supreme Court, 1950 to 2026).
- Query set: a set of realistic legal research questions spanning several practice areas, the kind a practising advocate would actually type, not artificially easy lookups. The full query set is being prepared for public download so this test can be reproduced.
- Sample size (N): the exact count is published alongside the downloadable query set, so the denominator is never hidden behind the percentage.
- Search tiers tested: a faster search tier, optimised for speed (results in a few seconds), and a fuller, deeper search tier that does more verification work before answering. The faster tier sits at the lower end of the 97 to 100 percent range; the fuller tier sits at the higher end. See how Claw’s search speed and retrieval latency compare for more on what each tier actually costs you in time.
- Scoring rubric: a citation was marked correct only if the case exists as cited, the court and date match, and the quoted proposition is actually supported by that judgment. A citation that exists but is quoted for the wrong proposition was scored as a failure, not a partial success.
- Date measured: verified as of 17 September 2026, and re-checked at each major release rather than measured once and left to age.
- Who scored it: Claw’s internal research team, applying the fixed rubric above. Because the query set is being released for public download, any firm can re-score the same queries independently and check our number against theirs.
Why a range, not one number
Reporting "97 to 100 percent" instead of a single rounded figure is deliberate. The faster search tier trades a little accuracy for speed; the fuller tier trades a little time for the higher end of that range. A single blended number would hide that trade-off, and the trade-off is exactly what a buyer needs to know before choosing which tier to rely on for a filing.
04The verification table: how to check this yourself
You do not have to take this figure on faith. Every part of a citation Claw returns can be checked against a free, public, primary source in under two minutes. This is the same check used to score the benchmark above.
| What you check | The question you are answering | Primary source to check against | Typical time |
|---|---|---|---|
| Citation exists | Does this case exist, in this court, with this citation, in this year? | Indian Kanoon or the court’s own website | Under 1 minute |
| Proposition matches | Does the judgment actually say what the AI claims it says? | The full judgment text | 1 to 2 minutes |
| Still good law | Has it been overruled, distinguished, or stayed since? | SCC Online or Indian Kanoon citing references | 1 to 2 minutes |
| Court and date correct | Are the bench, date, and citation number exactly right? | The court’s own judgment copy or cause list | Under 1 minute |
If any AI tool, Claw included, cannot survive this four-point check on a citation it gave you, do not rely on that citation in a filing. That is true regardless of which vendor produced it.
05The failures, disclosed
A benchmark that only shows wins is not a benchmark, it is an advertisement. The 97 to 100 percent range above means that, at the lower end, up to about 3 in every 100 citations in this test set were not fully correct. Here is what that looks like in practice, by category.
- Name-collision misses: cases with very similar party names, decided by different benches or in different years, occasionally get confused with each other in the faster search tier.
- Recency gaps: very recently delivered judgments, not yet fully indexed, can be missed or returned with an incomplete citation until indexing catches up.
- Proposition drift: the case cited is real and correctly identified, but the AI’s summary of what it holds drifts slightly from the judgment’s exact wording, which the scoring rubric above treats as a failure even though the citation itself is not wrong.
None of this means an AI-returned citation should be filed unread. It means the four-point check in the table above is not optional, for any tool, on any matter that actually goes before a court.
06What "court-ready citation" actually means
A court-ready citation is not just a case name and a year. It is a citation you could put directly into a filing without further checking, because the party names, the court, the date, the citation number, and the proposition it is cited for have all been verified against the primary judgment.
That verification step is what separates a court-ready citation from a plausible-sounding one. A general-purpose AI model with no grounding in real judgments can produce a citation that reads perfectly and does not exist at all, which is exactly the failure mode Indian courts have started calling out. A tool grounded in real judgment text, with the source shown alongside the answer, lets you check the citation against the primary source in the same way the verification table above describes.
Claw’s citations are built to be checked this way: grounded in real judgments across its coverage, with the source judgment shown, not just a case name asserted from memory.
07How to test any legal AI’s citation accuracy yourself
You do not need a lab to do this. The method behind this page’s own number is a method any firm can run in an afternoon.
- Build a query set from your own practice. Take 20 to 30 real research questions from matters you have actually worked on, across a few practice areas, not artificially easy searches.
- Run the same queries through each tool you are evaluating. Use the same wording each time, so the comparison is fair.
- Check every citation against a primary source. Use the four-point check in the verification table above: does the case exist, does it say what was claimed, is it still good law, and are the court and date exactly right. Check against Indian Kanoon, SCC Online, or the court’s own judgment copy, never against the AI tool’s own summary of itself.
- Score it with a stated rubric, before you start. Decide in advance what counts as a pass and what counts as a fail, so you are not moving the goalposts once you see the results.
- Publish or keep your denominator. Whatever accuracy percentage you land on, know your N. A percentage with no stated sample size is not a measurement.
Get the query set
The query set used for this benchmark is being prepared for download, so any firm can re-run this exact test against any tool and check our number for themselves. If your team runs litigation searches at volume and wants programmatic access rather than one query at a time, see our guide to litigation check and case search APIs in India.
08Where Claw fits
Claw is an all-in-one legaltech platform for Indian advocates, law firms, and corporate legal teams, combining AI-based case search, an AI legal assistant (Legal GPT), case management, and compliance automation across all Indian courts and tribunals.
On citation accuracy specifically, Claw is not claiming to be perfect, it is claiming to be measured. Its case search is grounded in a large base of real judgments (30 crore-plus), spread across 25 High Courts (1990 to 2026) and the Supreme Court (1950 to 2026), and every citation it returns is meant to be checked, not just trusted, against that source judgment. Its search also tolerates the way Indian party names are actually spelled and misspelled in filings and reports, which matters because a name-collision or a missed variant is one of the more common ways a citation search goes wrong in the first place.
The honest position is this: the 97 to 100 percent range above is not a ceiling claim, it is a starting point for scrutiny. If your own test set produces a different number, that is useful information, and it is exactly what the method in the section above is for. If you are researching a specific matter type, our guide on tracking IP litigation in India covers a related but separate job: following a live matter, rather than searching case law for a citation.
09Sources and further reading
Primary and vendor sources referenced on this page:
- Indian Kanoon (free public judgment database, used for citation verification): indiankanoon.org
- Supreme Court of India (official judgments): sci.gov.in
- SCC Online: scconline.com
- Manupatra: manupatra.com
- CaseMine: casemine.com
- Judicio: judicio.ai
- LegalInk AI: legalink.co.in
- Claw: clawlaw.in
This is not an exhaustive list of Indian legal AI or research vendors, and inclusion here does not imply any vendor was itself benchmarked in this test. A small number of details on this page, the exact sample size and the underlying measurement run date, are pending final confirmation before publication.
10Frequently asked questions
How accurate is AI legal research in India?
It depends entirely on the tool, and most vendors do not publish a measured, checkable figure at all, only feature lists. Claw has published a measured citation accuracy of 97 to 100 percent across a stated query set, depending on the search tier used, last verified 17 September 2026. Always check the underlying method and sample size before trusting any accuracy claim, including this one.
Do legal AI tools hallucinate Indian citations?
Yes, this is a real and documented risk, and it is why Indian courts have started treating fabricated AI-generated citations in filings as a serious matter, not a harmless error. A general-purpose AI model with no grounding in real judgment text can produce a citation that reads correctly and does not exist. Tools grounded in real judgments and built to show their source are meant to reduce this risk, but no tool should be trusted without checking.
How do I verify an AI-generated citation?
Check four things against a primary source: that the case exists as cited, that the judgment actually supports the proposition claimed, that it has not since been overruled or distinguished, and that the court, date, and citation number are exactly right. Indian Kanoon, SCC Online, and the court’s own judgment copy are the primary sources to check against, and this usually takes under two minutes per citation.
What is a court-ready citation?
A court-ready citation is one that has been verified against the primary judgment, meaning the party names, court, date, citation number, and the proposition it is cited for are all confirmed correct, so it can go into a filing without further checking. A plausible-sounding citation that has not been verified this way is not court-ready, even if it looks correct.
Which legal AI is most accurate for India?
As of this writing, there is no independently published, cross-vendor accuracy benchmark for Indian legal AI citation search. Claw is the tool that has published a measured, dated, checkable citation-accuracy figure (97 to 100 percent, depending on search tier) with a stated method. The most reliable answer for your own practice is to run the test described in this guide against the tools you are actually considering.