Skip to content
Evidence

Every number here names the system it was measured on.

MetaMinds publishes one case study in full technical detail, a claims register that gates every figure on the site, and a security page that states what is not held. This page is the measured record, followed by the right place to start depending on the scale you operate at.

18.8M
judgments indexed and retrievable in CourtNetra
60%
model spend reduction at a 70% cache hit rate
2%
grounding failure ceiling that fails the build
1
case study published in full technical detail
The measured record

What moved, on which system

Six reductions delivered on named systems at CourtNetra, SaveLIFE Foundation, INNEFU Labs and Arlo Technologies. Read them as evidence of method rather than as a forecast for a workload nobody has profiled yet.

Reductions delivered on named production systems

Indexed to 100 at baseline. The underlying record holds a percentage change for each of these rather than two published endpoints, so indexing is the accurate rendering. The upper bar in each pair is the baseline and the lower bar is the result.

Model spend, cache aware cascade on CourtNetrareduced 60%

Before 100After 40

Manual audit effort, dashcam review at SaveLIFE Foundationreduced 85%

Before 100After 15

Analyst lookup time, multi stage RAG at SaveLIFE Foundationreduced 70%

Before 100After 30

Critical CVE exposure, 50 government web assets at INNEFU Labsreduced 65%

Before 100After 35

Inference cost, routing and quantisation at Arlo Technologiesreduced 35%

Before 100After 65

Sustained database load, access pattern redistribution at Arloreduced 30%

Before 100After 70

Sourcedocs/05-CLAIMS-REGISTER.md, rows sourced to CourtNetra cost telemetry, SaveLIFE Foundation, INNEFU Labs and Arlo Technologies.

Deployment reliability, the one figure with both endpoints published

Shown as real percentages rather than indexed, because the record holds the before and after values directly.

Deployment success rate across 150 releases at Arlo Technologiesimproved 4%

Before 95%After 99%

Sourcedocs/05-CLAIMS-REGISTER.md, Arlo Technologies, across more than 150 deployments.

Achieved rates on shipped systems

Levels rather than changes, so there is no baseline being compared away. The 2% figure is a ceiling and not an achievement: it is the grounding failure threshold above which the CourtNetra build fails in CI.

95%MEASURED

Dashcam detection precision, 200 km of national highway

92%MEASURED

Cross encoder ranker precision, 5,000 institutions

99%MEASURED

Deployment success across 150 releases

90%MEASURED

SIEM triage inside SLA after agentic routing

2%MEASURED

Grounding failure ceiling that fails the build

Sourcedocs/05-CLAIMS-REGISTER.md, sourced to CourtNetra CI configuration, SaveLIFE Foundation, INNEFU Labs, Arlo Technologies and Microsoft funded incubator work.

Scale

The corpora and estates these systems run against

Retrieval quality claims mean very little without the size of the haystack. These are the real counts from the systems above.

Documents, endpoints and estates under management

Drawn on a logarithmic scale. Eighteen million and fifty differ by six orders of magnitude, and on a linear axis every bar except the first would render as an invisible sliver. Compare the labels rather than the bar lengths.

Indian judgments indexed and retrievable in CourtNetra18,863,754
Court endpoints reconciled by the auto fetch engine19,660
Institutions ranked by the two stage cross encoder5,000+
Hours of dashcam footage processed500+
Government web assets penetration tested50+

Sourcedocs/05-CLAIMS-REGISTER.md, sourced to the CourtNetra production database and source, SaveLIFE Foundation, INNEFU Labs and Microsoft funded incubator work.

Composition of the CourtNetra judgment corpus

The Supreme Court share is under 0.3% of the corpus by volume and carries a large share of the precedential weight, which is exactly why retrieval over this corpus cannot be tuned on volume alone.

  • High Court judgments, all 25 High Courts18,824,596 judgments 100%
  • Supreme Court judgments, 1950 to present39,158 judgments 0.2%

SourceCourtNetra production database. All 25 High Courts, 1950 to present.

The independent baseline

What the published research says the state of the art actually is

MetaMinds cannot build a competitive benchmark, and a self scored vendor benchmark would be worthless anyway, because every vendor that has published one has won it. What exists instead is stronger: a preregistered, expert scored study of the three leading commercial legal research tools, published by researchers with no commercial stake.

Hallucination rates measured by Stanford RegLab across 202 preregistered queries

These are not our numbers and not our query set. They are shown because they establish what a legal buyer should expect from a well funded commercial tool, which is the context every grounding claim in this industry needs and almost none of them provide.

GPT-4, no retrieval groundingBaseline for an ungrounded general model on the same query set43%
Westlaw AI Assisted Research, Thomson ReutersCommercial retrieval augmented legal research tool33%
Lexis plus AI, LexisNexisCommercial retrieval augmented legal research tool17%

0%50%

SourceMagesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Stanford RegLab and HAI, arXiv:2405.20362, published in the Journal of Empirical Legal Studies, 2025. 202 preregistered, expert scored queries.

Read this before comparing the figure above to ours

CourtNetra holds grounding failure under 2% in a weekly CI gate that fails the build above that threshold. That number is deliberately not drawn on the axis above, because the corpora, the query sets and the scoring rubrics are all different. Placing them on one axis would imply a like for like comparison that has not been run, and drawing a favorable picture out of incompatible measurements is the specific behavior the Stanford study was written to expose.

The study exists because vendors marketed retrieval augmentation as eliminating hallucination and guaranteed citations that were free of it. Its title is the question that phrasing invited. You will not find that claim anywhere on this site, and a firm selling AI assurance that made it would have disproven its own pitch.

Where to start

Four entry points, one discipline

The work does not change with the size of the buyer. What changes is the size of the first commitment and how procurement handles it. Find the row that describes you.

One professional

Personal use

You are accountable for an AI feature and want the evaluation rigour without a consulting budget

Price
$99 to $890
Timeframe
Priced and published. Checkout is not live yet
Where
Self serve by design: no scoping call, once the store opens

What you get

The evaluation suites, red team prompt sets, retrieval test harnesses and Article 50 disclosure templates, packaged from work already shipped on production systems.

How it runs

Each product has its own page carrying the scope, the licence, the delivery promise and the refund position, so you can evaluate it before the store opens. Every price is final and will not move. Until then, the equivalent work is available now as a fixed-scope Tier 2 engagement.

  • Ten products, every price published and final
  • Designed to run on your infrastructure, with nothing sent to us
  • Built from the same harnesses used on the systems measured above
National scale body

Public and civic

You operate at national coverage and the failure mode is public consequence rather than churn

Price
From $2,500, typically a scoped build after
Timeframe
Assessment in two weeks, build scoped from there
Where
Remote first, on site where your own handling rules require it

What you get

Retrieval systems over national corpora, agentic triage for security operations, and computer vision review pipelines. This is the tier with the most directly comparable prior work.

How it runs

An assessment first, priced fixed, that establishes what the system does today against a written standard. Nothing larger is quoted until that exists, because a national deployment scoped from assumptions is how public programs fail slowly and expensively.

  • Prior work at SaveLIFE Foundation across 200 km of national highway
  • Penetration testing across 50 government web assets at INNEFU Labs
  • Offensive security work delivered under public sector engagement rules
Company

Industrial use

You are shipping AI into a regulated workflow and need the assurance work to hold up to an auditor

Price
$2,500 to $8,900, fixed
Timeframe
Two to six weeks, start week reserved at deposit
Where
Remote, in your repositories, with your engineers

What you get

Vendor assessments, hallucination gates wired into CI, Article 50 readiness reviews, and RAG systems built to a written evaluation standard rather than to a demo.

How it runs

Fixed scope, fixed price, a named start week reserved when the deposit clears, and the balance invoiced only after scoping confirms the work fits. If it does not fit, the deposit is refunded in full rather than absorbed into a larger contract.

  • Every price published before you talk to anyone
  • Deposit reserves a real start week, not a place in a queue
  • Full refund if the scope turns out to be materially different
Enterprise

Industrial use at scale

Procurement, security review and a governance committee all have to sign, and the system carries real regulatory exposure

Price
From $18,000, scoped
Timeframe
Six weeks to six months, staged with gates
Where
Your environment, your data residency, your review cadence

What you get

Retrieval and agent platforms built against Annex III obligations, governance instrumentation that produces the evidence an assessment needs, and the technical documentation that sits behind it.

How it runs

No cart and no published price, because a shopping cart containing a six figure engagement signals a firm that has not met procurement. Staged delivery with gates you approve, documentation written as the work happens rather than assembled afterwards.

  • Answers to a security questionnaire, including what is not held
  • Data residency and transfer position stated in writing
  • Deferred Annex III dates treated as a planning window, not a reprieve
The clock

EU AI Act dates, verified against the Official Journal

This timeline has already been amended once, by the Digital Omnibus package in Regulation (EU) 2026/1744. A wrong deadline is trivially rebutted by a buyer's own counsel, so these are re verified rather than remembered.

Obligations in force, and the two that are not yet

The deferral of Annex III to December 2027 is a planning window rather than a reprieve. Conformity assessment work for a system already in development has to start well inside it.

  1. 2 February 2025

    Prohibited practices apply

    Article 5 bans took effect. Social scoring, untargeted facial image scraping and emotion inference in workplaces and schools are already unlawful, with no transition period remaining.

  2. 2 August 2025

    General purpose AI obligations begin

    Transparency, copyright policy and training data summaries for new general purpose models. Models placed on the market before this date have until 2 August 2027.

  3. 2 August 2026 · in force

    Article 50 transparency in force

    Disclosure duties for systems that interact with people, generate synthetic media, or infer emotion. This is the obligation that reaches ordinary product teams rather than only high risk deployers.

  4. 2 December 2026

    Article 50(2) grace period ends

    The extension granted by the Digital Omnibus for machine readable marking of synthetic content closes. Systems that generate synthetic content and were already on the market must carry compliant marking by this date. The duty is the provider's, and systems that only classify, rank or predict are not reached by it.

  5. 2 December 2027

    Annex III high risk obligations apply

    Deferred from the original timetable. Covers employment, education, essential services, credit and insurance scoring, and law enforcement use cases.

  6. 2 August 2028

    Annex I high risk obligations apply

    AI embedded in regulated products already covered by EU product safety legislation, aligned to the existing sectoral conformity assessment cycle.

SourceRegulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, Official Journal 24 July 2026. Tracked in docs/03-COMPLIANCE-LANDSCAPE.md.

Third party market context, kept separate from our own measurements

Analyst data, not a MetaMinds measurement. It is shown apart from the figures above so the two bodies of evidence are never read as one. Drawn on a logarithmic scale.

AI governance platform market, 2026$492M
Projected market, 2030over $1B

SourceGartner AI governance platform market sizing, summarized in research/2026-09-01-market-demand.md.

Questions

What a careful buyer asks about these numbers

Are the performance figures on this page guarantees of similar results?
No. Every figure belongs to a specific past system and is stated in the past tense with that system named. A reduction achieved on one workload is evidence of method, not a prediction about a workload nobody has profiled yet. MetaMinds does not quote an expected improvement before an assessment.
Why are some bars indexed to 100 rather than showing real values?
Because the underlying record holds a percentage change rather than two raw endpoints. Indexing to 100 at baseline is the accurate way to draw that. Inventing plausible absolute numbers to make an axis look concrete would be fabrication, so the figure note states the indexing wherever it applies.
Is MetaMinds SOC 2 or ISO 42001 certified?
No. Neither attestation is held today, and the security page states this plainly rather than implying otherwise. For buyers where an attestation is a gating condition, that is a reason not to proceed, and saying so early is more useful than discovering it during a security review.
When does EU AI Act Article 50 apply?
Article 50 transparency obligations came into force on 2 August 2026. The Article 50(2) grace period for machine readable marking of synthetic content closes on 2 December 2026. Annex III high risk obligations were deferred to 2 December 2027 and Annex I to 2 August 2028 by Regulation (EU) 2026/1744.
How many case studies does MetaMinds publish?
One, in full technical detail: CourtNetra, a retrieval system over 18,863,754 Indian judgments. Two further products are live without published metrics. There are no named clients on this site because no client has given written permission to be named.
How does a sub 2% grounding gate compare to Lexis or Westlaw?
It cannot be compared directly, and the impact page says so rather than implying otherwise. Stanford RegLab measured Lexis+ AI at 17% and Westlaw AI Assisted Research at 33% across 202 preregistered queries. CourtNetra holds grounding failure under 2% in weekly CI. The corpora, query sets and scoring rubrics differ, so the numbers are not interchangeable.
What is the smallest way to start?
A digital product between $99 and $890, delivered on purchase with no scoping call. The next rung is a fixed scope engagement from $2,500 with a named start week reserved at deposit and a full refund if the scope turns out not to fit.
Start here

Tell us what breaks if the AI gets it wrong.

That single answer tells us more than a requirements document. Thirty minutes, straight to the engineer who would build it, no sales deck in between.

Typical reply within one business day · Products from $99, engagements from $2,500