Every number here names the system it was measured on.
MetaMinds publishes one case study in full technical detail, a claims register that gates every figure on the site, and a security page that states what is not held. This page is the measured record, followed by the right place to start depending on the scale you operate at.
What moved, on which system
Six reductions delivered on named systems at CourtNetra, SaveLIFE Foundation, INNEFU Labs and Arlo Technologies. Read them as evidence of method rather than as a forecast for a workload nobody has profiled yet.
Indexed to 100 at baseline. The underlying record holds a percentage change for each of these rather than two published endpoints, so indexing is the accurate rendering. The upper bar in each pair is the baseline and the lower bar is the result.
Sourcedocs/05-CLAIMS-REGISTER.md, rows sourced to CourtNetra cost telemetry, SaveLIFE Foundation, INNEFU Labs and Arlo Technologies.
Shown as real percentages rather than indexed, because the record holds the before and after values directly.
Sourcedocs/05-CLAIMS-REGISTER.md, Arlo Technologies, across more than 150 deployments.
Levels rather than changes, so there is no baseline being compared away. The 2% figure is a ceiling and not an achievement: it is the grounding failure threshold above which the CourtNetra build fails in CI.
Dashcam detection precision, 200 km of national highway
Cross encoder ranker precision, 5,000 institutions
Deployment success across 150 releases
SIEM triage inside SLA after agentic routing
Grounding failure ceiling that fails the build
Sourcedocs/05-CLAIMS-REGISTER.md, sourced to CourtNetra CI configuration, SaveLIFE Foundation, INNEFU Labs, Arlo Technologies and Microsoft funded incubator work.
The corpora and estates these systems run against
Retrieval quality claims mean very little without the size of the haystack. These are the real counts from the systems above.
Drawn on a logarithmic scale. Eighteen million and fifty differ by six orders of magnitude, and on a linear axis every bar except the first would render as an invisible sliver. Compare the labels rather than the bar lengths.
Sourcedocs/05-CLAIMS-REGISTER.md, sourced to the CourtNetra production database and source, SaveLIFE Foundation, INNEFU Labs and Microsoft funded incubator work.
The Supreme Court share is under 0.3% of the corpus by volume and carries a large share of the precedential weight, which is exactly why retrieval over this corpus cannot be tuned on volume alone.
SourceCourtNetra production database. All 25 High Courts, 1950 to present.
What the published research says the state of the art actually is
MetaMinds cannot build a competitive benchmark, and a self scored vendor benchmark would be worthless anyway, because every vendor that has published one has won it. What exists instead is stronger: a preregistered, expert scored study of the three leading commercial legal research tools, published by researchers with no commercial stake.
These are not our numbers and not our query set. They are shown because they establish what a legal buyer should expect from a well funded commercial tool, which is the context every grounding claim in this industry needs and almost none of them provide.
0%50%
SourceMagesh, Surani, Dahl, Suzgun, Manning and Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Stanford RegLab and HAI, arXiv:2405.20362, published in the Journal of Empirical Legal Studies, 2025. 202 preregistered, expert scored queries.
Read this before comparing the figure above to ours
CourtNetra holds grounding failure under 2% in a weekly CI gate that fails the build above that threshold. That number is deliberately not drawn on the axis above, because the corpora, the query sets and the scoring rubrics are all different. Placing them on one axis would imply a like for like comparison that has not been run, and drawing a favorable picture out of incompatible measurements is the specific behavior the Stanford study was written to expose.
The study exists because vendors marketed retrieval augmentation as eliminating hallucination and guaranteed citations that were free of it. Its title is the question that phrasing invited. You will not find that claim anywhere on this site, and a firm selling AI assurance that made it would have disproven its own pitch.
Four entry points, one discipline
The work does not change with the size of the buyer. What changes is the size of the first commitment and how procurement handles it. Find the row that describes you.
Personal use
You are accountable for an AI feature and want the evaluation rigour without a consulting budget
- Price
- $99 to $890
- Timeframe
- Priced and published. Checkout is not live yet
- Where
- Self serve by design: no scoping call, once the store opens
What you get
The evaluation suites, red team prompt sets, retrieval test harnesses and Article 50 disclosure templates, packaged from work already shipped on production systems.
How it runs
Each product has its own page carrying the scope, the licence, the delivery promise and the refund position, so you can evaluate it before the store opens. Every price is final and will not move. Until then, the equivalent work is available now as a fixed-scope Tier 2 engagement.
- Ten products, every price published and final
- Designed to run on your infrastructure, with nothing sent to us
- Built from the same harnesses used on the systems measured above
Public and civic
You operate at national coverage and the failure mode is public consequence rather than churn
- Price
- From $2,500, typically a scoped build after
- Timeframe
- Assessment in two weeks, build scoped from there
- Where
- Remote first, on site where your own handling rules require it
What you get
Retrieval systems over national corpora, agentic triage for security operations, and computer vision review pipelines. This is the tier with the most directly comparable prior work.
How it runs
An assessment first, priced fixed, that establishes what the system does today against a written standard. Nothing larger is quoted until that exists, because a national deployment scoped from assumptions is how public programs fail slowly and expensively.
- Prior work at SaveLIFE Foundation across 200 km of national highway
- Penetration testing across 50 government web assets at INNEFU Labs
- Offensive security work delivered under public sector engagement rules
Industrial use
You are shipping AI into a regulated workflow and need the assurance work to hold up to an auditor
- Price
- $2,500 to $8,900, fixed
- Timeframe
- Two to six weeks, start week reserved at deposit
- Where
- Remote, in your repositories, with your engineers
What you get
Vendor assessments, hallucination gates wired into CI, Article 50 readiness reviews, and RAG systems built to a written evaluation standard rather than to a demo.
How it runs
Fixed scope, fixed price, a named start week reserved when the deposit clears, and the balance invoiced only after scoping confirms the work fits. If it does not fit, the deposit is refunded in full rather than absorbed into a larger contract.
- Every price published before you talk to anyone
- Deposit reserves a real start week, not a place in a queue
- Full refund if the scope turns out to be materially different
Industrial use at scale
Procurement, security review and a governance committee all have to sign, and the system carries real regulatory exposure
- Price
- From $18,000, scoped
- Timeframe
- Six weeks to six months, staged with gates
- Where
- Your environment, your data residency, your review cadence
What you get
Retrieval and agent platforms built against Annex III obligations, governance instrumentation that produces the evidence an assessment needs, and the technical documentation that sits behind it.
How it runs
No cart and no published price, because a shopping cart containing a six figure engagement signals a firm that has not met procurement. Staged delivery with gates you approve, documentation written as the work happens rather than assembled afterwards.
- Answers to a security questionnaire, including what is not held
- Data residency and transfer position stated in writing
- Deferred Annex III dates treated as a planning window, not a reprieve
EU AI Act dates, verified against the Official Journal
This timeline has already been amended once, by the Digital Omnibus package in Regulation (EU) 2026/1744. A wrong deadline is trivially rebutted by a buyer's own counsel, so these are re verified rather than remembered.
The deferral of Annex III to December 2027 is a planning window rather than a reprieve. Conformity assessment work for a system already in development has to start well inside it.
2 February 2025
Prohibited practices apply
Article 5 bans took effect. Social scoring, untargeted facial image scraping and emotion inference in workplaces and schools are already unlawful, with no transition period remaining.
2 August 2025
General purpose AI obligations begin
Transparency, copyright policy and training data summaries for new general purpose models. Models placed on the market before this date have until 2 August 2027.
2 August 2026 · in force
Article 50 transparency in force
Disclosure duties for systems that interact with people, generate synthetic media, or infer emotion. This is the obligation that reaches ordinary product teams rather than only high risk deployers.
2 December 2026
Article 50(2) grace period ends
The extension granted by the Digital Omnibus for machine readable marking of synthetic content closes. Systems that generate synthetic content and were already on the market must carry compliant marking by this date. The duty is the provider's, and systems that only classify, rank or predict are not reached by it.
2 December 2027
Annex III high risk obligations apply
Deferred from the original timetable. Covers employment, education, essential services, credit and insurance scoring, and law enforcement use cases.
2 August 2028
Annex I high risk obligations apply
AI embedded in regulated products already covered by EU product safety legislation, aligned to the existing sectoral conformity assessment cycle.
SourceRegulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, Official Journal 24 July 2026. Tracked in docs/03-COMPLIANCE-LANDSCAPE.md.
Analyst data, not a MetaMinds measurement. It is shown apart from the figures above so the two bodies of evidence are never read as one. Drawn on a logarithmic scale.
SourceGartner AI governance platform market sizing, summarized in research/2026-09-01-market-demand.md.
What a careful buyer asks about these numbers
- Are the performance figures on this page guarantees of similar results?
- No. Every figure belongs to a specific past system and is stated in the past tense with that system named. A reduction achieved on one workload is evidence of method, not a prediction about a workload nobody has profiled yet. MetaMinds does not quote an expected improvement before an assessment.
- Why are some bars indexed to 100 rather than showing real values?
- Because the underlying record holds a percentage change rather than two raw endpoints. Indexing to 100 at baseline is the accurate way to draw that. Inventing plausible absolute numbers to make an axis look concrete would be fabrication, so the figure note states the indexing wherever it applies.
- Is MetaMinds SOC 2 or ISO 42001 certified?
- No. Neither attestation is held today, and the security page states this plainly rather than implying otherwise. For buyers where an attestation is a gating condition, that is a reason not to proceed, and saying so early is more useful than discovering it during a security review.
- When does EU AI Act Article 50 apply?
- Article 50 transparency obligations came into force on 2 August 2026. The Article 50(2) grace period for machine readable marking of synthetic content closes on 2 December 2026. Annex III high risk obligations were deferred to 2 December 2027 and Annex I to 2 August 2028 by Regulation (EU) 2026/1744.
- How many case studies does MetaMinds publish?
- One, in full technical detail: CourtNetra, a retrieval system over 18,863,754 Indian judgments. Two further products are live without published metrics. There are no named clients on this site because no client has given written permission to be named.
- How does a sub 2% grounding gate compare to Lexis or Westlaw?
- It cannot be compared directly, and the impact page says so rather than implying otherwise. Stanford RegLab measured Lexis+ AI at 17% and Westlaw AI Assisted Research at 33% across 202 preregistered queries. CourtNetra holds grounding failure under 2% in weekly CI. The corpora, query sets and scoring rubrics differ, so the numbers are not interchangeable.
- What is the smallest way to start?
- A digital product between $99 and $890, delivered on purchase with no scoping call. The next rung is a fixed scope engagement from $2,500 with a named start week reserved at deposit and a full refund if the scope turns out not to fit.
Tell us what breaks if the AI gets it wrong.
That single answer tells us more than a requirements document. Thirty minutes, straight to the engineer who would build it, no sales deck in between.
Typical reply within one business day · Products from $99, engagements from $2,500