Skip to content
Operate and improve

An assessment is true on the day it is signed. Then things move.

Nothing in an AI system announces that its evidence has gone stale. The model version changes, the corpus grows, the question distribution shifts and the regulation gets amended, and every one of those happens without a single test going red. This is the work that keeps a system's documentation describing the system that actually exists.

2
EU AI Act deadlines moved in a single July 2026 amendment
768
checks this site runs against itself on every change
2%
the CI hallucination ceiling on CourtNetra, enforced weekly
$690
a month, where the retainer starts
What actually expires

Six ways a documented system becomes undocumented, with nothing failing

These are not hypotheticals. Each one is a mechanism by which a green test suite keeps being green while the thing it certifies has changed underneath it.

The evaluation set stops being representative

A golden set is a snapshot of the questions users asked when it was built. Six months later the distribution has moved, the set is still green, and the green means less than it did. Nothing in the system reports this, because from the inside nothing has changed.

A model version is deprecated under you

The replacement is not the model your grounding rate was measured on. Providers do not notify your evidence file. This is the single most common way a documented AI system becomes undocumented without anyone touching the code.

The corpus grows

Retrieval quality measured over 400,000 documents is a claim about 400,000 documents. Add a year of new material and the ranker faces a different problem. Recall is usually where it shows first, and recall is the one most teams never measure.

The law moves

Regulation (EU) 2026/1744 shifted two high-risk deadlines in July 2026. Guidance published before that date still states the old schedule, and a compliance file assembled against it is confidently wrong in a way that reads as diligence.

A new attack class arrives

Indirect prompt injection through retrieved documents was not in most threat models two years ago. A test suite written before a technique existed does not fail on it; it simply does not look.

The people change

The engineer who knew why a threshold was set to 0.82 leaves. The number survives, the reason does not, and the next person treats it as a constant rather than a decision that can be revisited.

Three levels

Priced by what is actually being kept true

Every price is published, as everywhere else on this site. The floor is the floor: work above it is scoped, and nothing here is quoted without a conversation about what your system does.

Watch

$690 per month

Evidence maintenance.

A dated brief each month covering what changed in the instruments that apply to your systems, what the change raises for systems of your class, and what your evidence file would need to stay current. Model deprecation notices tracked across the providers you use.

Fits when: You have an assessment and want it not to quietly expire.

Maintain

From $2,400 per month

The evaluation keeps running.

Everything in Watch, plus your evaluation suite executed on a fixed cadence against the current model and the current corpus, with the result and the diff delivered rather than a dashboard. Threshold regressions are reported with the failing cases attached.

Fits when: You have a production system whose accuracy is load-bearing.

Operate

From $6,000 per month

The full engagement, three-month minimum.

Everything in Maintain, plus monitoring, model and cost optimization, incident response inside an agreed window, and quarterly re-testing including an adversarial pass. The published floor in our catalog, and the arrangement most systems in regulated use end up needing.

Fits when: The system is in front of customers or regulators and cannot be wrong quietly.

The proof we can actually offer

We run this discipline on ourselves, in public

MetaMinds has no retainer client to point at, and will not invent one. What exists instead is this website, held to the standard we are describing, with the result published.

Every change to this site runs 768 layout checks, a no-JavaScript render of every page, a keyboard path over every route and a reduced-motion pass, and the numbers are published on our security page rather than asserted.

On 4 and 5 September 2026, four pages were added to this site. The gate caught the published audit table going stale every time, within minutes, and failed the build until the numbers matched. It also failed a new page of ours on a tap target 17 pixels tall, 69 times across three links and 24 viewports, against the WCAG minimum of 24.

That is the mechanism being sold here, pointed at your system instead of ours: a check that fails loudly when a published claim stops being true, rather than a report that was accurate in March.

Two honest qualifications, because a page about things going stale should not overstate its own instrument. The gate proves a published figure no longer matches the arithmetic; it does not re-run the measurement for you, and its failure message tells a developer to go and check. And it caught none of the three defects a human found by looking at a screenshot the same week, because an empty grid cell, a missing call to action and a caption wrapping are not overflows. A gate narrows what can go wrong silently. It does not replace someone looking.

How it starts

Usually at the end of something else

A retainer is rarely the first thing anyone buys, and it should not be. It is what makes sense once there is an assessment worth keeping current.

The normal path is an assessment first, so there is a documented baseline, and then Watch or Maintain to hold it. If you already have an assessment from someone else, that is fine and often better: we will read it, tell you what in it has a shelf life, and you can decide whether any of that is worth paying to maintain.

There is no minimum term on Watch and no notice period. Maintain and Operate carry a three-month minimum because an evaluation cadence shorter than a quarter does not produce a trend, and a trend is the deliverable.

Start here

What in your system would go stale first?

If the answer comes quickly, you already know what this is for. Thirty minutes with the engineer who would do the work, and no obligation to buy a cadence you do not need.

Retainers from $690 per month · Assessments from $2,500 · Typical reply within one business day