Every agentic system runs on written standards. Voice rules. Audience definitions. The do-not-say list. A description of what a good draft looks like, specific enough that a stranger could apply it. Teams treat these as setup work, finished once. They are not setup. They are inventory, and inventory ages.
Your business moved since you wrote them. A price changed. A product name was retired. Positioning got sharper after a competitor exited. Someone decided in a Tuesday meeting that the company no longer uses a particular word, and that decision lives in three people's memory and nowhere else. The system keeps producing correct work against an outdated specification, which looks exactly like a system producing wrong work.
Drift starts upstream of the model.
When quality slides, the instinct is to blame the model or reach for a newer one. That is almost always the wrong repair. The output is a faithful execution of the instructions it was given. If the instructions describe a version of the company that no longer exists, a better model produces a more polished version of the same error.
This is the part most teams underestimate when they design their operating layer. The model is interchangeable by design — that is the point of model-agnostic routing. The brief is not. It carries the accumulated judgment of everyone who has decided what this brand does and does not say, which makes it the most valuable and most perishable asset in the system.
Rot has a signature. Your gate already records it.
You do not need a survey to detect this. The human approval gate produces the dataset for free. Three signals matter.
First, rejection clustering: the same reason appearing across unrelated outputs, week after week. That is not an execution problem, it is a missing rule. Second, the identical edit: a reviewer rewriting the same paragraph, the same claim, the same closing line every single time. Whatever they keep typing belongs in the brief. Third, and hardest to catch, output that is correct against the standard and wrong against reality — approved reluctantly, with a shrug. That shrug is the most expensive signal in the operation, because nobody logs a shrug.
Audit-grade logs turn the first two into something you can query. Across more than 44,000 runs in 63 days, patterns that would be invisible in a single review become obvious in aggregate. The gate is not only a safety control. It is the instrument that tells you where your written standard has gone stale.
Schedule the maintenance or pay for the rebuild.
The cadence that works is unglamorous. Monthly, pull the rejections, cluster the reasons, and patch the briefs that caused them. Quarterly, read the whole set against current positioning and cut what no longer applies — deletion matters more than addition, because bloated briefs get skimmed and skimmed briefs get ignored. Version every change and note who made it, so the standard has the same audit trail as the work.
Do this and the system compounds: each cycle encodes judgment that used to live in one person's head, which is what actually delivers three to five times the throughput in ninety days. Skip it, and you spend the fourth quarter rebuilding what you already built, which is how a working deployment quietly turns into a failed one. The sequence is laid out on the journey page, and the agents page shows which operators own which briefs.
If your output got worse and you suspect the instructions rather than the technology, book the 30-minute strategy blueprint call.