AgentInvoice exceptions
OwnerFinance operations
PermissionsRead, flag, route
InterfaceUnchanged
ModelFamily default
Exceptions escalated1 in 8
If it can change the decision, it is not a technical upgrade.
A business approves an agent. Three months later the model behind it changes. The agent keeps the same name, the same owner, the same permissions and the same interface. Nobody rebuilt anything. Nobody filed a change request. And it may now reason differently.
The problem hiding in plain sight
Put the two records side by side and almost every line matches. That is precisely why the change goes unnoticed: an inventory, an access review and an architecture diagram would all report that nothing happened.
AgentInvoice exceptions
OwnerFinance operations
PermissionsRead, flag, route
InterfaceUnchanged
ModelFamily default
Exceptions escalated1 in 8
AgentInvoice exceptions
OwnerFinance operations
PermissionsRead, flag, route
InterfaceUnchanged
ModelSilently upgraded
Exceptions escalated1 in 30
Two lines moved. The second one is the one that matters: the agent is now resolving cases it used to hand to a person. Nobody widened its authority. Its judgement simply got more confident, and confidence is not the same thing as permission.
Technically it is the same application. Operationally it may no longer be the same capability.
A first for this series
Until now these essays have argued ahead of the platforms. This week one of them wrote the argument down. Microsoft's Copilot Studio guidance now states plainly that model choice is not a one-time design decision, that an agent can behave differently when its model changes even within the same family, and that such a change can alter instruction interpretation, tool selection, formatting, latency and consumption.
Their recommendation is the part worth noticing. Treat it as a migration: evaluate, approve, deploy, monitor, and fold the failures into regression testing. That is not a deployment checklist. That is an operating discipline, and it is the same one we set out in essay 005.
A change that alters the decision is not a version bump. It is a change to how the business operates.
The one change to make this quarter
Most agents are configured to follow whatever model their platform currently defaults to. For a drafting assistant that is sensible: you get improvements for free. For an agent that touches money, customers, cases or compliance it means the most consequential dependency in the system updates on a schedule you do not set and are not told about.
Pin it. Name a specific model for every consequential agent, so the next change has to pass through a decision instead of arriving as a surprise. It costs almost nothing and it converts a silent event into a governed one.
A baseline evaluation, a named business owner alongside the technical one, and thresholds for quality, cost and latency. Without a baseline there is nothing to compare against and no way to prove the change was safe.
Regression cases from real failures, monitoring of escalation and override rates, and a rollback path that someone has actually tested. An agent can be perfectly available and quietly deciding differently.
Why this matters now and not last year
Enterprise AI has left the technology department. On OpenAI's own enterprise figures, weekly active use of agentic tooling since February grew 108 times in legal, 41 times in sales, 41 times in recruiting and 26 times in marketing, against 5 times in engineering. Treat that as a depth-of-adoption signal rather than proof of value, which is how OpenAI itself frames it.
The direction is what counts. AI is moving into the functions where work is defined by procedure, authority, evidence and exception handling. A silent model change used to land on code, where a test suite would catch it. It now lands on a contract review, a candidate shortlist or a payment.
Where Praxis stands
Release management asks whether the system still works. An operating model asks whether it still decides the way the business agreed it should. Those are different questions, and only the second one catches a model change that quietly moved the line between what the agent settles and what a person sees.
So model change control belongs in the agent capability lifecycle, next to authority and evidence: a known dependency, named owners, a baseline, regression cases, explicit approval to migrate, a rollback that works, and monitoring afterwards that watches decisions rather than availability.
If changing the technology can change what the business decides, it does not belong in release management. It belongs in the operating model.
Begin
One conversation, no pitch deck. If the answer is not on hand, that is the finding, and it is the place to start.