Praxis AI Partners

If it can change the decision, it is not a technical upgrade.

Every upgrade is an operating change.

A business approves an agent. Three months later the model behind it changes. The agent keeps the same name, the same owner, the same permissions and the same interface. Nobody rebuilt anything. Nobody filed a change request. And it may now reason differently.

The problem hiding in plain sight

Same name. Same owner. Different capability.

Put the two records side by side and almost every line matches. That is precisely why the change goes unnoticed: an inventory, an access review and an architecture diagram would all report that nothing happened.

At approval March

AgentInvoice exceptions

OwnerFinance operations

PermissionsRead, flag, route

InterfaceUnchanged

ModelFamily default

Exceptions escalated1 in 8

In production June

AgentInvoice exceptions

OwnerFinance operations

PermissionsRead, flag, route

InterfaceUnchanged

ModelSilently upgraded

Exceptions escalated1 in 30

Two lines moved. The second one is the one that matters: the agent is now resolving cases it used to hand to a person. Nobody widened its authority. Its judgement simply got more confident, and confidence is not the same thing as permission.

Technically it is the same application. Operationally it may no longer be the same capability.

A first for this series

Until now these essays have argued ahead of the platforms. This week one of them wrote the argument down. Microsoft's Copilot Studio guidance now states plainly that model choice is not a one-time design decision, that an agent can behave differently when its model changes even within the same family, and that such a change can alter instruction interpretation, tool selection, formatting, latency and consumption.

Their recommendation is the part worth noticing. Treat it as a migration: evaluate, approve, deploy, monitor, and fold the failures into regression testing. That is not a deployment checklist. That is an operating discipline, and it is the same one we set out in essay 005.

A change that alters the decision is not a version bump. It is a change to how the business operates.

The one change to make this quarter

Stop letting the platform choose.

Most agents are configured to follow whatever model their platform currently defaults to. For a drafting assistant that is sensible: you get improvements for free. For an agent that touches money, customers, cases or compliance it means the most consequential dependency in the system updates on a schedule you do not set and are not told about.

Pin it. Name a specific model for every consequential agent, so the next change has to pass through a decision instead of arriving as a surprise. It costs almost nothing and it converts a silent event into a governed one.

Before a change

Know what good looked like

A baseline evaluation, a named business owner alongside the technical one, and thresholds for quality, cost and latency. Without a baseline there is nothing to compare against and no way to prove the change was safe.

After a change

Watch the decisions, not the uptime

Regression cases from real failures, monitoring of escalation and override rates, and a rollback path that someone has actually tested. An agent can be perfectly available and quietly deciding differently.

Why this matters now and not last year

Enterprise AI has left the technology department. On OpenAI's own enterprise figures, weekly active use of agentic tooling since February grew 108 times in legal, 41 times in sales, 41 times in recruiting and 26 times in marketing, against 5 times in engineering. Treat that as a depth-of-adoption signal rather than proof of value, which is how OpenAI itself frames it.

The direction is what counts. AI is moving into the functions where work is defined by procedure, authority, evidence and exception handling. A silent model change used to land on code, where a test suite would catch it. It now lands on a contract review, a candidate shortlist or a payment.

Where Praxis stands

Release management asks whether the system still works. An operating model asks whether it still decides the way the business agreed it should. Those are different questions, and only the second one catches a model change that quietly moved the line between what the agent settles and what a person sees.

So model change control belongs in the agent capability lifecycle, next to authority and evidence: a known dependency, named owners, a baseline, regression cases, explicit approval to migrate, a rollback that works, and monitoring afterwards that watches decisions rather than availability.

If changing the technology can change what the business decides, it does not belong in release management. It belongs in the operating model.

Begin

When did your model last change?

One conversation, no pitch deck. If the answer is not on hand, that is the finding, and it is the place to start.