Praxis AI Operating Partner

AI is making production cheap.

Making got cheap. Checking did not.

The cost of making code, analysis, research, campaigns, customer responses, documents, decisions and operational actions is falling quickly. The cost of checking all of it is not falling with it, and that gap is becoming one of the defining operating problems of the AI era.

What broke

Making and checking used to move at the same speed.

For decades most organisations used people to do both. A person created the work. Another person reviewed it. A manager approved it. A specialist checked the exception. The pace of production and the pace of verification were roughly matched, because both were constrained by the same thing: human effort.

AI breaks that balance. A machine can generate ten versions before a person has reviewed one. It can produce hundreds of code changes, thousands of responses or millions of decisions at a speed no human review team can realistically match.

That creates a new bottleneck. The constraint is no longer production. It is verification.

The answer is not more checking.

The instinct is to add human approval, and it recreates the old bottleneck at a higher cost. Teams create more than they can inspect. Approval queues grow. Managers start rubber-stamping, because nobody can read at the speed the work arrives. People who should be doing the work only humans can do spend their days checking machine output instead.

That is not transformation. It is moving the bottleneck, and paying more for it.

Essay 010 listed approvals as one of the operating primitives built on the assumption that a person does the work. This is what happens to that one primitive when production scales by an order of magnitude.

The distinction that matters

Checking is getting cheaper too. Judgement is not.

Put that more precisely. Machines check as well as make. They can run the tests, validate against the rules, gather the evidence and flag what looks wrong, and none of that needs a person. What stays expensive is the part that does: deciding whether something ambiguous is acceptable, whether a step that cannot be undone should happen, and who carries it if it goes wrong.

So the design job is separating the checks a machine can run from the judgement only a person can make, then spending that judgement only where it changes the outcome.

The verification model

Every AI-enabled workflow has to decide what reaches a person.

Better verification architecture starts with four operating questions for each workflow. What can the machine verify itself? What evidence must it produce? What can be sampled? What threshold triggers review? Answer them and the work arranges itself into stages, each one settling what it can before anything moves on.

ProducedEverything the system makes.

Verified by the machineTests, rules and validation it runs on its own output.

Arrives with evidenceCarries what a reviewer would need, or does not proceed.

SampledLow-risk work reviewed by proportion, not item by item.

Crosses a thresholdLow confidence, high value, or hard to undo.

Reaches a personWhere judgement changes the outcome.

Widths are illustrative, not measured.

Four more questions decide where the thresholds sit. What actions are reversible? What happens when confidence is low? Who has the authority to stop the process? And what happens when nobody does?

Proportion

Not all work deserves the same gate.

Low risk

A marketing draft

Automated quality checks and occasional sampling.

Rule-bound

A pricing change

Validation against the pricing rules before it goes live.

Explicit authority

A payment

Released only by someone entitled to release it.

Human judgement

A safety-critical decision

A person decides, with a complete evidence trail behind them.

That will sound familiar to anyone who read essay 003, which set how much authority an agent holds by consequence, reversibility, evidence and value at risk. The structure is the same. The reason is different. 003 tiers authority because some harms are too serious to delegate. This tiers verification because otherwise the people run out. The failure it prevents is not an incident. It is a queue.

The human gate should become smaller as AI gets stronger, not larger.

That changes how a company should think about scale. If AI makes production ten times faster, the organisation should not plan for ten times more reviewers. It should redesign the system so that only the right fraction, perhaps one per cent, reaches a person at all.

Most businesses are still asking how much faster AI can make them. That question has a ceiling, set by how fast their people can check. How much more they can safely trust the system to do without them has no such ceiling.

Where Praxis stands

Do not start by hiring reviewers or buying a review tool. Start with one workflow where AI already produces more than people can inspect, and draw its funnel: what the machine verifies itself, what arrives with evidence, what is sampled, what crosses a threshold and what reaches a person.

Then count what reaches a person today. In most organisations it is nearly everything, which is exactly why the queue is growing and why the pilot felt faster than the programme.

The work is unglamorous: deciding, workflow by workflow, what evidence counts, what can be sampled and where each threshold sits. It is also the difference between AI that multiplies output and AI that multiplies the backlog.

The point of AI is not to move more work into a human approval queue. It is to move humans out of routine production and place them precisely where judgement changes the outcome.

Begin

What still reaches a person in your busiest workflow?

One conversation, no pitch deck. Bring a workflow where AI already produces more than your people can check, and we will draw its funnel with you.