Research AI

Where AI is working in enterprise operations, and where it is not

The AI in production inside operations teams is mostly reading things. The AI that was asked to decide has mostly been switched off. Where the line sits today, the governance patterns that hold, and how to pick a first workflow.

August 27, 2026 · 5 min read · Principal Consultant · How work is changing

Where things stand

Two years of pilots have produced a fairly clear picture inside enterprise operations, and it is not the picture in the conference keynotes. The AI that is in production, doing real work every day, is mostly reading things. Documents, forms, emails, tickets, images of receipts. It extracts, classifies, routes, flags and summarizes. A person still decides.

The AI that has been stopped, quietly, is the AI that was asked to decide. Autonomous approvals. Agents that act on a system of record without a review step. Chat interfaces over regulated processes with no trail of what was said or why. Some of these were stopped by risk functions, some by the first incident, and some by the discovery that nobody could explain to an auditor what had happened.

We have built both kinds and we have had both experiences. What follows is where the line sits today, in our work, and how to choose a workflow that stays on the right side of it.

Where it is working

Document extraction

Reading a license, an invoice, a bank letter, a contract clause, and returning the fields that matter with a confidence score per field. This is the workhorse. It works because the output is checkable: the field is either right or a reviewer corrects it, and the correction improves the threshold. It fails on poor images and unfamiliar formats, and the fix for both is mostly product design at the upload step.

Classification and routing

What kind of document is this. What kind of request is this ticket. Which team should see this email. Which vendor category does this invoice belong to. High volume, low individual stakes, and a wrong answer is a delay rather than a loss. Classification models earn their place quickly, and the exception path (route anything below confidence to a person) is simple to build.

Exception detection

Not deciding what to do, but noticing what is unusual. A duplicate invoice. A vendor whose documents pass but whose address is in a jurisdiction that needs a second look. A timesheet that does not match the allocation. A reported volume that departs from the pattern. The model surfaces it with a reason. A person handles it. This is where much of the operational value sits, and it is the least discussed.

Summarization for reviewers

A compliance reviewer facing forty pages produces a better decision faster with a two-paragraph summary and links back to the source. The key word is reviewer. The summary supports a human judgment, and every claim in it points to the passage it came from. Summaries that cannot be traced to source are not used twice.

Where it is not

Autonomous approval in any regulated or financially consequential process. Not because the model is necessarily worse than a tired person at four o’clock, but because the organization cannot yet defend the decision. Defensibility requires a policy version, a decision log, an explanation and a named human who could have intervened. Until those exist, the model recommends and a person signs.

Anything without an audit trail. A chat interface that answers questions about policy is useful. The same interface acting on a system of record, with the conversation as the only record of intent, is a finding.

Free-form generation into customer-facing or regulator-facing channels without review. Drafts, yes. Sends, no.

Continuous self-improvement in a governed workflow. A model that retrains itself on its own corrections sounds attractive and makes last month’s decisions unexplainable. Threshold and model changes should be deliberate, logged and reviewed.

Governance patterns that hold

Per-field confidence, not per-document. A document should not pass or fail as a whole because one field is unreadable.

Thresholds owned by the business function, not the engineering team. Compliance sets and adjusts them, with a log of every change.

An exception queue that is prioritized, not first-in first-out, and instrumented so that queue depth by type and region tells you about a new failure mode before anyone reports it.

Human sign-off on anything consequential, with the screen showing what the system found and the click logged with person, time and rules version.

A decision log that can reconstruct any outcome a year later: inputs, model version, confidence, threshold, reviewer, action.

Evaluation sets that run before every change, with failures examined rather than just scored.

How to choose a first workflow

Pick a workflow that is document-heavy, high volume, and currently slow because it waits for a person rather than because the person’s work is hard. Verification, intake, invoice processing and request triage all qualify. Avoid anything where the interesting part is the judgment.

Time it before you touch it. The instrumented baseline is the most useful artifact in the program, and it usually shows the delay is in queueing, not in checking. That finding shapes the whole design.

Build the exception path first and the model second. The queue, the reviewer screen, the flag reasons in plain language, the logging. A good exception path with a mediocre model beats the reverse, and it is still there when the model improves.

Keep humans on everything for the first weeks even where the model scores well, so the correction data exists to lower thresholds with evidence. It is slower to show results and easier to defend.

Platforms of this type can cut onboarding cycle time by up to 45% in partner networks and vendor processing time by up to 42%. In every case we have run, more of that came from the queue and the exception routing than from the model.

What to watch

Regulatory guidance is arriving in most sectors and it will ask for the trail, the log and the human. Build them now. Watch the cost of inference at production volume, which can turn a cheap pilot into an expensive operation. And treat any proposal for an autonomous agent in a regulated workflow as a proposal to remove the human from the audit trail, because that is what it is. Sometimes that is acceptable. Usually it is not yet.

Read more at /capabilities/ai-digital-transformation/, or see how it worked in practice: How we cut KYC verification from days to hours.

Newsletter

Get the next piece by email.

Engineering notes, playbooks and research. No promotions, unsubscribe in one click.

Next step

Working on something like this?

We are happy to compare notes. Tell us where you are stuck.

Your first seven days are on us. Plan, strategy and solution architecture, before any commitment.

Not ready to talk? Take the 8 minute readiness assessment