LibraryFoundational AI: Do's and Don'ts8 min read
AI Without the Hype
An AI first pass can shorten the work between a customer request and a useful answer. Measure that gain without hiding the checking and cleanup.

An operations lead can use AI to turn incoming customer requests into a first-pass work queue, replacing the opening hour of reading, copying, and sorting. The useful question is how much of that hour stays saved after someone checks the results. A convincing demo answers the first half. A business has to pay for both.
That distinction gets lost in arguments about whether AI is impressive. The more useful thing to establish is whether it improves a specific piece of work at a quality level your team can actually accept. You don't need a position on the future of intelligence to measure that.
_Revised September 10, 2026. The original publication date is retained._
The useful change is a cheaper first pass
The work between an input and a decision often contains a lot of language handling. Someone reads a request, finds the account, separates the complaint from the background, and writes a short summary for the next person. AI can help with that preparation. Whether it should also make the decision is a separate choice.
In its engineering guidance on effective agents, Anthropic distinguishes predefined workflows from systems in which a model chooses its own steps. It recommends starting with the simplest solution that meets the need and adding complexity when results justify it. More autonomous processes carry tradeoffs in cost, delay, and error handling. Buying more steps doesn't automatically buy a better outcome.
For an operator, the implication is practical. A fixed sequence that reads an email and proposes three fields might be enough. It doesn't need permission to search the whole company, negotiate with the customer, update the account, and invent its own escalation procedure. Start with the piece of work whose improvement you can see.
Consider a hypothetical distributor receiving a hundred order questions a day. The team wants an account name, order reference, request category, and proposed next action. A useful first version shows those fields beside the original message. It also shows when the order reference is absent. Nobody has to accept a guessed order number to keep the queue moving.
This is a design decision, not a prompting trick. The system is allowed to be incomplete in an obvious way. That makes it easier to trust the parts that are present because the reviewer can see where the machine stopped. A polished paragraph that quietly fills every gap gives the reviewer a harder job.
The output format matters for the same reason. A paragraph saying that a customer seems unhappy invites interpretation. A category, supporting sentence from the message, and suggested destination can be checked against a rule. The business doesn't need to reconstruct every step inside the model to decide whether that routing proposal is acceptable.
Measure a finished task, including the cleanup
Before trying the tool, watch the existing process. Measure several ordinary cases and several difficult ones. Record how long the person spends preparing the work, how long they spend checking it, and how often it comes back. The baseline includes the errors in the current process too. Human work deserves the same honest accounting.
Here's a worked example with invented numbers. Suppose preparing a request takes six minutes today. The AI produces a first pass in twenty seconds, and a reviewer spends two minutes checking it. The routine case saves three minutes and forty seconds. That sounds worthwhile, but it isn't the final calculation.
If one request in ten needs another eight minutes of cleanup, spread that additional work across the batch. The average cleanup cost is forty-eight seconds per request. The net saving becomes two minutes and fifty-two seconds, before administration, software charges, or rework discovered later. The smaller number is still useful. It's also a number you can defend.
Then ask where the saved time goes. If the same employee can clear the morning queue before calls start, that is a real operating improvement. If the work merely moves into five scattered interruptions, the team may experience less benefit than the arithmetic suggests. Time freed in usable blocks is different from seconds shaved off a screen.
Separate labor capacity from cash savings. Ten hours released each month doesn't automatically remove ten hours from payroll. It may let the team answer more requests with the same headcount, reduce overtime, or avoid a delayed hire. Name the actual outcome. The case for adoption becomes weaker when every minute is presented as money immediately returned to the bank.
Also measure the consequence of an error, not just the error count. An incorrect label that a reviewer changes in two seconds is different from sending a replacement item to the wrong address. A single blended accuracy percentage can hide that distinction. Count harmless corrections, expensive corrections, and errors that reached someone outside the team separately.
Use examples from the real queue, with appropriate permission and data handling. Include abbreviated messages, forwarded threads, missing details, and requests containing two different issues. A clean sample is useful for getting the workflow running. A representative sample is what tells you whether to keep it.
A pilot needs a stopping rule
I would start this kind of pilot with a fixed batch and a named reviewer. Fifty requests might be a manageable first exercise for a small team. That isn't a statistically magical number or proof that rare failures can't happen. It's enough to reveal whether the proposed output is even helping people do the job.
Write the acceptance conditions before running the batch. For this example, the team might require that every account assignment is checked against a record, missing order references remain empty, and no response reaches a customer without approval. The time target belongs beside those conditions. Faster work that violates the operating rules doesn't pass.
Keep a simple result sheet. Record the original request identifier, proposed category, reviewer correction, elapsed review time, and final disposition. You don't need a separate analytics product to learn that half the requests are being sent to the wrong queue. You need enough detail to see which half and why.
Then divide the failures into things you can fix deliberately. Some come from an unclear business rule. Some come from missing context. Some come from an unsuitable output format. Some are model mistakes despite clear instructions. Those categories point to different changes, so don't bundle everything into a vague complaint about the prompt.
For example, the team may disagree about whether a delayed shipment belongs with customer service or purchasing. The AI cannot resolve an operating policy that the company hasn't settled. Writing a longer instruction only hides the disagreement in more text. Decide the policy, add an example, and rerun the affected cases.
Keep difficult cases as a repeatable check. When someone changes the instructions or switches the model, run those cases again alongside fresh examples. Otherwise every revision gets judged on a different sample, and the team can mistake a friendlier batch for an improvement. A small collection of known failures is often more useful than another impressive demonstration.
Set a stop condition for the pilot too. If review consistently takes longer than the original work, stop expanding it. If the output cannot make uncertainty visible, redesign it before connecting actions. A pilot that tells you a proposed workflow is unsuitable has done its job. It doesn't owe the vendor a rollout.
The honest catch is ownership
A human checking the result sounds reassuring until nobody has time to check it. If reviewing the queue is necessary, assign the work, account for the time, and make skipped review visible. A person who is technically responsible but never sees the result isn't part of a control. They're a name in a document.
The same applies to authority. Reading a customer message, suggesting an answer, changing a record, and sending a message are different permissions. Give the workflow the smallest set it needs for the stage being tested. Expanding from suggestions to actions should follow evidence about the actual action, not excitement about the quality of the prose.
A fallback also needs to be usable. If the service is unavailable, can the team still open incoming requests and process them normally? Can a reviewer reject a suggestion without fighting the tool? Does a failed run stay visible, or does it disappear into a generic success message? These are operating questions that deserve answers before dependency grows.
Budget for small ongoing work. Someone will maintain categories, update examples when the business changes, check bills, and investigate complaints. That doesn't make the project a bad idea. It makes the cost estimate complete. Software that takes an hour a month to maintain can still save twenty; pretending it needs zero maintenance is how the owner ends up surprised.
Watch adoption after the novelty fades. Ask the people doing the work which suggestions they accept, which they rewrite, and which they avoid entirely. A high usage count can mean the tool is useful, or it can mean everyone is required to open it. Review behavior and finished work give that count meaning.
There is room to be ambitious here. A reliable first pass can expose consistent patterns, shorten training, and make a growing queue manageable. Those gains are enough to justify a sensible project without claiming the whole department is about to disappear. Once the narrow workflow works, the next improvement has a baseline to beat.
The AI worth keeping is the one whose savings survive the work of checking it.
Sources
Every claim above traces back to one of these. Go read them yourself.
- 01Building effective agents
Anthropic / anthropic.com / retrieved Sep 10, 2026
Suggested reading
Selected articles based on topic, tags, and skill focus across the library.
Foundational AI: Do's and Don'ts
Which part of your monthly report is actually an AI job?
A recurring management report does three separate jobs: it pulls numbers, it computes them, and it explains them. Only one of those is an AI job, and Microsoft retiring a spreadsheet function on the fourteenth is a useful reminder of which one.
Foundational AI: Do's and Don'ts
Your Automations Are Logged In As A Person
The ops lead left in July and the Monday invoice chase stopped in August, quietly, with no error and no alert. The handover doc was never the deliverable. The list of what runs under their name was, and three of the four platforms you use will not tell you it exists.
Foundational AI: Do's and Don'ts
Your Team Uses Personal AI Accounts. Ban It Or Buy It?
Twelve people on a paid AI workspace runs about $2,880 a year. Twelve people on their own logins runs zero, and what that zero buys is a set of data terms somebody on your payroll agreed to on their phone.

