Choosing the first workflow in a regulated firm
The first agentic workflow should be chosen for reviewability and reversibility, not for impressiveness. A practical selection framework.
The first production workflow sets the internal narrative about AI for the next two years. Choose a spectacular, high-blast-radius process and one bad week ends the program. Choose a trivial one and nobody notices it shipped.
The right first candidate is high-volume, well-bounded, easy to verify, and cheap to reverse.
A five-question screen
Score every candidate workflow against these:
- Volume: does this run at least weekly, ideally daily? Rare processes never generate enough signal to improve.
- Verifiability: can a competent reviewer confirm correctness in under two minutes? If not, quality is unmeasurable.
- Reversibility: if an output is wrong and reaches a human, what is the cost? Prefer workflows where a mistake is caught, not compounded.
- Boundedness: are the inputs a known set of formats and the outputs a known structure? Open-ended scope is where pilots die.
- Ground truth: does a historical record of correct outputs already exist? That is your evaluation set, already paid for.
Start in shadow mode
Run the workflow against live volume in parallel with the existing process, producing outputs nobody acts on. Compare. You gain a real accuracy measurement on real data, at zero operational risk, and the comparison report is the most persuasive artifact you can bring to a risk committee.
Two to four weeks of shadow running usually replaces months of hypothetical debate about whether the system is trustworthy.
Instrument the baseline first
Measure the current process before you change it: cycle time, touch count, rework rate, error rate, cost per item. Firms that skip this step can never prove value afterward, because there is nothing to compare against.
