“Where can we use AI?” produces a portfolio before it produces a decision. Every department can name repetitive work, every vendor can demonstrate a model, and none of that tells you which workflow should be first.
Start with the work. Choose a small but consequential workflow whose need, system path, evidence, failure boundary and owner can be described before implementation. AI should earn its place inside that workflow rather than supplying the reason it exists.
The UK Government’s AI Playbook starts from business and user needs, asks whether AI has a meaningful advantage, and allows the conclusion that AI is not the best answer. Its mandatory language belongs to the government context it describes. The sequence is still useful outside government: need, advantage, impact, feasibility and ownership before procurement or build.
Name the finished work
A candidate needs a trigger and a finished outcome. “Email intelligence” is a category. “Read historic correspondence, separate people from machine traffic, and let a person approve which relationships enter the CRM” is a workflow.
The second description gives the team something to inspect:
- a known input estate;
- a read path and a later write path;
- a classification that is not mistaken for a business decision;
- a named approval before the CRM changes;
- an observable finished outcome.
That does not make the workflow ready. It makes the decision small enough to test.
Need
Name the repeated delay, error, hand-off or constraint and the people affected. Use an observed baseline where one exists. Write “unknown” where it does not, followed by the measurement that would make it known.
Frequency alone is not a reason to automate. A frequent task may be cheap, safe, satisfying or already better served by a rule. A rare task may carry enough consequence to deserve attention. The useful question is what the current path costs in time, error, delay, control or missed work, and how confidently that cost is known.
Avoid borrowed percentages and generic departmental claims. The first workflow belongs to the organisation’s own queue, not to a list of popular use cases.
Path
Draw the trigger, inputs, systems, permissions, writes and finished outcome. Confirm which parts are accessible now and which require a security, legal, vendor or data decision.
This is where many exciting candidates become ordinary integration work. That is useful information. A model that can classify a document does not create permission to read the source estate, a stable identifier for the record, an API that supports the intended write or a recovery path when the write is repeated.
The path can be intentionally narrow. Read-only discovery before any write, a review queue instead of direct posting, or one document class before an entire archive can turn an unbounded proposal into an operable first step.
Judgement
The team needs examples and criteria for distinguishing an acceptable output from an unacceptable one. This is not always one numeric metric. It may combine deterministic checks, expert review, pairwise comparison and explicit refusal cases.
NIST’s voluntary AI Risk Management Framework asks teams to document testing considerations, intended context, error costs, human oversight and risk tolerance. OpenAI’s evaluation process similarly begins with an objective, dataset and metrics tied to the task. We use those principles as evidence for one disqualifier: if nobody can say how the output will be judged, the workflow is not ready to build.
Consequence
Describe what a wrong action changes, whether it can be reversed and who can intervene. A classification placed in a review queue has a different consequence from an autonomous payment, even if both use the same model.
Human review is not a complete control by itself. Name where it occurs, what information the reviewer sees, what they may change, what happens when they disagree and whether the action has already taken effect. Approval before a write, monitoring after a write and human takeover are different arrangements.
The acceptable failure boundary should be written before implementation. If the team cannot describe one, pause. A promise to “add a human in the loop” later is not a boundary.
Advantage
Explain why AI is preferable to a rule, form, search, integration or process change. The answer may be a hypothesis, but it should be specific enough to test.
AI does not have to be the only possible technique. It has to offer a defensible advantage inside this workflow. That might be handling variable language, reducing the number of human decisions, or sorting a large unstructured estate before a person acts. It is not “because the task is manual”.
Do not turn advantage into a fake ROI calculator. Use a measured baseline when available, keep assumptions beside any modelled figure and preserve “unknown” when the organisation has not measured the current path.
Ownership
Name the domain owner, technical owner and post-handover home. They may be different people. Together they need authority over the business decision, the system path and the operating response.
A candidate without an owner is not waiting for engineering capacity. It is waiting for a decision about responsibility. We would not start it until that responsibility exists or until establishing it is the explicit first piece of work.
Ownership also includes the capability to maintain the result. The UK Playbook asks whether the required skills and infrastructure exist and whether training or a partner is needed. That does not force every capability in-house on day one. It makes the capability gap part of the plan instead of a surprise at handover.
Four reasons to wait
Pause or disqualify a candidate when any of these remains true:
- No owner. Nobody can accept the business decision, technical path and operating responsibility.
- No usable access. The required source system, data or permission cannot be reached on acceptable terms.
- No acceptable failure boundary. A wrong action cannot be contained, reversed or placed behind suitable authority.
- No credible evaluation. The team cannot distinguish an acceptable output from an unacceptable one before real work is affected.
“Wait” is not a verdict on AI. It names the missing work. A discovery phase may close the gap, a process change may remove the need, or another workflow may be ready first.
A worked candidate
Our correspondence-to-CRM pipeline began with a narrow unit: a decision per relationship rather than per message. The system read and classified; a person decided which relationships were worth reviving; nothing reached the CRM before sign-off.
The need was visible in an archive nobody could inspect at message level. The path was split into read access first and write access later. Judgement was bounded to person, organisation or machine traffic, while the business decision remained with the client. The write effect was reversible and approval-bound. The client owned every write.
The case page records the measured classification and approval counts and keeps modelled review arithmetic separate from client results. It does not publish an accuracy percentage or claim revenue from revived relationships. Those absences are part of the selection discipline: evidence that does not exist should not be invented to make a candidate look stronger.
The one-meeting canvas
| Screen | Write this down | Evidence state |
|---|---|---|
| Need | Repeated pain and people affected | Observed baseline, or unknown plus measurement action |
| Path | Trigger, systems, permissions, writes and outcome | Confirmed, constrained or inaccessible |
| Judgement | Examples and criteria for acceptable output | Testable, expert-reviewable or not observable yet |
| Consequence | Wrong action, reversibility and intervention | Contained, approval-bound or unacceptable |
| Advantage | Why AI beats the simpler alternative here | Evidenced hypothesis, not promised saving |
| Ownership | Domain, technical and receiving owners | Named, capability plan required or absent |
The canvas is not a predictive score. Its job is to expose a reason to build, a reason to wait and the evidence behind either decision.
Sources and limits
The screen is Esya’s delivery practice. It is informed by the UK Government AI Playbook and NIST’s voluntary AI RMF Core. Neither source validates this exact canvas, and neither says a candidate that passes it will succeed.
The visual uses a cleared Esya case as an explanatory worked example. It contains no client document, client identity or invented saving.