Team mapping one workflow to assess where AI can help, where data sits and where human review is needed
Team mapping one workflow to assess where AI can help, where data sits and where human review is needed

What is an AI workflow assessment?

Workflow, adoption and value

An AI workflow assessment is a structured review of one specific workflow to decide whether AI can help, where it should fit, and what conditions must be in place before anything is built. It maps how the work really moves, identifies friction, checks data and decision points, defines human review and sets measures for success. Unlike an AI opportunity assessment, it does not compare many ideas. It diagnoses one chosen workflow.

Reviewed by Jackie, Head of Learning & Development, Levellers - Last reviewed 8 June 2026

What this means

Once a leadership team has chosen a candidate use case, the next question is narrower and more practical: is this particular workflow genuinely suitable for AI? That is the job of an AI workflow assessment. It is a diagnostic, not a sales exercise.

The assessment looks at the real flow of work from trigger to completed result. It asks where information arrives, where it gets stuck, where judgement happens, where data is weak, and which steps are repetitive enough for AI support. It also asks which steps should remain human because the cost of error, the need for explanation or the rights at stake make that sensible.

This article sits between the portfolio level AI opportunity assessment and the later stage of workflow redesign. If you are still deciding which work to pursue, move up to the opportunity assessment. If you already know the workflow is worth changing and need to reshape the process itself, move on to workflow redesign.

Why it matters

An AI workflow assessment matters because many failed AI efforts were never workflow ready. A model may parse text well, classify images well or draft language well, yet still fail in practice because the workflow around it is too variable, too poorly instrumented, too dependent on tacit judgement or too risky to hand over without strong review. The assessment prevents leaders from mistaking technical capability for operational fit.

It matters commercially because one workflow can hide several different jobs. A team may describe a process as "complaints handling" or "invoice processing", but inside that label sit very different tasks, some repetitive, some uncertain, some judgement heavy and some controls related. The assessment breaks that bundle apart so the organisation can target the right work rather than forcing AI into the whole chain.

It matters for quality because improvements often come from much more than speed. Better routing, cleaner data capture, clearer exception handling, better review packs and stronger traceability may matter as much as reduced manual effort. An assessment makes those levers visible.

It matters for governance because this is where consequential issues become concrete. If personal data is involved, if decisions affect individuals in significant ways, or if the workflow touches higher scrutiny activity, the organisation needs to understand exactly where human authority must sit, how exceptions are escalated and what evidence must be retained.

Finally, it matters for change management because staff will judge AI on whether it makes their work clearer and safer, not just faster. A workflow assessment is one of the best ways to answer the frontline question, "what exactly changes for me?"

How it works

Define the workflow boundary properly

The first discipline is boundary setting. Leaders need to identify one workflow and define where it starts, where it ends and what counts as a case. Without that, the review becomes vague very quickly.

A workflow boundary should include the trigger, the finished state, the owner, the main systems touched and the main case types. It should also say what sits outside scope. For example, "incoming supplier invoices from receipt to approved posting" is a workflow. "Finance admin" is not. "New client onboarding from first accepted proposal to activated account" is a workflow. "Client set up" is not.

Case types matter because many workflows have a quiet standard lane and a noisy exception lane. If those are mixed together, the average picture can mislead. A workflow may look unsuitable for AI because the exceptions dominate everyone's memory, even though the standard cases are highly structured and frequent. The reverse can also happen. The workflow looks orderly on paper, but in practice the exceptions consume most of the time and managerial attention.

Map the workflow as it actually runs

Many teams have a formal process document. What they need here is the real process. That usually requires a mix of methods. Interviews are useful. So is screen sharing. So is short shadowing. Where event data exists, process mining or log analysis can reveal handoffs, bottlenecks, loops and waiting time far more reliably than memory.

The map should show each step, each actor, each decision point, each handoff, each key information input and each review gate. It should also show where work waits, where it returns for rework and where cases diverge into alternative paths. The aim is not to create a pretty diagram. The aim is to expose friction.

Leaders should pay special attention to four points. First, handoffs, because delay, loss of context and duplication often arise here. Second, queue points, because work that sits idle usually signals either unclear ownership or poor sequencing. Third, review steps, because many workflows contain layers of checking that exist because upstream information is weak. Fourth, exception handling, because a workflow can look highly standard until the true variants become visible.

Separate the workflow into task types

AI rarely improves an entire workflow in one move. It usually improves particular task types inside that workflow. That is why the assessment should decompose the work.

Typical task types include extracting fields from documents, classifying requests, matching records, summarising case history, drafting first responses, retrieving policy content, recommending next steps, flagging anomalies, preparing review packs and routing cases to the right queue. Some of these are strong AI candidates. Others remain squarely human.

This task level view helps in two ways. It stops leaders from overpromising broad automation where only a few steps are genuinely suitable. It also reveals where value might come from a relatively small intervention. A workflow can remain mostly human while still improving materially because one badly timed or repetitive task is lifted out or supported.

A useful question is this: which exact pieces of work consume time before a person applies judgement? The more of that preparation burden a workflow contains, the more likely it is to benefit from AI support. Another useful question is this: where does human judgement genuinely add value that should stay central? The more consequential and context dependent the decision, the higher the bar for AI involvement.

Locate friction and describe the real problem

The next step is to translate the map into specific friction statements. "The workflow is slow" is too broad. "Cases wait in queue because incoming information is incomplete and triage rules are unclear" is useful. "Senior reviewers spend time compiling evidence that already exists across four systems" is useful. "Staff draft near identical responses but must search for precedent manually" is useful.

Good friction statements are important because they stop organisations from prescribing AI before the problem is clear. A workflow might be slow because of missing data rather than because staff are reading too much. It might be error prone because of inconsistent templates rather than because decisions are hard. It might have a review bottleneck because approval thresholds are badly set rather than because people need drafting help.

This diagnosis stage should also note what would happen if the organisation did not use AI at all. Sometimes the right answer is a simpler workflow rule, a standard form, better field validation, clearer ownership or a small integration. That does not make the assessment a failure. It means the assessment has done its job honestly.

Test suitability for AI, step by step

A workflow is more likely to suit AI when several conditions are present together. The work is frequent enough to matter. Inputs are mostly digital. Patterns recur often enough for comparison. Quality can be judged with reasonable clarity. Error can be detected and corrected. Exceptions are identifiable. There is enough structure that AI support would fit the operational rhythm rather than confuse it.

Suitability falls when the workflow depends heavily on local context, rare edge cases, face to face interpretation, contested evidence, or legally significant judgement that must be explained case by case. It also falls when the volume is too low to justify the change effort.

This step works best if teams rate each major task inside the workflow rather than applying one blanket judgement to the workflow as a whole. A single workflow may contain one very strong candidate step, one borderline step and one clearly unsuitable step. The assessment should say so plainly.

Reviewability is a key part of suitability. If staff cannot tell when the AI has gone wrong, the task is a weak candidate. If they can review quickly, override safely and learn from exceptions, the fit is stronger.

Check data, access and rights

Once a task or step looks promising, the assessment needs a stricter data view. Where exactly do the inputs come from? Are they complete, current and linked to the case? Are key fields free text or structured? Are documents machine readable? Are there historic examples? Who owns the data? What permissions apply? Can the organisation lawfully use the material in the way proposed?

This is also where privacy and rights screening becomes practical rather than theoretical. If the workflow uses personal data, the team needs to examine legal basis, transparency, minimisation, retention, access and whether a data protection impact assessment is likely to be required. If the workflow informs or makes decisions with legal or similarly significant effects on individuals, meaningful human review becomes critical.

Teams should also check intellectual property and confidentiality. A workflow might appear attractive until leaders realise it depends on sending client, employee or commercially sensitive material into a setting the organisation cannot justify. In many cases the assessment reveals that the workflow is not impossible, but it does require a narrower design or a different technology choice.

Design the human-AI pattern before any build work

The right question is not simply "can AI do this step?" It is "what role should AI play here?" In practice there are several patterns. AI can assist by retrieving, summarising or drafting. It can triage by routing obvious standard cases. It can recommend by producing a ranked view or a proposed next step. It can automate only the lowest risk slice of work. Each pattern carries a different control burden.

The assessment should define who remains responsible, what review is needed, what must be logged, when a case escalates, and what happens if the AI response is unavailable or clearly wrong. It should also be honest about training. A workflow that uses AI well often requires staff to review differently, challenge differently and document differently.

Human oversight is not a token sign off at the end. It must be meaningful. That means the reviewer understands the task, has authority to challenge or override, and is not set up to rubber stamp material at unrealistic speed.

Set the baseline and decide the next move

An AI workflow assessment is incomplete without a baseline. Leaders need a current measure of the workflow before anything changes. Depending on the case, useful measures include cycle time, queue age, first pass accuracy, rework rate, exception rate, escalation rate, staff effort, complaint rate, override rate and service level performance.

The point is not to measure everything. It is to select a small set of indicators that reflect the actual reason the workflow was chosen. A team trying to improve triage should not judge success mainly by a soft satisfaction score. A team trying to reduce risky inconsistency should not judge success only by average handling time.

At the end, the workflow should land in one of four states. Proceed to pilot. Redesign first, then pilot. Fix prerequisites such as data or ownership, then reassess. Or stop, because the workflow is a poor fit right now. The ability to land on each of these states is what makes the assessment worthwhile.

Examples

A managed services firm chooses its service desk ticket triage workflow for assessment. The review finds that issue intake, categorisation and first draft response are promising AI supported steps, but root cause diagnosis is far more variable and depends on tacit system knowledge. The result is a narrow design: AI helps classify and prepare context, while engineers retain diagnosis and client facing decisions.

A wholesale distributor reviews supplier invoice handling from receipt to posting. The workflow map shows that the real bottleneck is not approval time. It is the effort spent matching inconsistent supplier documents to purchase records and chasing missing fields. That makes document extraction, matching support and exception flagging the strongest AI candidates. Full no touch processing remains out of scope because the exception lane is still too messy.

A care provider examines recruitment administration for frontline hiring. Uploading documents, checking completeness, drafting routine candidate messages and scheduling are reasonably structured. Assessing suitability for employment remains human and tightly governed. The assessment therefore supports a bounded recruitment admin use case rather than an AI hiring system.

A housing organisation reviews repair request handling. The workflow appears straightforward until the assessment separates standard repairs from safeguarding related cases and vulnerable residents. Standard triage becomes a plausible AI supported step. Cases involving risk to wellbeing, disputed responsibility or complex resident history remain on a stronger human path with clearer escalation.

Common misunderstandings

One misunderstanding is that a workflow assessment is basically a model test. It is not. Model quality matters, but this stage is mostly about work design, decision points, data and control.

Another is that if one task inside a workflow looks suitable, the whole workflow is suitable. In reality, mixed fit is common. Strong assessments identify which parts should change and which should not.

Another is that the formal process document is enough. It rarely is. Many important delays, loops and workarounds live outside the official description.

Another is that the goal is always automation. Often the most useful pattern is preparation, triage or drafting support that improves a human step rather than replacing it.

A final misunderstanding is that if staff say "every case is different", AI is off the table. Sometimes that is true. Sometimes it hides the fact that a large share of cases still follow repeatable patterns and only a minority genuinely differ.

Risks and boundaries

The first risk is shallow mapping. If the organisation assesses the workflow from memory rather than evidence, it may miss the handoffs, exceptions and rework that determine whether AI will actually help.

The second risk is poor decomposition. When leaders treat a workflow as one block of work, they either overstate automation potential or reject a promising intervention because one step is too difficult.

The third risk is underestimating rights and governance. Workflows involving personal data, employee matters, customer eligibility, pricing exceptions, legal exposure or other significant decisions require a much more careful human review design.

The fourth risk is automation bias. If reviewers are set up to approve AI generated material under time pressure, the organisation may create the appearance of control without its substance.

The scope boundary is also important. This article is not the right tool for choosing among many possible use cases across the firm. That is the job of the AI opportunity assessment. Nor is it the same as changing the whole process shape. If the workflow is viable but structurally poor, the next step is workflow redesign.

What to do next

First, choose one shortlisted workflow only. Resist the urge to assess several at once. Focus creates better judgement.

Second, define the workflow boundary, case types and current baseline. Make sure everyone means the same thing when they refer to the workflow.

Third, map the real path of work with operational owners and frontline users. Where data exists, use it. Where it does not, observe the workflow directly.

Fourth, break the workflow into task types and identify where friction actually sits. Frame the problem in operational terms, not in abstract AI terms.

Fifth, rate the fit of each candidate task for AI support, then add a hard check on data, rights, oversight and exception handling. Be especially careful where personal data or consequential decisions are involved.

Sixth, define the human AI operating pattern in plain language. Say what AI will do, what humans will do, when cases escalate and how review will work.

Seventh, decide the correct next move. That may be a pilot, a redesign effort, a data clean up project or a no go. The value lies in reaching the right decision early, not in forcing every assessed workflow into build.

Have a question or a suggestion, or want to understand how we research and review these guides? Read about our editorial standards and how to reach us.

FAQs

How long should an AI workflow assessment take?

For a bounded workflow in an SME, it can often be done in days rather than months if the owner, data access and frontline participation are available. More complex regulated workflows take longer because mapping, rights checks and review design are heavier.

Do we need process mining software?

No. It is useful where good event data exists, but many assessments start with interviews, observation, document review and simple timestamps. The key is evidence, not a specific tool.

Can one workflow contain more than one AI use case?

Yes. That is common. A workflow may contain extraction, triage and drafting opportunities at the same time. The assessment helps decide which of those are sensible and in what order.

What usually makes a workflow a poor fit?

Low volume, highly variable cases, unclear quality standards, weak data, hard to detect errors, or decision steps where contextual human judgement and explanation are central.

Is this only for generative AI?

No. The assessment works for classification, prediction, optimisation, extraction and generative tasks. The method is about workflow fit, not one AI type.

Does every promising workflow need redesign before a pilot?

Not always. Some workflows only need a narrow intervention. But if delay, rework and handoffs are the real problem, redesign should come before or alongside a pilot.

Who should own the assessment?

Usually the operational owner of the workflow, supported by someone who understands process analysis, data and governance. Ownership should stay close to the work rather than sit only with an innovation team.

Can we use AI in decisions about people automatically?

That is an area for extra caution. Where personal data and significant decisions are involved, leaders should assume stronger governance, meaningful human review, and careful privacy screening are needed before proceeding.

Sources