PwC — Document Extraction

Making enterprise AI understandable, reviewable and trustworthy.

Organizations process thousands of contracts, invoices and reports every day. Machine learning can extract the structured information in seconds — but business users still need confidence before they act on it. The challenge wasn’t building a smarter model. It was designing a workflow that helped people verify AI-generated information quickly, understand where it came from, and confidently correct it when necessary.

At a glance

Client
PwC — enterprise document intelligence platform
Role
Product design, UX strategy, interaction design, prototyping — alongside product managers and machine learning engineers. Mine: the review workflow end to end — evidence linking, uncertainty display, and the correction loop.
The problem
Users re-checked even high-confidence predictions by hand. The bottleneck wasn’t model accuracy — it was verification.
The decision
Design the entire workflow around one question — “Why did the AI make this prediction?” Evidence one interaction away. Uncertainty visible. Corrections feeding training.
The outcome
Verification stopped being the slowest part of the workflow — specifics under NDA.

Client documents are confidential — which is how I work inside regulated organisations. The demo below is a faithful reconstruction of the interaction model, built on fictional data.

The thinking is real; the invoice is not.

01 — The product

A platform that reads business documents so people don’t have to.

The model comes pre-trained and keeps learning from every correction. Instead of manually reading every page, users receive predicted values — and before those values flow into downstream business systems, a person reviews and approves them. That person is the product’s real user, and their confidence is the product’s real feature.

02 — The challenge

Users didn’t distrust the AI because it made mistakes.

They distrusted it because they couldn’t quickly verify whether a prediction was correct. Even high-confidence predictions were manually checked against the original document. The same loop, in every review session:

  • Read the extracted value. Search through the document. Find the matching sentence.
  • Confirm the prediction. Return to the extraction list.
  • Repeat — for a document with dozens of fields.

That context-switching became the slowest part of the workflow. The problem wasn’t accuracy. It was verification.

So the design goal was to reduce the effort required to answer one simple question — and nothing else:

“Why did the AI make this prediction?”

03 — Design decision 01

Keep every prediction connected to its source.

Instead of forcing users to search through long documents, selecting an extracted value immediately highlights its original location. Prediction and evidence stay visible at the same time — so instead of hunting for information, users simply verify it. The demo below carries all three design decisions. Try clearing the queue:

1 cleared automatically · 4 awaiting review

Invoice_2025-0417_Nordwind.pdf

Invoice issued by , Hamburg, for freight and warehousing services rendered in February 2025.

Invoice reference number: .

Invoice date: — payment due within 14 days.

The total amount payable is including VAT.

Extracted values

Select a value ↔ see its source. Confidence decides how much attention each one asks for — and corrections feed back to training.

04 — Design decision 02

Make uncertainty visible.

A confidence percentage alone rarely helps a user decide where to focus. So the interface uses visual priority instead — guiding attention toward the predictions where human judgement adds the most value, and letting the confident ones clear quietly.

97 %

Clears quietly

Model and document agree — approved without a human ever opening it.

79 %

Worth a glance

Plausible, but not certain — surfaced for a quick human look.

Review
51 %

Requires confirmation

The model is unsure — exactly where human judgement adds the most value.

Confirm

What each prediction asks of a human — example values from the demo above, where the interface shows priority, not raw scores.

Rather than asking users to interpret a score, the interface says this one deserves your attention — and lets the confident ones clear quietly.

05 — Design decision 03

Make corrections part of the learning process.

Review isn’t the end of the workflow. It’s how the model improves. When a prediction is wrong, users update the value while the original evidence is preserved — you did it in the demo when you corrected Nordwind to Nordwind Logistics GmbH. The correction stays linked to the document: high-quality feedback for future training, without ever interrupting the flow. Correcting feels like review, not debugging.

06 — Design principles

Four principles guided every interaction.

  • Every prediction should be explainable.
  • Evidence should always be one interaction away.
  • The interface should guide attention instead of demanding it.
  • Correcting should feel like review — not debugging.

07 — Outcome

Rather than asking people to trust AI, the product gives them the tools to verify it.

By reducing context switching, connecting every prediction to its source, and making uncertainty actionable, the review workflow becomes faster, easier to understand, and better suited for enterprise environments where accuracy is a liability question, not a preference. Interaction design can make complex AI systems feel transparent — not by hiding uncertainty, but by giving people the information they need to make confident decisions.

Specific figures are under NDA — the review model above is the part that travels.

Reflection

I used to think trust in AI was a model-quality problem.

It isn’t. Trust is verified, not claimed — and correcting AI should feel like reviewing a document, not debugging software.

Putting AI in front of people who can’t afford to act on a wrong answer?

Then the review workflow is the product. I design the verification loop — evidence, uncertainty, correction — as a working prototype your actual reviewers can test, before the model ships. The cheapest place to earn trust is before launch.