Two conversations that get confused
There are two separate conversations about AI and design. The first is about tooling: generating variants, producing mockups faster, summarising research verbatims. The second, discussed far less, is about the artefact itself: what happens to an interface when the system behind it produces a variable result — sometimes wrong, usually plausible? That second question is what genuinely reshapes the designer's and the researcher's craft.
Designing with AI tools
AI speeds up production: screen variants, UI copy, interview synthesis. The product itself stays deterministic — the same actions produce the same result.
Designing a product that uses AI
The product's behaviour depends on a probabilistic system. You have to design the expected output, the doubtful output, the wrong output, the way users take back control, and the means of verification.
The first changes how you work. The second changes what you are designing — and that is the subject of this guide.
Designing with uncertainty as material
In a conventional interface uncertainty is an exception: a loading state, a network error, an invalid field. In an AI-backed interface it is permanent. The system does not always know it is wrong, and the user cannot always check. Design therefore means making that uncertainty workable rather than hiding it.
- Show what the answer rests on, whenever the source exists and can be opened.
- Visually separate what the system produced from what the user wrote.
- Treat the absence of an answer as a designed state, not as emptiness.
- Make editing as reachable as accepting.
- Avoid authoritative wording when the system cannot back the claim.
A useful rule: the interface should never display more certainty than the system actually has. That is a product credibility question as much as an ethical one.
Communicating what the system can and cannot do
Most adoption failures come from an expectation gap: the user assumes the system understands everything, hits an arbitrary limit, and walks away. Design's job is to set the scope at the moment of use, not in a help page.
- A first screen showing three real uses instead of an empty input.
- Clickable examples that teach the scope without long explanatory copy.
- A clear, non-blaming message when the request falls outside the scope.
- A freshness indicator when the answer depends on dated data.
A free-text box is the broadest promise an interface can make. If the system only keeps part of that promise, it is better to constrain the input than to disappoint on the output.
Aim for calibrated trust, not maximum trust
The goal is not for users to trust the system as much as possible, but to trust it at the right level. Excess trust produces errors accepted without review. Insufficient trust produces abandonment, or double-checking that cancels the benefit. Calibration is built through consistent, honest signals.
- Consistency: the system behaves the same way in similar situations.
- Verifiability: a user can confirm an answer in one move, not five.
- Admitted limits: the system says what it could not find instead of filling the gap.
- Repairability: an error can be fixed without losing prior work.
Errors, imperfect answers and fallback states
Designing an AI product mostly means designing its degraded states. Most review decks show the ideal case; the real experience happens elsewhere. Four families of states deserve to be drawn explicitly, with the same care as the nominal one.
- Partial answer: the system found some of the information and says so.
- Uncertain answer: the answer exists but deserves checking, and the checking path is provided.
- No answer: the system abstains and offers something useful instead (search, contact, blank template).
- Wrong answer caught by the user: immediate reporting, easy correction, kept trace.
The last state is the most neglected and the most decisive. A user who spots an error and can do nothing about it loses trust for good; a user who fixes it in one move becomes a contributor to system quality.
Keeping the user in control
Taking back control is not a safety net, it is a core part of the experience. It is designed on three planes: before the action (choose the scope, adjust the request), during it (interrupt, rephrase, narrow) and after it (edit, undo, return to the previous state).
- Every irreversible action goes through explicit approval, never an automatic chain.
- Every generated output is editable in the surface where it appears.
- History explains what was produced, when, and from what.
- Users can turn assistance off without losing the underlying feature.
Human-in-the-loop: designing the act of supervision
When an organisation decides a human validates system outputs, that validation becomes a real user journey with its own cognitive load and its own risk. A poorly designed setup produces rubber-stamping: the user clicks accept without reading, and the control exists only on paper.
- Surface what must be checked first rather than the whole output.
- Distinguish high-stakes elements (amounts, dates, commitments) from the rest of the text.
- Cap the volume reviewed per session: beyond it, attention collapses.
- Make rejection as easy as acceptance, and capture the reason when useful.
A control that can be performed in one second without reading is not a control, it is a signature. Design decides which of the two actually exists.
User feedback as a product loop
Thumbs up and down is the zero degree of feedback: it signals dissatisfaction without saying which kind. A useful mechanism captures the nature of the problem with minimal effort and connects to the evaluation work done by product and engineering.
- Offer two or three concrete reasons rather than a free-text field: missing information, wrong information, wrong tone, off-topic.
- Capture the actual correction when the user rewrites the output — the richest signal available.
- Close the loop: say the report was taken into account when that is true.
- Do not ask for feedback on every interaction; sampling is enough.
Useful transparency without interface clutter
Explaining everything amounts to explaining nothing: a permanent warning banner becomes invisible within a week. Relevant transparency is contextual and progressive — a short marker on the answer, one-click access to the source, and full detail only for whoever asks.
Decorative transparency
A generic "answers may contain errors" banner everywhere, never read, legally protective but of no help to anyone deciding anything.
Operative transparency
At the level of the sentence involved: the cited source, its date, and direct access to the document. The user can verify that specific point without leaving their context.
Conversational: a choice, not a reflex
Chat has become the default shape of AI features, often for no reason. Conversation fits unpredictable, exploratory or iterative requests. It fits far less when the task is repetitive and well understood: there, a form, a contextual button or an inline suggestion is faster and more reliable.
- Known, bounded task: contextual action or structured field.
- Exploratory task with free phrasing: conversation.
- High-volume repeated task: automation with review by exception.
- High-stakes task: guided flow with explicit verification.
Copilot or full automation
Choosing between assisting and automating is a design decision, not a maturity level. A copilot leaves the decision to the user and gains in acceptability what it loses in time saved. Full automation delivers the largest gain but assumes errors are rare, detectable and reversible. In between, automation with review by exception handles only the confident cases and routes uncertain ones to a human.
The designer's role is to make that choice visible to the team, along with its consequences: who owns the output, what the user sees of the work done, and what happens when the system is wrong and nobody notices.
Running research on a non-deterministic journey
Classic user research relies on a stable path: several people perform the same task in the same interface. On an AI product, two participants asking the same question may get different answers. The protocol has to adapt without losing rigour.
- Test a task and a goal, not a frozen screen: what matters is whether people get there.
- Prepare several possible outputs in advance — good, partial, wrong — and observe the reaction to each.
- Deliberately introduce an error case: it is the only way to learn whether users detect it.
- Separate misunderstanding of the system from misunderstanding of the interface.
- Document the exact context of each session: version, available data, answer received.
Two measures deserve particular attention: error detection — does the user spot a wrong answer? — and over-trust — do they accept a plausible answer without checking? Those two signals say more about a feature's viability than a satisfaction score.
A concrete case: a recommendation that gets it wrong
An internal contract management tool suggests a standard clause based on the case context. In the first version the clause is inserted straight into the document with no indication of origin. During testing a participant accepts a clause unsuited to their case: it was well written, it looked coherent, and nothing invited them to check. The problem is not model quality — it is the absence of design around error.
What the revised version provides:
- The suggestion appears in a state visually distinct from approved text until it is accepted.
- It names the reference template used and the date it was last updated.
- Variable elements (duration, amount, jurisdiction) are highlighted as items to verify.
- When the context is incomplete, the system offers an empty clause rather than a plausible one.
- One-click rejection with an optional reason feeds the weekly quality review.
- History records who accepted what, and from which suggestion.
None of these decisions belong to the model. All of them belong to experience design, and they are what determines whether the feature is usable in a high-stakes context.
Working with product and AI engineering
On this kind of product the usual boundary between design, product and engineering moves. System behaviour is partly a design decision: what happens under doubt, what is shown, what is refused. Conversely, a technical constraint — latency, cost, data freshness — becomes an experience constraint.
- Take part in defining test cases: product edge cases are usually interface edge cases.
- Ask to see real outputs, including bad ones, before drawing screens.
- Express interface needs as expected behaviour, not as components.
- Anticipate the experience cost of latency: anything taking several seconds is designed differently.
A designer who has seen a hundred real outputs designs a radically different interface from one who has seen three demo examples. That is probably the most useful method change on this whole list.
Before signing off an AI feature
Eight design checks before treating an AI experience as shippable.
- The system's scope is understood at the moment of use, not in a help page.
- Partial, uncertain, empty and wrong states are designed as carefully as the nominal one.
- Generated output is visually distinct from human-approved content.
- Verifying an answer takes one move.
- Editing is as reachable as accepting.
- The supervision task is realistic: bounded volume, risky elements highlighted.
- Error reporting exists and is actually used by the team.
- Transparency is contextual, not a permanent generic banner.
Key points
- Designing with AI tools and designing a product that uses AI are two different crafts.
- The goal is calibrated trust, not maximum trust.
- Degraded states are the bulk of the design work on an AI product.
- A human review nobody can perform seriously is not a control.
- Chat is not the default form: it is chosen per task.
- Research tests a task against several possible outputs, not a frozen path.
- AI experience
- Human-in-the-loop
- Copilots