UX/UI Design
AI in the Interface
A working guide to designing interfaces where part of the output is generated at runtime and cannot be predicted. When to use a model, the four patterns, the prompt and context surfaces, streaming and cost, confidence and provenance, permissions, failures, accessibility, the data boundary, evals and the shipping checklist. Written for designers shipping AI into real products, not demos.
Muhammad Saud Musaddiq14 min read
The Shift
Deterministic UI made two promises: the same input gives the same output, and the output exists before the screen renders. Generation breaks both. The result arrives over time, differs on retry and can be wrong while looking finished. Every pattern in this guide is a response to one of those losses, so design the failure and the wait first and the happy path last.

When AI Earns Its Place
A model earns its place when the task is fuzzy, the input is unstructured and a good enough answer is worth more than a perfect one that takes an hour. If a rule, a filter or a form can do it, do that: it is cheaper, faster and never confidently wrong. Four tells that a feature should not use a model: the answer must be exact, the user cannot check it, the cost of a wrong answer is high, or the task is rare.

Choosing the Pattern
Four patterns cover almost everything: inline generation inside the content the user is editing, an assistant panel beside it, ambient suggestions that appear without being asked, and agentic runs that take multiple steps. Pick by how much attention the task deserves and how much the user needs to see. Chat is the fallback when you have not decided, not a pattern.

Naming, Voice and the Stated Limit
Name the entry point after the job, not the technology: Summarise, Draft reply, Clean up, never Ask AI. The voice is a colleague who is fast, occasionally wrong and says so. State the limit at the entry point in one line, what it is good at and what it will not do, so the first disappointment happens before the first request rather than after.

Inline Generation and Rewriting
Inline output is ghosted until accepted. Accept is a deliberate key, Tab or a button, never a click that also does something else. Rewrites show a diff at the granularity the user edits at, sentence for prose, block for code, and every accept has an undo that survives the next action. The original text is never destroyed until the user says so.

The Assistant Panel
A panel is 360 to 420 wide, resizable, and remembers its state per route. It works on the object the user has open and its answers point back into it: click a claim, the source scrolls into view. On mobile the panel becomes a bottom sheet that never covers the thing it is talking about. The panel is a tool beside the work, not a second app.

Ambient Suggestions
Unasked suggestions pay an interruption cost, so gate them: only when the signal is strong, only where the user is already looking, never mid typing. Each suggestion gets three dismissals, not now, not this, never this kind, and the last two are remembered. Placement is beside the content, in the flow, with a fixed height so nothing jumps when it appears or leaves.

Agentic Work
A multi step run shows the current step live, the steps done and the steps to come, in one list. Stop halts the run immediately; undo reverts what was done. They are different buttons. Steps with consequences pause for approval with the exact payload shown. Every run leaves a log the user can open later, because the question is always what did it change.

Tool Access and Permissions
An agent that can act needs tools, and each tool is a permission. Grant per tool, scoped to a workspace, split read from write, and show the grant at the moment of first use, not buried in settings. Every grant is listed in one place with what it can touch and a revoke that works immediately. Connected via MCP or an integration, the rule is the same: the user sees what the agent can reach before it reaches.

The Prompt Surface
The prompt box is two to three lines tall, grows to a limit and never hides the send button. Starters are editable, not buttons, so the user learns the shape of a good request. Where the input has structure, tone, length, audience, use fields, not prose. Show the capability edge at the surface: attachments accepted, what is not, and the model in use if it changes the result.

The Second Message and Memory
What the model carries into the second message must be visible: the document, the selection, the earlier turns, any remembered facts. Each item is a chip the user can remove. When memory is written, say so with a receipt the user can open and delete. Offer a clean start that is obviously clean. Hidden context is the fastest way to lose trust in the second week.

Steering Controls
Regenerate is for a bad random draw, nothing else, so give it one button and no menu. Everything else is a one dimension nudge, shorter, warmer, more formal, applied to the output the user already has. Keep previous attempts reachable, and make editing the output cheaper than re asking. Users who reprompt three times have been failed by the controls.

Streaming, Latency and the Wait
Show the first token within a second or show a named stage: reading the document, searching, drafting. Stream text; hold structured output until it is valid. Past ten seconds the wait becomes a background task with a notification, so the user can leave. Never show a spinner with no words for more than three seconds.

Cost, Quota and Model Choice
If a request costs the user, show it before the click in human units, 3 of 50 today, not tokens. Offer at most two speeds, fast and careful, and explain the trade in one line. Retries after a failure are free. Quota warnings arrive at 80 percent, and hitting the limit never destroys the draft in progress.

Confidence, Consequence and the Gate
Do not show raw confidence numbers; users cannot calibrate on them. Hedge in the copy where the model is unsure and be plain where it is not. Gate on consequence: reversible actions run and offer undo, irreversible ones show the exact payload and wait. The undo window is visible and long enough to read what happened.

Provenance and Authority
Every claim that came from somewhere shows where, inline, at the sentence. Three grades of citation: a source the user can open, a source the system has but the user cannot, and no source, which is stated. Generated text and verified data are visually distinct so a number pulled from the database does not look like a number the model invented.

Refusals, Errors and the Way Out
Seven failures need seven exits: refused, truncated, timed out, rate limited, tool failed, empty result and wrong format. Each says what happened, what the user can change and offers the manual route. A refusal explains the boundary without lecturing. A truncation offers continue. None of them ends in a dead end with a sad icon.

Accessibility of Generated Content
Streaming text goes into a polite live region announced once when complete, not token by token. Focus stays where the user was; the output does not steal it. Generated images get generated alt text the user can edit. Cancel is reachable from the keyboard while streaming. Ghost text meets contrast rules or gets an off screen equivalent.

The Data Boundary
Say what leaves the workspace, where it goes and how long it stays, at the point where the user first shares it. Content the model reads, documents, web pages, emails, is untrusted: instructions inside it are quoted back, not followed. Anything that sends data outward, an email, a post, a payment, is confirmed by the user in the interface, never by text in a document.

After the Answer
Outputs persist with the inputs that made them, so a result can be reopened, versioned and shared as a link that shows the prompt. Feedback is the action the user took, kept, edited, discarded, not a thumbs up they will never press. Those actions are the only signal that tells you whether the feature works.

Evals and the Quality Loop
A feature that generates is a feature you cannot test by clicking through once. Build a golden set of real requests with graded answers, score with a rubric the design team wrote, and run it on every prompt or model change as a regression gate. For agents, grade the trajectory, the steps taken, not only the final answer. If quality is not measured it is drifting.

The Shipping Checklist
Before launch, force every state by hand: empty, streaming, long, truncated, refused, timed out, rate limited, wrong format, tool down, memory full, quota hit, offline, first run, and the manual fallback. Inject each fault in staging. Confirm the kill switch removes the feature without breaking the page. Ship the fallback route first, the model second.
