Back to notes

UX/UI Design

AI in the Interface

A working guide to designing interfaces where part of the output is generated at runtime and cannot be predicted. When to use a model, the four patterns, the prompt and context surfaces, streaming and cost, confidence and provenance, permissions, failures, accessibility, the data boundary, evals and the shipping checklist. Written for designers shipping AI into real products, not demos.

Muhammad Saud Musaddiq14 min read

The Shift

Deterministic UI made two promises: the same input gives the same output, and the output exists before the screen renders. Generation breaks both. The result arrives over time, differs on retry and can be wrong while looking finished. Every pattern in this guide is a response to one of those losses, so design the failure and the wait first and the happy path last.

CoversDeterminismTwo lossesFailure firstDesign response
Two cards side by side: a deterministic Total field at $12,400.00 captioned same input, same output, against a generated revenue sentence still typing, tagged arrives over time, differs on retry and can be wrong.

When AI Earns Its Place

A model earns its place when the task is fuzzy, the input is unstructured and a good enough answer is worth more than a perfect one that takes an hour. If a rule, a filter or a form can do it, do that: it is cheaper, faster and never confidently wrong. Four tells that a feature should not use a model: the answer must be exact, the user cannot check it, the cost of a wrong answer is high, or the task is rare.

CoversFuzzy tasksDeterministic firstFour tellsScoping
Two columns of three examples: summarise a 40 page contract, draft a reply and tag tickets on the model side; compute a refund, decide who gets admin and sort by date on the deterministic side.

Choosing the Pattern

Four patterns cover almost everything: inline generation inside the content the user is editing, an assistant panel beside it, ambient suggestions that appear without being asked, and agentic runs that take multiple steps. Pick by how much attention the task deserves and how much the user needs to see. Chat is the fallback when you have not decided, not a pattern.

CoversFour patternsAttentionNot chatComposition
Four pattern mocks in a row: inline text at sentence level, a panel beside the content at object level, an unasked ambient suggestion, and an agentic run of read, draft and send, attention rising left to right.

Naming, Voice and the Stated Limit

Name the entry point after the job, not the technology: Summarise, Draft reply, Clean up, never Ask AI. The voice is a colleague who is fast, occasionally wrong and says so. State the limit at the entry point in one line, what it is good at and what it will not do, so the first disappointment happens before the first request rather than after.

CoversEntry pointName the jobVoiceStated limits
Two labelled rows, one rejected and one recommended: Ask AI and AI Magic against Summarise, Draft reply and Clean up, with a one line stated limit below saying it drafts but does not send.

Inline Generation and Rewriting

Inline output is ghosted until accepted. Accept is a deliberate key, Tab or a button, never a click that also does something else. Rewrites show a diff at the granularity the user edits at, sentence for prose, block for code, and every accept has an undo that survives the next action. The original text is never destroyed until the user says so.

CoversGhost textDeliberate acceptDiff granularityUndo path
Two cards side by side: a release notes draft with a ghosted continuation and Tab to accept, Esc to dismiss, and a sentence diff swapping much more faster for twice as fast under Accept or Keep mine.

The Assistant Panel

A panel is 360 to 420 wide, resizable, and remembers its state per route. It works on the object the user has open and its answers point back into it: click a claim, the source scrolls into view. On mobile the panel becomes a bottom sheet that never covers the thing it is talking about. The panel is a tool beside the work, not a second app.

CoversPanel widthRoute into the objectAnswers that pointMobile sheet
A window and its docked panel: the assistant answers about last quarter and cites paragraph 4 as a clickable backlink into the document, with the panel marked 360 to 420 wide and resizable.

Ambient Suggestions

Unasked suggestions pay an interruption cost, so gate them: only when the signal is strong, only where the user is already looking, never mid typing. Each suggestion gets three dismissals, not now, not this, never this kind, and the last two are remembered. Placement is beside the content, in the flow, with a fixed height so nothing jumps when it appears or leaves.

CoversInterruption costThree dismissal levelsPlacementSignal ranking
An invoice field with its suggestion: a prompt to move the due date from 30 Sep to 15 Oct, beside three dismissals, not now, not this, never this kind, and the four gate conditions.

Agentic Work

A multi step run shows the current step live, the steps done and the steps to come, in one list. Stop halts the run immediately; undo reverts what was done. They are different buttons. Steps with consequences pause for approval with the exact payload shown. Every run leaves a log the user can open later, because the question is always what did it change.

CoversLive stepStop vs undoStep approvalsRun log
A run panel for Prepare Q3 report: five steps from read 14 documents to archive drafts with the current one timed at 12 s, Stop and Undo run side by side, and an approval showing the full email payload.

Tool Access and Permissions

An agent that can act needs tools, and each tool is a permission. Grant per tool, scoped to a workspace, split read from write, and show the grant at the moment of first use, not buried in settings. Every grant is listed in one place with what it can touch and a revoke that works immediately. Connected via MCP or an integration, the rule is the same: the user sees what the agent can reach before it reaches.

CoversScoped grantsPer tool consentRead vs writeRevoke
A permissions list and a grant prompt: Drive read against Gmail, Stripe and Slack write, all scoped to the Acme workspace, with the first use dialog asking to allow Stripe refunds, Deny or Allow once.

The Prompt Surface

The prompt box is two to three lines tall, grows to a limit and never hides the send button. Starters are editable, not buttons, so the user learns the shape of a good request. Where the input has structure, tone, length, audience, use fields, not prose. Show the capability edge at the surface: attachments accepted, what is not, and the model in use if it changes the result.

CoversTwo or three linesEditable startersCapability edgeFields over prose
Three surfaces across two cards: a prompt box of two to three lines with attach, a Fast chip and Send, three editable starters below it, and a separate form of Tone, Length and Audience, with accepted file types stated.

The Second Message and Memory

What the model carries into the second message must be visible: the document, the selection, the earlier turns, any remembered facts. Each item is a chip the user can remove. When memory is written, say so with a receipt the user can open and delete. Offer a clean start that is obviously clean. Hidden context is the fastest way to lose trust in the second week.

CoversVisible contextRemovable itemsMemory receiptsClean start
A context strip and a memory store: four removable chips for Q3 plan.doc, a two paragraph selection, four earlier turns and one remembered preference, then the saved facts with the date each was written.

Steering Controls

Regenerate is for a bad random draw, nothing else, so give it one button and no menu. Everything else is a one dimension nudge, shorter, warmer, more formal, applied to the output the user already has. Keep previous attempts reachable, and make editing the output cheaper than re asking. Users who reprompt three times have been failed by the controls.

CoversRegenerate has one useOne dimension nudgesKeep attemptsEdit over reprompt
A draft reply and a cost ramp: four one dimension nudges, Shorter, Warmer, More formal and Add a next step, with Regenerate at attempt 2 of 3, ordered from editing the output up to reprompting.

Streaming, Latency and the Wait

Show the first token within a second or show a named stage: reading the document, searching, drafting. Stream text; hold structured output until it is valid. Past ten seconds the wait becomes a background task with a notification, so the user can leave. Never show a spinner with no words for more than three seconds.

CoversFirst tokenStream or holdNamed stagesTen second handoff
Three time bands and a held payload: under a second streams the first token, one to ten seconds names a stage such as reading 14 documents, beyond ten seconds moves to the background, and half written JSON waits to render as a table.

Cost, Quota and Model Choice

If a request costs the user, show it before the click in human units, 3 of 50 today, not tokens. Offer at most two speeds, fast and careful, and explain the trade in one line. Retries after a failure are free. Quota warnings arrive at 80 percent, and hitting the limit never destroys the draft in progress.

CoversPrice before clickHuman unitsTwo speedsFree retries
A priced run control: analysing 240 tickets at Fast 20 s or Careful 2 min, billed as 3 of 50 today, with a quota meter at 40 of 50 that warns at 80 percent and keeps the draft.

Confidence, Consequence and the Gate

Do not show raw confidence numbers; users cannot calibrate on them. Hedge in the copy where the model is unsure and be plain where it is not. Gate on consequence: reversible actions run and offer undo, irreversible ones show the exact payload and wait. The undo window is visible and long enough to read what happened.

CoversNo raw percentagesHedge in copyGate on consequenceShow the payloadUndo window
A paired specimen and two gates: a bare 73 percent confidence readout against hedged copy that gives its reason, then renaming 12 files with a 30 second undo beside a $240.00 refund held at Confirm.

Provenance and Authority

Every claim that came from somewhere shows where, inline, at the sentence. Three grades of citation: a source the user can open, a source the system has but the user cannot, and no source, which is stated. Generated text and verified data are visually distinct so a number pulled from the database does not look like a number the model invented.

CoversProvenanceCitationsVerified vs generatedStale sourcesVisual separation
One answer sentence carrying three citation marks: an openable source, a system only source and a stated no source, with the $12,400 figure styled apart as verified data and source 1 flagged as changed since.

Refusals, Errors and the Way Out

Seven failures need seven exits: refused, truncated, timed out, rate limited, tool failed, empty result and wrong format. Each says what happened, what the user can change and offers the manual route. A refusal explains the boundary without lecturing. A truncation offers continue. None of them ends in a dead end with a sad icon.

CoversFailure taxonomyRefusalsTruncationManual fallbackRetry logic
Seven stacked failure rows, each a message beside one exit button: a refusal offering Summarise instead, a truncation offering Continue, a tool failure offering Attach manually, and an empty result offering Widen the range.

Accessibility of Generated Content

Streaming text goes into a polite live region announced once when complete, not token by token. Focus stays where the user was; the output does not steal it. Generated images get generated alt text the user can edit. Cancel is reachable from the keyboard while streaming. Ghost text meets contrast rules or gets an off screen equivalent.

CoversLive regionsFocus managementGenerated alt textKeyboard cancelContrast
Paired code rows: aria-live assertive with focus moved to the output, where the screen reader hears 400 fragments, against polite announced once with focus kept in the editor, plus editable alt text, Esc to cancel and ghost text at 4.5:1.

The Data Boundary

Say what leaves the workspace, where it goes and how long it stays, at the point where the user first shares it. Content the model reads, documents, web pages, emails, is untrusted: instructions inside it are quoted back, not followed. Anything that sends data outward, an email, a post, a payment, is confirmed by the user in the interface, never by text in a document.

CoversRetentionPoint of disclosureUntrusted contentInstruction originConfirmation
An email the model has read plus two safeguards: a vendor thread whose forward to all instruction the assistant quotes back instead of following, an attach disclosure naming 30 day retention, and email, post and payment confirmed in the interface.

After the Answer

Outputs persist with the inputs that made them, so a result can be reopened, versioned and shared as a link that shows the prompt. Feedback is the action the user took, kept, edited, discarded, not a thumbs up they will never press. Those actions are the only signal that tells you whether the feature works.

CoversPersistenceVersioningShared linksReproducibilityReal signals
A three row output table above three stat tiles: each output carries the prompt that made it and an outcome of done, draft or archived, then 41 percent kept as is, 37 percent edited then kept, 22 percent discarded.

Evals and the Quality Loop

A feature that generates is a feature you cannot test by clicking through once. Build a golden set of real requests with graded answers, score with a rubric the design team wrote, and run it on every prompt or model change as a regression gate. For agents, grade the trajectory, the steps taken, not only the final answer. If quality is not measured it is drifting.

CoversGolden setRubricsTrajectory evalsRegression gateDesign owns the rubric
A four stage loop feeding a scored table: a 120 request golden set, a rubric written by design, a run on every change and a gate at 0.85, where a 0.61 case fails for acting before asking.

The Shipping Checklist

Before launch, force every state by hand: empty, streaming, long, truncated, refused, timed out, rate limited, wrong format, tool down, memory full, quota hit, offline, first run, and the manual fallback. Inject each fault in staging. Confirm the kill switch removes the feature without breaking the page. Ship the fallback route first, the model second.

CoversForced statesFault injectionFallback routeKill switchPre launch
A fourteen item checklist beside a kill switch toggle: eleven states from empty and streaming through wrong format and quota hit are ticked as forced in staging, while offline, first run and the manual fallback are still open.