Agent Product Surface

The agent product surface is the user-facing layer where a human supervises autonomous work: intent, progress, tool calls, approvals, artifacts, errors, cost, pause/resume, and final proof.

Why this is a concept

Many pages already describe pieces of agent UI: AI-Native Design Patterns, Agent Machines, Loop, Browser Agent Console, Agent Trajectory Evaluation, and Latency Budgets for Agents. The missing concept was the product boundary: what a user must see and control for agent work to feel usable instead of spooky.

An agent product surface is not just chat. Chat is the intent channel. The product surface is the whole operating room.

Required surfaces

Surface User question it answers
Intent What did I ask this agent to do, and can I revise it?
Plan/state What phase is it in now?
Tool activity What is it reading, running, changing, or calling?
Approvals What side effect needs my consent, and what happens if I say yes/no?
Artifacts What did it produce: diff, file, screenshot, report, deployment, message?
Provenance Which sources, logs, traces, and commands support the result?
Control Can I stop, pause, resume, retry, fork, or cancel honestly?
Cost/latency How long is this taking and what is it spending?
Memory What will persist after the run?

Vercel's AI SDK and AI Elements ecosystem is useful IRL evidence that these are becoming reusable product primitives: chat transport, message parts, tool states, citations, reasoning, branch controls, prompts, and conversation components are becoming component surfaces rather than bespoke UI each app reinvents. Source: Vercel AI SDK docs, https://ai-sdk.dev/docs, 2026-06-19; Vercel AI Elements docs, https://ai-sdk.dev/elements, 2026-06-19

The durable user-facing contract splits into four concept pages:

Concept Surface implication
Human-in-the-Loop Control Approval cards need approve/edit/reject/respond/cancel semantics, not a generic "continue."
State Resumability Runs need visible state, resume/cancel/retry behavior, and stale-run handling.
Proof-Carrying Work Final artifacts need sources, traces, commands, validations, and unresolved risks.
Agent Unit Economics Expensive or long-running work needs cost, budget, and value visibility.

Product states

If the UI only supports "generating..." and "done," the product has no honest model of agent work.

Design rules

  • Progress should be factual: "Reading 7 files" beats "thinking."
  • Show product-level progress, tool activity, searches, and receipts—not private chain-of-thought or a synthetic “reasoning” transcript.
  • Stop must be honest. If cancellation is best-effort, say what will still finish.
  • Approval cards should show target, material inputs, diff/preview, affected system, authority/risk reason, timeout, exact proposal identity, and rollback/cancel path. Ask/questionnaire cards are a separate surface.
  • Tool results should have compact and expanded forms.
  • Artifacts should be copyable, downloadable, linkable, and attributable to a run.
  • Long-running work should survive refresh, navigation, and device changes.
  • Errors should explain the failed invariant and the next viable action.
  • The surface should expose the trajectory without making users read raw logs.

Ambient agent visuals are secondary state

The saved 25-orb post opens into two different, useful sources. AIcss is a machine-readable registry of 14 agent-conversation components—nine free and five licensed—covering thinking, reasoning, tool/action states, text, citations, diffs, tasks, tables, and agent input in React, Vue, and Svelte. Its free code may be used in personal and commercial apps, but source may not be redistributed as-is through a competing component library. Treat it as a job-shaped component shelf, not a bundle to mirror or install wholesale. The registry currently emits broken localhost:3000 page fields, so canonical URLs must be reconstructed from the public origin. [Sources: X 2085012911798292828; AIcss registry, AIcss machine guide, and usage terms, reviewed 2026-08-12]

For a persistent React status indicator, first evaluate the canonical thinking-orbs package rather than copying the whole AIcss gallery. Version 0.3.1 exposes nine truthful indeterminate verbs—working, searching, solving, listening, connecting, weaving, composing, breathing, and shaping—with separate 20px/64px designs, static reduced-motion frames, hidden-tab and offscreen suspension, live theme following, DPR control, and no production dependency beyond React. A discovered third-party skill still described only six states; canonical package state therefore outranks imported skill prose. The exact revision passed typecheck and production build, but has no executable upstream test suite, so host behavior still needs fixtures. Source: Jakubantalik/thinking-orbs@e04f3e87075faa6dd7d42f3073198434d26ba730, npm thinking-orbs@0.3.1, reviewed 2026-08-12

Neither source is an agent status protocol. An orb may reinforce a real running phase only when the same state is already expressed by durable runtime events, visible text, accessible status semantics, and an exact run identity. Never infer completion or authority from color, pulse, or speed; never let perpetual motion disguise a stalled run. The AIcss gallery's pinned browser audit found 375 CSS animations and 350 still running after their rendered groups were scrolled offscreen; reduced motion correctly dropped the count to zero, but a WCAG AA contrast audit still found 38 affected nodes. Adopt one status visual, not the catalog page, and prove contrast, non-color cues, cleanup, suspension, terminal-state removal, and frame/battery cost in the host. Kevin Wiki should not add an orb to its public agent until /api/system/agent exposes the real intermediate phase events the visual would represent. Source: reviews/source-explorations/aicss-agent-ui-registry-2026-08-12.json; reviews/source-explorations/thinking-orbs-canonical-package-2026-08-12.json

Selection Before Conversation

Chat is useful after the operator knows which goal, run, workflow, object, or artifact is in scope. The compact default should therefore expose a stable selection hierarchy before opening conversation:

goal/loop -> workflow/run -> owner/object -> source/proof -> action

The selected object keeps its identity while the user changes list, board, timeline, review, or graph views. A peek panel answers “what is this?” without losing place; an inspector answers “why is it in this state?”; the command palette answers “what can I legally do next?” Filters and URLs make the current cohort shareable. Agent chat consumes that same selection and capability context—it does not infer a hidden second context from the current screen.

Circle is a strong reference for this interaction grammar: compact grouped lists, URL-backed filters, hidden-item counts, an anchored multi-scale project timeline, contextual command routes, and an overview/guide/diff review sequence. It is also a useful counterexample. Its agent is keyword-matched canned text; its reviews, checks, commits, diffs, and deployments are seeded; its review checkboxes are not durable; and it has no backend, auth, database, API, or tests. Borrow the hierarchy and interaction patterns, never its proof semantics. Source: X/@ln_dev7, 2026-08-03; complete source, dependency, lint, typecheck, and production-build review of ln-dev7/circle@778598503e680b4c658d694dd9f65351ee48b3d3, 2026-08-12

Conversation History Is A Navigable Source

An agent-product history is not a flat transcript archive. The official Claude Agent SDK cookbook's session-browser example exposes the minimum durable operations: list sessions with branch/title/modified metadata, replay stored messages without starting an agent, rename and tag, filter, fork at a selected point, and resume the fork as a live run. Hosting examples preserve the same session identity across Docker, Modal, and Kubernetes behind one message/SSE interface. Source: anthropics/claude-cookbooks commit f65eb122a51e9710d4db3f4893016879c65c77d6, notebooks 05 and 07

For Kevin's system, the UI implication is direct: 535 parent conversations and their delegated traces should be browsable by harness, workspace, project, branch, time, tags, and learned/writeback state. A parent conversation is the navigation unit; subagent and workflow traces are expandable evidence, not fake extra conversations. Forking creates a new related source identity. A conversation only becomes brain knowledge after replay, interrogation, and owner writeback, so the surface must distinguish captured, reviewed, integrated, held, and superseded rather than showing one misleading "indexed" badge.

Recent Product Evidence

The 2026-06-30 bookmark review added three concrete surface patterns:

Source Product-surface lesson
Copilot Design System Multi-surface AI products need a visible focus-transfer protocol: global memory, local canvas execution, contextual entry point, and handoff state.
Clicky Screen-aware assistants can point, speak, and respond to the current desktop context; this is a different surface from a transcript-only chat.
Agent Engineering Skills Agent-assisted project updates should draft from recent product activity and Slack context, but the user still verifies, edits, and posts.
Nous Portal Model/agent products need team-level budget and usage controls, not only model-prompt UX.
Termany High-concurrency operators benefit from nested session location, working/done/attention state, inline diff/file/browser views, worktree and remote-host identity, port ownership, and usage accounting; those gains also expose shell, transcript, credential, process, and network authority that must stay visible and revocable.

Source: X bookmark artifact audit, 2026-06-30

The 2026-07-02 Clicky workflow demo sharpens the screen-aware row: the agent surface can be a compact desktop control panel beside the user's current app, with Home/Agents navigation, explicit voice/control affordances, cursor controls, app integrations, and an unlock/permission path. The surface lesson is not "agent speaks"; it is that voice needs visible control state and app context so proactive nudges remain inspectable. Source: X/@FarzaTV, 2026-05-16; Source: local video contact sheet, 2026-07-02

Product-Owned Agent Substrate

Rhys Sullivan's self-thread adds a product-surface rule for apps that expose agents: build the in-app agent on the same public skills and MCP surface that power users can bring to their own agents. The X article links were inaccessible through anonymous fetches, but the self-thread states the durable thesis clearly: the product should support users who want the built-in agent and power users who want their own agent, without splitting the underlying capability substrate. Source: X/@RhysSullivan self-thread, 2026-06-27

The practical implication is that product UI and external agent API should share one capability contract: skills, MCP servers, resources, and state boundaries. Otherwise the built-in assistant becomes a proprietary wrapper while external agents get a weaker or differently named tool surface. That divergence makes evaluation, docs, permissions, and support worse.

Public Agent Sequencing

An agent on the product's own domain is a convenience and operations surface, not the capability source. Build it in this order:

  1. ship composable primitives through OpenAPI or another typed schema, SDKs, CLI commands, MCP tools/resources, and machine-readable documentation;
  2. publish a shared context layer—skills, plugins, or equivalent guidance—that any supported harness can load progressively;
  3. let the on-domain agent reuse those same primitives and skills for people who do not arrive with a configured harness; and
  4. admit proactive work only through a separate identity, trigger, authority, approval, execution-isolation, and audit contract.

Vercel's current product separation is useful evidence. Its Plugin injects current platform knowledge and skills; its MCP server exposes public and OAuth-authenticated tools; and Vercel Agent combines platform context with a dashboard surface. The Agent documentation says it runs under its own identity, is bounded by the requester's permissions, remains read-only by default, asks for a scoped plan before elevated work, validates generated code in a Sandbox, and attributes elevated action to the agent, requester, and approver. Those are admission controls for action—not reasons to weaken the API available to an external agent. Source: Vercel Agent documentation, Vercel MCP documentation, and Vercel Plugin, reviewed 2026-08-11

The implementation details are versioned evidence, not marketing constants. The bookmarked Plugin page reported 26 skills, the current page rendered 28, and the exact repository tree at 12d0770 contained 32 top-level SKILL.md directories plus three agent definitions and five commands. Preserve the revision and derive counts from the source instead of copying a number into a durable product promise. The same route-level caution applies to "Markdown-over-the-wire": current live checks returned Markdown for Vercel docs and changelogs but HTML for the homepage and Plugin page, so content negotiation must be tested per route and represented honestly. Source: vercel/vercel-plugin@12d077072a7dac431c2fd4f675c2dd4627f7a196 and live Accept: text/markdown checks, 2026-08-11

The strongest reply to the original post states the product rule directly: make it the same agent whichever entrance the user chooses—one tool/context substrate over MCP for bring-your-own harnesses and native generative UI on the site. A built-in RAG widget with a cheaper model and no real capability fails that test. Source: X/@lukerramsden reply captured with the root post, 2026-07-17

Agent As Direct User

Simon Last's AGI-era product sketch adds the customer-shape version of the same rule: the direct product user may be an agent acting for a human. The reviewed image stack-ranks the work as: define the goal/problem, design general and well-engineered primitives, design the ideal agent interface, then design the human interface around talking to the agent plus direct access to low-level primitives. Source: X/@simonlast and local image review, 2026-07-04

The implication is not "hide the UI behind chat." The product still needs primitive-level affordances so the human can inspect, correct, and operate below the agent layer. Chat becomes the primary intent channel, while the rest of the surface shows the primitives, state, and evidence the agent is using.

Ivan Zhao's "APIs first, UI last" note is the short version of the product order. Define the primitive operations and agent-facing API before designing the human screen. Then keep the UI dumb: expose the primitive, show state, and let the agent/human compose it without burying core behavior inside a clever interface. Source: X/@ivanhzhao, 2026-06-26

Dylan Mikus' reply adds the pre-AI reason this rule was already good product design: an API lets users invent uses the product designer did not hard-code into the UI. Agents make that constraint more visible, because the direct user increasingly composes primitives through an API rather than a hand-authored screen. Source: X/@dbmikus, 2026-06-26

App-Layer Differentiation

Ivan Burazin's app-layer thesis is the negative product-surface test: many AI apps are "nice UI + a harness that gets models to the data + a sales team." That can close deals, but it is not defensible unless the surface encodes a differentiated workflow, state model, domain data boundary, evaluation loop, or distribution wedge. The product surface is where that differentiation must become visible: what the user can supervise, reuse, verify, and trust that a generic harness cannot provide. Source: X/@ivanburazin, 2026-06-24

Kevin project implications

Project Product-surface implication
Agent Machines Needs fleet/run views where machines, sessions, traces, approvals, artifacts, and cost are one inspectable surface.
Loop Refreshes need version history, diff, source traces, and review controls, not only a regenerated output.
Ariadne Operator surfaces must expose event state and intervention paths because real guests are affected.
Sigil UI Agent-authored UI should be constrained by tokens and components so generated surfaces stay coherent.
wiki UI Search and graph surfaces should show source, freshness, and backlink context rather than only page hits. Wikibot remains public-source Q&A until it has the same callable substrate plus identity, scoped authority, approval, audit, and recovery proof; chat presence alone does not make it a public actor.

Anti-patterns

  • Chat transcript as the only history.
  • Spinners with no last real event.
  • Tool calls hidden until failure.
  • Approval prompts with no diff or rollback story.
  • A "cancel" button that only hides the UI.
  • Cost/latency discovered after the invoice.
  • Generated UI that bypasses the design system.

Timeline