Agent Product Surface
The agent product surface is the user-facing layer where a human supervises autonomous work: intent, progress, tool calls, approvals, artifacts, errors, cost, pause/resume, and final proof.
Why this is a concept
Many pages already describe pieces of agent UI: AI-Native Design Patterns, Agent Machines, Loop, Browser Agent Console, Agent Trajectory Evaluation, and Latency Budgets for Agents. The missing concept was the product boundary: what a user must see and control for agent work to feel usable instead of spooky.
An agent product surface is not just chat. Chat is the intent channel. The product surface is the whole operating room.
Required surfaces
| Surface | User question it answers |
|---|---|
| Intent | What did I ask this agent to do, and can I revise it? |
| Plan/state | What phase is it in now? |
| Tool activity | What is it reading, running, changing, or calling? |
| Approvals | What side effect needs my consent, and what happens if I say yes/no? |
| Artifacts | What did it produce: diff, file, screenshot, report, deployment, message? |
| Provenance | Which sources, logs, traces, and commands support the result? |
| Control | Can I stop, pause, resume, retry, fork, or cancel honestly? |
| Cost/latency | How long is this taking and what is it spending? |
| Memory | What will persist after the run? |
Vercel's AI SDK and AI Elements ecosystem is useful IRL evidence that these are becoming reusable product primitives: chat transport, message parts, tool states, citations, reasoning, branch controls, prompts, and conversation components are becoming component surfaces rather than bespoke UI each app reinvents. Source: Vercel AI SDK docs, https://ai-sdk.dev/docs, 2026-06-19; Vercel AI Elements docs, https://ai-sdk.dev/elements, 2026-06-19
The durable user-facing contract splits into four concept pages:
| Concept | Surface implication |
|---|---|
| Human-in-the-Loop Control | Approval cards need approve/edit/reject/respond/cancel semantics, not a generic "continue." |
| State Resumability | Runs need visible state, resume/cancel/retry behavior, and stale-run handling. |
| Proof-Carrying Work | Final artifacts need sources, traces, commands, validations, and unresolved risks. |
| Agent Unit Economics | Expensive or long-running work needs cost, budget, and value visibility. |
Product states
If the UI only supports "generating..." and "done," the product has no honest model of agent work.
Design rules
- Progress should be factual: "Reading 7 files" beats "thinking."
- Show product-level progress, tool activity, searches, and receipts—not private chain-of-thought or a synthetic “reasoning” transcript.
- Stop must be honest. If cancellation is best-effort, say what will still finish.
- Approval cards should show target, material inputs, diff/preview, affected system, authority/risk reason, timeout, exact proposal identity, and rollback/cancel path. Ask/questionnaire cards are a separate surface.
- Tool results should have compact and expanded forms.
- Artifacts should be copyable, downloadable, linkable, and attributable to a run.
- Long-running work should survive refresh, navigation, and device changes.
- Errors should explain the failed invariant and the next viable action.
- The surface should expose the trajectory without making users read raw logs.
Ambient agent visuals are secondary state
The saved 25-orb post opens into two different, useful sources. AIcss
is a machine-readable registry of 14 agent-conversation components—nine free
and five licensed—covering thinking, reasoning, tool/action states, text,
citations, diffs, tasks, tables, and agent input in React, Vue, and Svelte.
Its free code may be used in personal and commercial apps, but source may not
be redistributed as-is through a competing component library. Treat it as a
job-shaped component shelf, not a bundle to mirror or install wholesale. The
registry currently emits broken localhost:3000 page fields, so canonical URLs
must be reconstructed from the public origin. [Sources: X
2085012911798292828; AIcss registry,
AIcss machine guide, and
usage terms, reviewed 2026-08-12]
For a persistent React status indicator, first evaluate the canonical
thinking-orbs package rather
than copying the whole AIcss gallery. Version 0.3.1 exposes nine truthful
indeterminate verbs—working, searching, solving, listening,
connecting, weaving, composing, breathing, and shaping—with separate
20px/64px designs, static reduced-motion frames, hidden-tab and offscreen
suspension, live theme following, DPR control, and no production dependency
beyond React. A discovered third-party skill still described only six states;
canonical package state therefore outranks imported skill prose. The exact
revision passed typecheck and production build, but has no executable upstream
test suite, so host behavior still needs fixtures. Source: Jakubantalik/thinking-orbs@e04f3e87075faa6dd7d42f3073198434d26ba730,
npm thinking-orbs@0.3.1, reviewed 2026-08-12
Neither source is an agent status protocol. An orb may reinforce a real running
phase only when the same state is already expressed by durable runtime events,
visible text, accessible status semantics, and an exact run identity. Never
infer completion or authority from color, pulse, or speed; never let perpetual
motion disguise a stalled run. The AIcss gallery's pinned browser audit found
375 CSS animations and 350 still running after their rendered groups were
scrolled offscreen; reduced motion correctly dropped the count to zero, but a
WCAG AA contrast audit still found 38 affected nodes. Adopt one status visual,
not the catalog page, and prove contrast, non-color cues, cleanup, suspension,
terminal-state removal, and frame/battery cost in the host. Kevin Wiki should
not add an orb to its public agent until /api/system/agent exposes the real
intermediate phase events the visual would represent. Source: reviews/source-explorations/aicss-agent-ui-registry-2026-08-12.json;
reviews/source-explorations/thinking-orbs-canonical-package-2026-08-12.json
Selection Before Conversation
Chat is useful after the operator knows which goal, run, workflow, object, or artifact is in scope. The compact default should therefore expose a stable selection hierarchy before opening conversation:
goal/loop -> workflow/run -> owner/object -> source/proof -> action
The selected object keeps its identity while the user changes list, board, timeline, review, or graph views. A peek panel answers “what is this?” without losing place; an inspector answers “why is it in this state?”; the command palette answers “what can I legally do next?” Filters and URLs make the current cohort shareable. Agent chat consumes that same selection and capability context—it does not infer a hidden second context from the current screen.
Circle is a strong reference for this interaction grammar: compact grouped
lists, URL-backed filters, hidden-item counts, an anchored multi-scale project
timeline, contextual command routes, and an overview/guide/diff review
sequence. It is also a useful counterexample. Its agent is keyword-matched
canned text; its reviews, checks, commits, diffs, and deployments are seeded;
its review checkboxes are not durable; and it has no backend, auth, database,
API, or tests. Borrow the hierarchy and interaction patterns, never its proof
semantics. Source: X/@ln_dev7, 2026-08-03; complete source, dependency, lint,
typecheck, and production-build review of
ln-dev7/circle@778598503e680b4c658d694dd9f65351ee48b3d3, 2026-08-12
Conversation History Is A Navigable Source
An agent-product history is not a flat transcript archive. The official Claude
Agent SDK cookbook's session-browser example exposes the minimum durable
operations: list sessions with branch/title/modified metadata, replay stored
messages without starting an agent, rename and tag, filter, fork at a selected
point, and resume the fork as a live run. Hosting examples preserve the same
session identity across Docker, Modal, and Kubernetes behind one message/SSE
interface. Source: anthropics/claude-cookbooks commit
f65eb122a51e9710d4db3f4893016879c65c77d6, notebooks 05 and 07
For Kevin's system, the UI implication is direct: 535 parent conversations and their delegated traces should be browsable by harness, workspace, project, branch, time, tags, and learned/writeback state. A parent conversation is the navigation unit; subagent and workflow traces are expandable evidence, not fake extra conversations. Forking creates a new related source identity. A conversation only becomes brain knowledge after replay, interrogation, and owner writeback, so the surface must distinguish captured, reviewed, integrated, held, and superseded rather than showing one misleading "indexed" badge.
Recent Product Evidence
The 2026-06-30 bookmark review added three concrete surface patterns:
| Source | Product-surface lesson |
|---|---|
| Copilot Design System | Multi-surface AI products need a visible focus-transfer protocol: global memory, local canvas execution, contextual entry point, and handoff state. |
| Clicky | Screen-aware assistants can point, speak, and respond to the current desktop context; this is a different surface from a transcript-only chat. |
| Agent Engineering Skills | Agent-assisted project updates should draft from recent product activity and Slack context, but the user still verifies, edits, and posts. |
| Nous Portal | Model/agent products need team-level budget and usage controls, not only model-prompt UX. |
| Termany | High-concurrency operators benefit from nested session location, working/done/attention state, inline diff/file/browser views, worktree and remote-host identity, port ownership, and usage accounting; those gains also expose shell, transcript, credential, process, and network authority that must stay visible and revocable. |
Source: X bookmark artifact audit, 2026-06-30
The 2026-07-02 Clicky workflow demo sharpens the screen-aware row: the agent surface can be a compact desktop control panel beside the user's current app, with Home/Agents navigation, explicit voice/control affordances, cursor controls, app integrations, and an unlock/permission path. The surface lesson is not "agent speaks"; it is that voice needs visible control state and app context so proactive nudges remain inspectable. Source: X/@FarzaTV, 2026-05-16; Source: local video contact sheet, 2026-07-02
Product-Owned Agent Substrate
Rhys Sullivan's self-thread adds a product-surface rule for apps that expose agents: build the in-app agent on the same public skills and MCP surface that power users can bring to their own agents. The X article links were inaccessible through anonymous fetches, but the self-thread states the durable thesis clearly: the product should support users who want the built-in agent and power users who want their own agent, without splitting the underlying capability substrate. Source: X/@RhysSullivan self-thread, 2026-06-27
The practical implication is that product UI and external agent API should share one capability contract: skills, MCP servers, resources, and state boundaries. Otherwise the built-in assistant becomes a proprietary wrapper while external agents get a weaker or differently named tool surface. That divergence makes evaluation, docs, permissions, and support worse.
Public Agent Sequencing
An agent on the product's own domain is a convenience and operations surface, not the capability source. Build it in this order:
- ship composable primitives through OpenAPI or another typed schema, SDKs, CLI commands, MCP tools/resources, and machine-readable documentation;
- publish a shared context layer—skills, plugins, or equivalent guidance—that any supported harness can load progressively;
- let the on-domain agent reuse those same primitives and skills for people who do not arrive with a configured harness; and
- admit proactive work only through a separate identity, trigger, authority, approval, execution-isolation, and audit contract.
Vercel's current product separation is useful evidence. Its Plugin injects current platform knowledge and skills; its MCP server exposes public and OAuth-authenticated tools; and Vercel Agent combines platform context with a dashboard surface. The Agent documentation says it runs under its own identity, is bounded by the requester's permissions, remains read-only by default, asks for a scoped plan before elevated work, validates generated code in a Sandbox, and attributes elevated action to the agent, requester, and approver. Those are admission controls for action—not reasons to weaken the API available to an external agent. Source: Vercel Agent documentation, Vercel MCP documentation, and Vercel Plugin, reviewed 2026-08-11
The implementation details are versioned evidence, not marketing constants.
The bookmarked Plugin page reported 26 skills, the current page rendered 28,
and the exact repository tree at 12d0770 contained 32 top-level SKILL.md
directories plus three agent definitions and five commands. Preserve the
revision and derive counts from the source instead of copying a number into a
durable product promise. The same route-level caution applies to
"Markdown-over-the-wire": current live checks returned Markdown for Vercel docs
and changelogs but HTML for the homepage and Plugin page, so content negotiation
must be tested per route and represented honestly. Source: vercel/vercel-plugin@12d077072a7dac431c2fd4f675c2dd4627f7a196 and live
Accept: text/markdown checks, 2026-08-11
The strongest reply to the original post states the product rule directly: make it the same agent whichever entrance the user chooses—one tool/context substrate over MCP for bring-your-own harnesses and native generative UI on the site. A built-in RAG widget with a cheaper model and no real capability fails that test. Source: X/@lukerramsden reply captured with the root post, 2026-07-17
Agent As Direct User
Simon Last's AGI-era product sketch adds the customer-shape version of the same rule: the direct product user may be an agent acting for a human. The reviewed image stack-ranks the work as: define the goal/problem, design general and well-engineered primitives, design the ideal agent interface, then design the human interface around talking to the agent plus direct access to low-level primitives. Source: X/@simonlast and local image review, 2026-07-04
The implication is not "hide the UI behind chat." The product still needs primitive-level affordances so the human can inspect, correct, and operate below the agent layer. Chat becomes the primary intent channel, while the rest of the surface shows the primitives, state, and evidence the agent is using.
Ivan Zhao's "APIs first, UI last" note is the short version of the product order. Define the primitive operations and agent-facing API before designing the human screen. Then keep the UI dumb: expose the primitive, show state, and let the agent/human compose it without burying core behavior inside a clever interface. Source: X/@ivanhzhao, 2026-06-26
Dylan Mikus' reply adds the pre-AI reason this rule was already good product design: an API lets users invent uses the product designer did not hard-code into the UI. Agents make that constraint more visible, because the direct user increasingly composes primitives through an API rather than a hand-authored screen. Source: X/@dbmikus, 2026-06-26
App-Layer Differentiation
Ivan Burazin's app-layer thesis is the negative product-surface test: many AI apps are "nice UI + a harness that gets models to the data + a sales team." That can close deals, but it is not defensible unless the surface encodes a differentiated workflow, state model, domain data boundary, evaluation loop, or distribution wedge. The product surface is where that differentiation must become visible: what the user can supervise, reuse, verify, and trust that a generic harness cannot provide. Source: X/@ivanburazin, 2026-06-24
Kevin project implications
| Project | Product-surface implication |
|---|---|
| Agent Machines | Needs fleet/run views where machines, sessions, traces, approvals, artifacts, and cost are one inspectable surface. |
| Loop | Refreshes need version history, diff, source traces, and review controls, not only a regenerated output. |
| Ariadne | Operator surfaces must expose event state and intervention paths because real guests are affected. |
| Sigil UI | Agent-authored UI should be constrained by tokens and components so generated surfaces stay coherent. |
| wiki UI | Search and graph surfaces should show source, freshness, and backlink context rather than only page hits. Wikibot remains public-source Q&A until it has the same callable substrate plus identity, scoped authority, approval, audit, and recovery proof; chat presence alone does not make it a public actor. |
Anti-patterns
- Chat transcript as the only history.
- Spinners with no last real event.
- Tool calls hidden until failure.
- Approval prompts with no diff or rollback story.
- A "cancel" button that only hides the UI.
- Cost/latency discovered after the invoice.
- Generated UI that bypasses the design system.
Timeline
-
2026-08-12 | Added the agent-orb boundary: animation may reinforce a real run event but never replace text, accessibility semantics, authority, or durable state; reduced motion, cleanup, frame cost, and exact snippet rights remain required. Source: X
2085012911798292828; aicss -
2026-08-12 | Added conversation history as a first-class navigable source: parent sessions, expandable evidence traces, replay without execution, branch/tag/filter/fork/resume controls, and visible capture-to-writeback state. Source:
anthropics/claude-cookbooks@f65eb122; local multi-harness session inventory -
2026-08-12 | Added selection-before-conversation as the compact agent-surface hierarchy and calibrated Circle as an interface reference only: its scan/detail/inspector, filters, timeline, command palette, and review sequence are useful, while its agent, diffs, checks, deployment, persistence, and authority are mock or absent. Source: X
2084328572160704943;ln-dev7/circle@778598503e680b4c658d694dd9f65351ee48b3d3 -
2026-08-12 | Added Termany as a retained agent-cockpit reference: nested location, attention states, worktrees, remote hosts, inline review, ports/processes, and cost are valuable surfaces, while transcript/credential/shell/network authority and AGPL deployment remain explicit adoption boundaries. Source: X
2079431675873038816;thinkany-ai/termany@e121820 -
2026-08-11 | Clarified that AI-native progress surfaces expose product-level work and receipts rather than hidden reasoning, and that approval cards require target-bound proposal identity, authority, timeout, and rollback—not merely a polished questionnaire treatment. Source: Beautiful UI exact-site/source replay; X
2082479500944904432 -
2026-06-19 | Created to bridge design, projects, agent runtime, observability, and AI SDK UI primitives into one product concept. Source: whole-wiki concept review, 2026-06-19
-
2026-06-19 | Added explicit control, state, proof, and economics concept links as the durable user-facing contracts for agent products. Source: whole-wiki concept review, 2026-06-19
-
2026-06-30 | Added concrete evidence from Copilot Design System, Clicky, Agent Engineering Skills, and Nous Portal: focus handoff, screen-aware agents, verified agent drafts, and team spend controls are product-surface requirements. Source: X bookmark artifact audit, 2026-06-30
-
2026-07-02 | Added Farza's newer Clicky workflow demo as desktop companion evidence: voice-first interaction still needs visible control state, app integrations, cursor controls, and permission/unlock affordances. Source: X/@FarzaTV, 2026-05-16; Source: local video contact sheet, 2026-07-02
-
2026-07-03 | Added the product-owned agent substrate rule from Rhys Sullivan's self-thread: an app's built-in agent and users' external agents should share public skills/MCP capability contracts where possible. Source: X/@RhysSullivan, 2026-06-27
-
2026-07-04 | Added Simon Last's "agent as direct user" product sketch: build strong primitives and an agent interface first, then give humans chat plus direct primitive access for supervision. Source: X/@simonlast and local image review, 2026-07-04
-
2026-07-04 | Added Ivan Zhao's API-first/UI-last rule as the terse product-order version of the agent-as-direct-user surface: primitives and agent-facing APIs come before the human UI. Source: X/@ivanhzhao, 2026-06-26
-
2026-07-04 | Added Dylan Mikus' API-first reply: APIs let users and agents invent product usage beyond the designer's first UI path. Source: X/@dbmikus, 2026-06-26
-
2026-07-04 | Added Ivan Burazin's app-layer critique: UI plus a model/data harness plus sales is not durable product differentiation unless the surface carries workflow, state, data, eval, or distribution advantage. Source: X/@ivanburazin, 2026-06-24
-
2026-08-11 | Added the public-agent sequence: composable APIs and tools first, shared progressive context second, an own-domain convenience surface third, and proactive work only after identity, least privilege, scoped approval, sandbox validation, attribution, and recovery. Current Vercel primary surfaces also proved that counts and Markdown delivery must be tested per revision and route. Source: X/@rauchg and replies; Vercel Agent, Plugin, MCP, and repository replay