UX Variations — monitoring-start

Alignment Status: review · Tier: prototype · Category: product-design · Skill: ux-variations · Date: 2026-06-18

Summary. This is a review (pre-approval) alignment page for five contrasting UX progression-mode variations of the gblock-party personal-workstation monitoring-start surface — the board where a solo “Agent Conductor” (5–12 concurrent agents across 1–3 repos) monitors their fleet (S3 Agent Board), launches new agents one-click (S4 Start Agent), and drills into a single agent in a folded, tool-adaptive detail surface (merged S5/S6).

The five variations deliberately spread four axes — V1 Operator Console (dense table + split-pane), V2 Card Mosaic (control), V3 Kanban Pipeline (status columns + drawer), V4 Command-First (palette + focus mode), V5 Glanceable Stream (mobile-first feed). Every variation’s full 23-section build-grade spec is rendered below with no context loss. The gates ask you to confirm the surfaced assumptions, the fixed-vs-variable scope, the concept set, the evaluation method, the proposed flow-map amendment, the artifact paths, coverage, and the post-approval route, before any /ui-interview routing. Awaiting your review.

2. Scope & Decision Surface

This plan varies the monitoring-start surface of the gblock-party personal workstation: the board where a solo “Agent Conductor” (5–12 concurrent agents across 1–3 repos) monitors their fleet and launches new agents one-click, then drills into a single agent to work it.

What is being varied

Two-layer surface. Every variation must serve both an orchestration layer (S3/S4 — fleet-level: scan, fan out parallel agents, launch fast) and a deep-work layer (folded S5/S6 — a single agent). What the variations contrast is how the orchestration layer connects to the deep-work layer.

Run mode. These are default progression-mode UX variations (how the user advances through the flow), not --layout-mode variations. Layout/density/nav differences fall out of each variation’s progression thesis rather than being the primary object of variation.

Parent flow branch: monitoring-start of design/user-flow-personal-workstation.md (Approved) — Happy Path steps 3–6 + Path C “Parallel Sprint”. Screens in scope: S3 Agent Board · S4 Start Agent · S5/S6 folded tool-adaptive Detail. Mode: Flat (single-product). Topic: monitoring-start.

3. Confirmed Reframe (applies to ALL variations)

Three load-bearing decisions from the brief are carried by every variation:

  1. Two-layer surface, not one job. S3/S4 = orchestration (fleet-level: scan, fan out, launch fast); folded S5/S6 = deep-work (single agent). Every variation serves both layers; they contrast only in how the layers connect.
  2. Folded + tool-adaptive detail. S5 (Structured) and S6 (Terminal) collapse into one detail surface whose primary rendering is determined by the tool’s integration pattern, not a user choice: terminal-primary (xterm.js over WebSocket) for headless-CLI tools (Claude Code CLI, Codex CLI, Cursor worker, Aider, Continue, Cline, Copilot CLI, Augment), structured for API-native tools (Claude Agent SDK). The old “Mode A vs Mode B user toggle” framing is dropped. Terminal latency still matters for CLI tools (it is the live interface) but it is intrinsic, not an optional feature toggle.
  3. All web PWA. S3–S6 all render in the browser PWA; even the terminal is xterm.js in the browser over WebSocket. PWA installable on mobile + desktop.

Proposed flow-map amendment (NOT silently edited)

The approved flow-map (design/user-flow-personal-workstation.md) lists S5 and S6 as two separate screens and frames Mode A/B as a user-selectable access mode (decision D2). The confirmed reframe contradicts that on two points:

Per the brief, variations carry the merged tool-adaptive model, and this canonical plan proposes it back to the flow-map as an amendment pending the flow-map owner’s acceptance. The approved flow-map document is not edited by this plan. (IDE / Mode C / code-server remains out of scope unless a variation makes a deliberate case; none does.)

4. The Four Variant Axes

The five concepts deliberately spread four axes so nothing is pre-decided:

AxisWhat it variesSpread across the set
a) Board representationhow the fleet is showndense table (V1) · card grid (V2) · status columns (V3) · quiet list (V4) · vertical feed (V5)
b) Orchestrate↔deep-work connectionhow drill-in relates to the boardpersistent split-pane (V1) · full-page route (V2) · side drawer over board (V3) · focus mode + back-stack (V4) · full-screen sheet (V5)
c) Launch affordancehow a new agent is started (S4)inline top-of-table row + ⌘K (V1) · modal overlay (V2) · “+” add-card per column (V3) · command palette ⌘K (V4) · duplicate-from-template sheet (V5)
d) Device posture / densitydesktop-dense ↔ mobile-firstdesktop-dense (V1) · co-equal responsive (V2) · desktop reflow→stacked (V3) · keyboard-desktop / mobile bottom-bar (V4) · mobile-first (V5)

Device posture is a deliberate axis, not an accident. “Desktop-primary” was explicitly rejected as a global default: instead, the set spans from V1’s desktop-dense extreme to V5’s mobile-first extreme, with V2 co-equal in the middle, so the right posture is tested empirically across the comparison rather than assumed up front.

5. Variation Comparison Matrix

#VariationBoard representationConnection modelLaunch affordanceDevice postureComplexityBest-fit userPrimary tradeoff
V1Operator Consoledense sortable data tablepersistent split-pane (board L / detail R)inline top-of-table row + ⌘Kdesktop-densemediumkeyboard power user at max density (8–12 agents, wide monitor)intimidating / weak on mobile; one-line cell truncates rich stdout
V2Card Mosaic controlresponsive card grid w/ live stdoutboard → full-page detail routemodal overlayco-equal responsivelow (baseline)balanced default; cross-device; recognition-over-recallcards waste space at 12 agents; route-switch drops board context
V3Kanban Pipelinecolumns = lifecycle statusboard → side drawer (board stays behind)“+” add-card in Queued columndesktop reflow → stacked sectionsmedium-hightriage-first / kanban mental modelcard motion/instability; weak deep-work continuity; slower pure launch
V4Command-Firstquiet minimal listpalette nav; detail = focus mode + back-stackcommand palette (⌘K)keyboard-desktop / mobile bottom-barmediumkeyboard-native dev (vim/tmux/Raycast)low discoverability; weaker fleet glanceability; command-typing awkward on phone
V5Glanceable Streamprioritized vertical feedfeed → full-screen sheetduplicate-from-template quick-start sheetmobile-firstmedium“closed laptop, checked from phone” conductor; launches mostly known recipesweak dense-desktop multitasking; feed reorders; template-launch less flexible for novel prompts

Axis Explorer interactive preview

Pick a value on each of the four axes; the variation(s) that best match your selection highlight. This is a preview — an exploration aid over the same data as the comparison matrix, not production code. No data is stored. Reset restores the initial state. Requires JavaScript; with JS disabled, the comparison matrix above and the “View as table” fallback below carry the same data.

a) Board representation

b) Connection

c) Launch

d) Device posture

V1 Operator Console
V2 Card Mosaic
V3 Kanban Pipeline
V4 Command-First
V5 Glanceable Stream

§6 below embeds the complete text of each intermediate spec, unsummarized. Each variation’s internal heading hierarchy nests under its variation heading. Source files: design/ux-variations-monitoring-start/{v1-operator-console,v2-card-mosaic,v3-kanban-pipeline,v4-command-first,v5-glanceable-stream}.md.

V1 — Operator Console (v1-operator-console)

Branch: monitoring-start · Variation id: v1-operator-console · Run mode: default progression-mode (NOT layout-mode). Parent flow-map: design/user-flow-personal-workstation.md (Approved) · Shared brief: design/_working/ux-variations-monitoring-start-brief.md. Screens in scope: S3 Agent Board · S4 Start Agent · S5/S6 folded tool-adaptive Detail. Status: Draft spec (build-grade) — not yet UI-interviewed.

1. Name & Thesis

Operator Console. The agent fleet is a NOC. One dense, sortable, filterable data table — one row per agent, 5–12 rows — with a persistent split-pane: the table holds the left column, the selected agent’s tool-adaptive detail holds the right. The fleet never leaves view. You triage like an SRE reading a monitoring dashboard: scan the status column, sort errors to the top, act inline (approve / unblock / stop) without leaving the row, and keep one agent open in the detail pane while the other eleven keep streaming beside it.

One-line thesis: Maximum-density, keyboard-driven fleet command — every agent on one screen, every action one keystroke away, drill-in without ever losing the fleet.

The design bet: a solo conductor running 12 agents does not want twelve cards to scroll through; they want a spreadsheet of their fleet and a terminal stapled to its side. Density is the feature, not a regression to apologize for.

2. Parent User Flow & Branch Relationship

This variation is one of five sibling treatments of the monitoring-start branch (Happy Path steps 3–6 + Path C “Parallel Sprint”). All five serve the same two-layer surface and contrast only in how the orchestration layer (S3/S4) connects to the deep-work layer (folded S5/S6).

idboard representationconnectionlaunchdevice posture
v1-operator-console (this)dense data tablepersistent split-paneinline top-of-table row + ⌘Kdesktop-dense
v2-card-mosaic (control)responsive card gridboard → full-page routemodal overlayco-equal responsive
v3-kanban-pipelinelifecycle status columnsboard → side drawer“+” add-card per columndesktop reflow → stacked
v4-command-firstquiet minimal listpalette nav + focus modecommand palette ⌘Kkeyboard-desktop / mobile bottom-bar
v5-glanceable-streamprioritized vertical feedfeed → full-screen sheetduplicate-from-template sheetmobile-first

Relationship to siblings: V1 sits at the explicit desktop-dense extreme of the device-posture axis (the brief’s deliberate counterweight to v5’s mobile-first). It shares v4’s keyboard-native instinct but diverges on glanceability: v4 hides the fleet behind a palette and shows a quiet list; V1 puts the entire fleet on screen at all times as structured tabular data. It is the antithesis of v2’s roomy cards. Mobile is V1’s deliberate degrade case, not its design center — that territory belongs to v5.

This spec carries the brief’s confirmed reframe (two-layer surface; folded tool-adaptive detail; all-web PWA) and does NOT silently edit the approved flow-map. Where it departs from flow-map structure, that is recorded in §11.

3. Target User Fit

Best fit: the keyboard-driven power user at peak load — the conductor running 8–12 concurrent agents across 2–3 repos from a wide desktop monitor, who lives in tmux/htop/k9s/Grafana and reads tables faster than cards. They want information per pixel, not whitespace; they sort and filter reflexively; they expect j/k to move and Enter to open.

Poor fit: a first-time user meeting agents for the first time (the table reads as intimidating cockpit chrome before there is anything to monitor); a phone-primary “quick-check” user (the table must collapse to survive small screens); a 1–3 agent light user (the density buys nothing — a card grid would serve them better, which is why v2 exists).

Fit signal: if the evaluator, during UAT, instinctively sorts the status column and uses keyboard nav to triage a 10-agent fleet faster than they did in v2/v5, this variation is earning its density.

4. Onboarding / Activation Model (first meeting + empty state)

V1 owns only the monitoring-start surface; first-run provisioning (S2 Setup Wizard, Path A) is upstream and out of scope. The activation question here is: how does a first-time user meet a NOC table with zero rows?

Empty state (zero agents): the table chrome is suppressed, not shown empty. A spreadsheet of nothing is hostile. Instead the split-pane collapses to a single centered column:

First-agent transition (activation moment): when the first agent launches, the empty column animates into the two-region split-pane. The first row appears with starting status and the detail pane auto-selects it (the only sensible selection), immediately showing live stdout. This is the “the console is now live” beat — the table header (sort/filter controls) fades in only once ≥1 row exists, so the user is never shown empty machinery.

Density ramp: columns are progressive. At 1–3 agents the table shows the full column set comfortably. The user is not asked to configure columns on day one; column customization is a power affordance discovered later (§14), never a setup gate.

5. Typical Workflow Sequence (orchestrate → drill-in loop)

The core loop. Numbered, keyboard-first, with mouse equivalents in parentheses.

  1. Open PWA → S3 console. Returning user lands on the populated table, default sort error > waiting > running > idle. The previously-selected agent is restored into the right pane (server-canonical; survives device switch).
  2. Scan the status column. Errors and approval-waits are sorted to the top and color-coded. A header chip shows counts: 2 waiting · 1 error · 7 running. Triage is a single visual sweep of the leftmost columns.
  3. Triage inline without drilling in. For a clear approval, press a on the focused row (or click the inline Approve affordance) — the row’s diff micro-preview expands one tier in place; confirm. For a blocked agent, u to unblock. For a runaway, x to stop (guarded, §15). The fleet stays put; no navigation.
  4. Drill in when a row needs attention. j/k (or arrow keys / click) to move the row cursor; Enter (or click the row) selects it into the right detail pane. The other 11 rows stay visible and keep streaming. This is the signature move: drill-in is selection, not navigation.
  5. Work the agent in the folded detail pane. The pane renders tool-adaptively (terminal-primary for CLI tools, structured for API-native — §13). Type into the terminal, read the structured step stream, inject a prompt, approve inline. The fleet table on the left keeps updating in your peripheral vision.
  6. Fan out a parallel sprint (Path C). To launch agent #2/#3 mid-flow: n (or click the inline + New Agent row pinned to the top of the table) — the top row becomes editable inputs (repo · tool · prompt). Fill and ⌘Enter to launch. New starting row drops into the table. Alternatively ⌘K → “New agent…” for a full quick-launch palette without leaving the keyboard.
  7. Round-robin approvals. As more agents hit decision points they sort to the top and the header count increments. Walk them: focus top waiting row → a → confirm → focus next. Or ⌘K → “Approve all clear” for the batch path (§15).
  8. Return to monitoring. Esc deselects the detail pane (collapses it wider on demand, §13) and returns focus to the table. Loop back to step 2.

The loop never leaves the single console route. Steps 2–3 are pure orchestration; steps 4–5 are deep-work; the split-pane is what fuses them.

6. Progression Model — and how it differs from the 4 siblings

How the user advances through the flow in V1: progression is lateral, not navigational. The user does not move between screens to advance; they move the selection cursor down the table and the detail pane updates underneath them. Advancement = re-sorting (state changes float rows up), selecting (binds a row to the pane), and acting inline (mutates state, which re-sorts). The flow’s S3→S5/S6 transition is collapsed into a pane-bind on one persistent route. There is no back-stack to manage because you never left the board.

Explicit contrast with each sibling:

Net: V1’s progression model is “sort + select + act-in-place on one persistent dense surface.” No other sibling keeps both the full fleet and a working agent on screen simultaneously — that simultaneity is V1’s defining progression property.

7. Sharing & Collaboration Model

N/A — single-user personal workstation. This is a solo dogfood environment; there is exactly one user (the Agent Conductor) and no second human to share with, hand off to, or collaborate alongside. There are no shared boards, comments, mentions, presence indicators, or co-editing. The only “handoff” is the same user across their own devices (cross-device continuity), which is server-canonical state, not collaboration — covered under §9 return-use. No sharing affordances are designed and none should be invented.

8. Permissions Model

N/A — single-user personal workstation. There are no roles, no per-row access control, no viewer/editor tiers, no team admin. The single user has full control of every agent in the table. The only authorization in the system is GitHub OAuth, which scopes what the agents may do to the user’s repos (pull/push/PR) — an upstream identity/repo-access concern owned by S1, not a board-level permissions model. The console exposes no permission UI and none should be invented. (The OAuth-expiry failure path — agents losing repo access mid-run — is a recovery concern, handled in §10, not a permissions model.)

9. Return-Use & Notification Model

Resume: on return, the console restores server-canonical state: same rows, same statuses, same sort, and the last-selected agent re-bound into the detail pane (its stdout buffer replayed on WebSocket attach). No reconnection ceremony — this is the flow-map’s core promise. If the user returns after a gap exceeding the recap threshold (flow-map D7), a dismissible recap strip docks above the table header — not a full-screen S8 takeover, because in the Operator Console the fleet itself is the recap. The strip summarizes “while you were away: 3 completed · 2 PRs · 1 error” with each item as a chip that, when clicked, selects the relevant row into the detail pane. Esc or “Got it” dismisses the strip and drops you onto the live table.

Approval surfacing within the board: approvals are first-class table citizens. A waiting agent sorts to the top with a waiting badge, its Approve / Reject affordances rendered inline in the row’s action cell, and a one-tier diff micro-preview expandable in place. The header carries a live approval-queue chip (2 waiting) that, when clicked or via ⌘K → “next approval”, jumps the cursor to the top waiting row. No approval ever requires leaving the console.

Error surfacing within the board: errors sort above everything, render a red error badge, and put the failure’s first line into the one-line stdout cell. The header error chip mirrors the count. Selecting the row binds the detail pane to the full error / stack trace with Retry / Restart / Open terminal actions (mapping flow-map S5 Error state).

Push notifications (PWA): web-push for approval-needed and error/completion events deep-links into the console with the relevant row pre-selected (not to a standalone approval screen — V1’s whole premise is that the board is the approval surface). On a desktop where the console is already open, the notification is redundant with the live header chips; on mobile it routes into the collapsed-list degrade view (§17) with the row expanded. Notification preferences themselves live in S10 Settings (out of scope here).

10. Failure Recovery Behavior (mapped to flow-map failure paths)

How the Operator Console surfaces and recovers each flow-map failure. The design principle: failures are table-and-pane states, surfaced where the user is already looking (header chips + row badges + detail pane), never as modal interrupts that hide the fleet.

Flow-map failureHow V1 surfaces itHow V1 recovers it
Backend unreachable (WS heartbeat >10s / health-check fail)A persistent top banner spans the full console width: “Reconnecting to workspace…” with a spinner. The table dims to ~70% opacity and renders last-known state with a “stale since HH:MM:SS” tag in the header. Inline mutate actions (approve/unblock/stop/launch) are disabled and visually greyed in every row.Auto-reconnect with exponential backoff. On reattach, banner clears, table un-dims, buffered WS deltas replay and re-sort the rows. Agents kept running on managed infra throughout — the dim signals “you’ve lost sight, not lost the fleet.”
Agent process crash (supervisor exit code ≠ 0)The agent’s row flips to red error badge, floats to the top via sort, and its one-line stdout cell shows the crash’s first line + a crashed tag. Header error chip increments.Supervisor auto-restarts transient crashes (≤3 retries) — the row shows a restarting (2/3) micro-state inline so the user sees the retry budget burn down. On persistent failure the row offers inline Retry / Restart; selecting the row binds the detail pane to the full crash output with Open terminal for manual intervention (flow-map Path D).
Network tunnel dropDistinct from backend-unreachable: a header pill reads “Tunnel reconnecting…” while the WebSocket itself may still be healthy. Table stays live but flagged.Tunnel auto-reconnects; pill clears. Agents unaffected (running on managed infra). Less severe than full backend loss → table is not dimmed, only flagged, to avoid crying wolf.
Model key / subscription expiry or rate-limitAffected agent’s row shows an amber paused badge with reason auth expired / rate-limited in the stdout cell. It does not sort as an error (it’s user-actionable config, not a crash) but sits above running.Inline “Update credential” affordance deep-links to S10 (Settings, out of scope). The agent’s session state is preserved; on credential update the row auto-resumes to running with no re-launch. Rate-limit case shows a cooldown countdown in the cell and self-resumes.
GitHub OAuth token expired / revoked (agent git op 401/403)Affected rows show amber paused — github auth badge and sort above running. Because all agents share the one OAuth identity, a header-level banner appears: “GitHub access expired — re-authorize to resume N agents.”One “Re-authorize GitHub” action in the banner re-runs OAuth; all paused-for-auth rows auto-resume on success. No per-agent action needed (single identity) — this is the one place a fleet-wide banner is warranted. No session state lost.
Git conflict during agent workThe agent pauses itself and sorts up as a waiting row (it is requesting approval, per flow-map). The stdout cell reads merge conflict — N files.Selecting the row binds the detail pane to the conflict detail. The user either resolves manually via the terminal in the folded detail pane (CLI tools) or injects an instruction (“resolve by taking theirs”) via the structured prompt box. Resumes inline.
Disk exhaustion (>90%)A header warning chip (disk 92%) appears before critical; at critical the backend pauses agents and their rows show paused — disk badges.Chip links to a cleanup action (old sessions / build artifacts). Pre-emptive, so the fleet rarely reaches forced-pause.
Browser / PWA crashNone visible to the running fleet (server-side). On reload the console rebuilds from server-canonical state.PWA service worker reloads to the console with current state and the prior selection restored. No data loss.
Concurrent device conflict (same agent acted on from two devices)Last-write-wins for approvals; the row’s action cell briefly flashes a “updated elsewhere” micro-toast so the user knows their stale click was superseded.Terminal sessions are separate tmux windows (independent). Console state is eventually consistent via WS — the table simply re-renders to truth.

Cross-cutting failure principle for V1: because the fleet is never hidden behind a route or modal, all of these failures remain visible in aggregate (header chips) and per-agent (row badges) at the same time. The split-pane means the user can be deep in one agent’s terminal and still see a second agent error appear at the top of the table — a property no route-switching sibling (v2) preserves.

11. Page & Flow Changes vs. the Flow-Map Baseline

The flow-map lists S3, S4, S5, S6 as four distinct screens/routes and frames Mode A (structured) vs Mode B (terminal) as a user-selectable access mode (flow-map D2). V1 carries the brief’s confirmed reframe and restructures as follows:

  1. S3 + S5 + S6 collapse onto ONE persistent route. There is no navigation from board to detail; the detail pane is a region of the console, bound by selection. The flow-map’s S3→S5/S6 “Downstream” handoff becomes an in-page pane-bind.
  2. S5/S6 fold into one tool-adaptive detail pane. The flow-map’s separate “S5 Structured” and “S6 Terminal” screens — and its Mode A/B user toggle (D2) — are dropped. The detail pane’s primary rendering is determined by the tool’s integration pattern, not a user choice: terminal-primary (xterm.js/WS) for headless-CLI tools, structured for API-native tools. There is no “switch to terminal” mode button as a primary affordance (an escape-hatch terminal toggle survives only for structured/API tools that also expose a shell, as a secondary control — §14).
  3. S4 changes from modal to inline. The flow-map’s centered modal-with-backdrop launch (S4) is replaced by an inline editable top-of-table row plus a ⌘K quick-launch palette. Launch happens in the table’s coordinate space, not over a dimmed board — preserving the never-lose-the-fleet principle even during launch.
  4. S8 Async Recap demotes from screen to strip. The full-screen recap becomes a dismissible docked strip above the table (§9), because the live table already is the recap surface in this layout.
  5. S7 Approval has no standalone screen. Approvals render inline in rows + in the detail pane; the flow-map’s dedicated S7 mobile-card is realized only in the mobile degrade view (§17), not on desktop.
  6. Default sort retained but extended. Flow-map’s error > waiting > running > idle is kept as the default, extended with starting/paused/stopped/done tiers and made user-resortable by any column (the flow-map sort was fixed; V1 makes it a table affordance).

These are variation-level proposals consistent with the brief’s amendment note; they do not edit the approved flow-map doc.

12. Navigation Model

Single-route, selection-driven, keyboard-first. The console is effectively a one-page app for the monitoring-start surface.

13. Screen-by-Screen Layout (regions & proportions)

All within one persistent route. Reference frame: desktop ≥1280px wide.

S3 — Console (the table) + persistent split-pane

┌──────────────────────────────────────────────────────────────────────────────┐
│ HEADER  gblock·console   [7 running ·2 waiting ·1 error]  [⌘K] [⟳tunnel] ⚙ │  ~52px
├──────────────────────────────────────────────────────────────────────────────┤
│ [filter: repo▾ tool▾ status▾]   sort: status▾            recap-strip(opt.) │  ~40px
├──────────────────────────────────────────┼────────────────────────────┤
│  + New Agent  (inline launch row)          │                              │
│ status·name·tool·repo·el·tok·act table       │   DETAIL PANE                │
│  ●E auth… aider api 12m 41k ⟳⏹          │   (folded, tool-adaptive)    │
│  ●W pay…  cc    web  4m 88k ✓✗         │   agent header / status      │
│  ●R index codex web  9m 120k ⏹  ◀───── binds selected row           │
│  ●R migr… sdk   api  2m 30k ⏹            │   PRIMARY RENDER:            │
│  ●I docs  cont… web 1h  9k ▶            │   terminal (xterm.js)        │
│  LEFT PANE ~55–62% width                     │   OR structured stream       │
│  table scrolls independently                 │   RIGHT ~38–45% width        │
│                                             │   footer: inject ▸ / stop    │
└──────────────────────────────────────────┼────────────────────────────┘

S4 — Start Agent (folded into the table, two ways in)

Folded S5/S6 Detail

Already described as the right pane above. Key point: there is no separate S5 vs S6 screen and no Mode A/B user toggle — one detail pane, rendering chosen by tool integration pattern. A secondary “Open full terminal” escape hatch exists for structured/API tools that happen to expose a shell, and a “Pop out” control can promote the pane to a focused full-width terminal for heavy interactive debugging (still the same route, the table temporarily yields width).

14. Key Components & Controls

15. Button & Link Behavior

All destructive or fleet-affecting actions are guarded or reversible; all resume actions are non-destructive and preserve session state.

16. Spatial Density, Sizing, Hierarchy

V1 is the explicit high-density end of the posture axis — density is the thesis, tuned, not careless.

17. Responsive Behavior at 3 Breakpoints

Mobile is V1’s deliberate degrade case (the design center is desktop-dense; v5 owns mobile-first).

18. Visual Tone

NOC / server-monitoring console. Dark, dense, instrument-panel calm-under-load. Monospace data, color-coded status as the dominant signal, tight gridlines, minimal chrome, no decorative illustration once the fleet is live (illustration appears only in the empty state). The aesthetic reference set: htop / k9s / btop / Grafana / a trading terminal — tools that pack maximum signal into a glance and reward a trained eye. Motion is restrained: rows re-sort with a quick settle, they don’t fly between regions (the deliberate anti-v3 choice); status badges may pulse subtly when waiting. The dark house palette is the baseline; status colors are the accent vocabulary. Confidence over friendliness — this screen should feel like a cockpit you’ve mastered, not a dashboard you’re being onboarded to.

19. Strengths

20. Risks & Failure Modes (of the design)

21. Implementation Complexity

Medium (matches the brief’s rating). Heavier than v2 (the card-grid control), lighter than v3 (kanban drag/animation/column state machine).

What drives the cost:

No multi-user, sharing, or permissions code (N/A) — that subtracts real complexity other product types would carry.

22. Prototype Scope (lightweight, evaluable build)

A lightweight prototype must include enough to let a solo evaluator run the §5 loop and feel the density tradeoff. Minimum:

  1. Populated fleet table with ~8–10 mock agents spanning every status (error / waiting / running / idle / paused / done / starting), default error > waiting > running > idle sort, at least sort-by-status and sort-by-elapsed working, and a status/repo filter.
  2. Live-ish one-line stdout per row (mock tail loop is fine) + status badges with color + text labels + status-tinted row edges.
  3. Keyboard navj/k/Enter row cursor + selection, Esc deselect, ⌘K palette (at least New-agent + Jump-to + Filter verbs).
  4. Persistent split-pane with a working selection bind: selecting a row renders the detail pane while the table stays visible and keeps updating. This simultaneity is the single most important thing to demonstrate — it’s V1’s whole thesis.
  5. Tool-adaptive detail — at least two mock agents: one CLI-type rendering a (mock or real) xterm.js terminal, one API-type rendering a structured step stream. Prove the auto-render-by-tool with no user toggle.
  6. Inline launch row producing a new starting row (mock launch) + ⌘K quick-launch equivalent.
  7. Inline actions — Approve/Reject on a waiting row (with one-tier diff peek), guarded Stop, Retry on an error row; action-cell clicks must not disturb the current pane selection.
  8. At least two failure surfaces rendered as table/pane states: backend-unreachable banner + table dim, and an agent-crash row with retry budget — to prove failures stay visible without hiding the fleet.
  9. Responsive proof — resize to ≤640px and show the table collapse to the stacked list + full-screen detail sheet (the honest degrade).

Mock WebSocket/stdout and mock approvals are acceptable; the evaluation target is the interaction model and density feel, not backend fidelity.

23. Readiness Signal — what makes this branch ready for /ui-interview

This branch is ready to advance to /ui-interview [v1-operator-console] when, during /uat --variant-evaluation against its siblings, the solo evaluator gives a clear signal such as:

If instead the evaluator wants V1’s split-pane fleet-visibility but v2’s card richness, or v5’s mobile posture, that is a consolidation signal (feed /consolidate-variations), not a ready-for-interview signal for V1 standalone. Readiness for /ui-interview specifically means: the Operator Console’s dense-desktop, split-pane, inline-action model is the chosen direction for the monitoring-start surface and is ready to be pinned down visually.

V2 — Card Mosaic (v2-card-mosaic) baseline / control

Parent flow branch: monitoring-start (“Board Monitoring & One-Click Start”). Run mode: default progression-mode UX variation (NOT layout-mode). Role in the set: BASELINE / CONTROL — closest to the approved flow-map’s low-fi wireframe notes; the other four variations are measured against this. Screens in scope: S3 Agent Board, S4 Start Agent, S5/S6 folded Agent Detail. Upstream: design/user-flow-personal-workstation.md, brief design/_working/ux-variations-monitoring-start-brief.md.

1. Name & Thesis

Card Mosaic. A familiar SaaS dashboard: the fleet is a responsive grid of rich agent cards, each a self-contained status tile with a live stdout preview. You scan the mosaic, launch from a modal, and drill into a single agent by navigating to a full-page detail route. The thesis is zero learning curve — anyone who has used Vercel, Railway, GitHub Actions, or a CI dashboard already knows how to read this board. It deliberately trades cleverness for legibility so it can serve as the control the other four concepts must beat.

Design center of gravity: glanceability per agent. Each card carries enough rich context (badge, tool, repo, elapsed, tokens, multi-line stdout) that the board answers “what is each agent doing right now?” without a drill-in. The cost — accepted on purpose — is fleet density: cards are large, so 12 agents scroll.

2. Parent User Flow & Branch Relationship

This variation realizes the monitoring-start branch: Happy Path steps 3–6 (S3 Board → S4 Start → S3 new card → S5/S6 Interact) plus Path C “Parallel Sprint” (launch agent #2, #3; monitor all; round-robin approvals).

It is the baseline/control. Where the brief’s four axes are spread across siblings, V2 takes the flow-map-native position on every axis:

AxisV2 positionFlow-map source
a) Board representationResponsive card grid w/ live stdoutS3 low-fi notes: “Responsive grid of agent cards (2–3 columns desktop, 1 column mobile)”
b) Orchestrate↔deep-work connectionBoard → full-page detail routeD3 “click existing → S5/S6” rendered as navigation
c) Launch affordanceModal overlayS4: “Access Mode A — modal overlay”; low-fi “Centered modal with backdrop”
d) Device posture / densityCo-equal responsive (balanced middle)S3 low-fi “2–3 columns desktop, 1 column mobile”

Because it tracks the flow-map most literally, V2 introduces the fewest net-new interaction concepts. The one deliberate divergence from the approved doc is the folded tool-adaptive detail (see §11) — a constraint carried by every variation per the brief’s confirmed reframe, not unique to V2.

3. Target User Fit

The balanced default Agent Conductor. Specifically:

Best-fit summary: the persona before they have strong opinions. V2 is what you reach for when you don’t yet know whether you’re a power-keyboard user (→ V1/V4), a triage-first thinker (→ V3), or a phone-up checker (→ V5). It degrades gracefully to 1-column mobile, so it is never wrong on any device, just never maximally optimized for one.

Weak fit: a max-density NOC operator babysitting 12 agents on a single 27" screen (cards waste vertical space — that’s V1’s job).

4. Onboarding / Activation Model

Entry into this surface is post-setup: Path A (onboarding) provisions the workspace and redirects to S3 with an empty board. V2 owns the empty → first-agent moment on the board.

Empty state (S3 Empty):

Activation = the first card appears. When the user completes the S4 modal, a card animates into the grid with a “starting” badge and flips to “running” within seconds, its stdout preview beginning to stream. That first live preview is the aha for this surface — the user sees the agent working without any SSH/tmux ceremony. The empty illustration is replaced by a real 1-card mosaic, which implicitly teaches the grid model.

Progressive disclosure: the modal’s “optional branch” field stays collapsed under a “More options” toggle on first run, so the first launch is the minimal repo + tool + prompt.

5. Typical Workflow Sequence

Parallel Sprint (Path C), the representative flow:

  1. Open PWA → lands on S3 Agent Board (returning session). Mosaic renders all current agents, sorted error > waiting > running > idle. Header shows approval-queue badge.
  2. Scan the mosaic. Each card’s badge + stdout preview answers “what’s happening” at a glance; no drill-in required for routine monitoring.
  3. Click Start Agent (header button) → S4 modal opens over the dimmed board.
  4. In the modal: pick repo (dropdown), pick tool (dropdown), type the prompt (textarea), optionally expand branch. Click Start Agent.
  5. Modal closes; a new “starting” card appears in the grid, transitions to “running,” stdout begins streaming. (Happy Path step 5.)
  6. Repeat 3–5 for agent #2 and #3 (different repos/branches). Board now shows 3+ live cards.
  7. An agent hits a decision point → its card flips to a waiting badge, pulsates subtly, and re-sorts toward the top; the header approval-queue count increments.
  8. Click the waiting card → navigate to the full-page detail route for that agent. The detail is folded + tool-adaptive: terminal-primary for CLI tools (xterm.js), or structured-primary for API-native tools. Handle the inline approval here.
  9. Click Back to Board → return to S3 (board re-reads server-canonical state; the card is now “running” again).
  10. Round-robin steps 7–9 across agents as approvals arrive.
  11. When agents finish, cards show a completion summary; a Review diffs CTA surfaces (handoff to S9, out of this branch’s deep scope but linked).

6. Progression Model — and How It Differs From the 4 Siblings

How the user advances in V2: by navigation and recognition. The orchestration layer (mosaic) and the deep-work layer (detail) are two distinct places. You progress by clicking a card to travel to its detail route, and you back out to return to the board. State is server-canonical, so the round trip is lossless — but it is a round trip: the board is replaced by the detail page, not held alongside it. New work enters via the modal, which is a temporary overlay on top of wherever you are.

This is the closest-to-baseline progression: a classic two-place dashboard → detail pattern with a modal for creation. Explicit contrast:

One-line positioning: every sibling deviates from the dashboard mental model on exactly one axis to test a hypothesis; V2 holds all four axes at the flow-map default so those deviations have something to be measured against.

7. Sharing & Collaboration Model

N/A — single-user personal workstation. Per the flow-map’s explicit non-goals (“Multi-tenant / team features (personal-workstation scope only)”) and the brief’s locked constraint, this product has exactly one user — the Agent Conductor who owns the managed VM and their own GitHub OAuth + model subscriptions. There are no shared boards, no invited collaborators, no co-viewing, no comments, and no handoff-to-another-person flow. The only “handoff” in scope is the same user moving between their own devices (cross-device continuity), which is server-canonical state, not collaboration. No sharing UI is designed.

8. Permissions Model

N/A — single-user personal workstation. There are no roles, no per-resource ACLs, no seat management, and no permission tiers, because there is exactly one principal. The only authorization in the system is external: GitHub OAuth grants the single user repo read/write + PR scopes so their agents can operate on their repos (flow-map S1), and the user’s own model keys/subscriptions authorize model traffic (BYO-client; the board never proxies). Those are credentials the user holds, not in-app permissions V2 grants to others. No permissions surface is designed; failures of those external credentials are handled as recovery paths (§10), not as a permissions UI.

9. Return-Use & Notification Model

V2 is a return-heavy surface (the Conductor checks in repeatedly across a day and across devices). Return mechanics:

10. Failure Recovery Behavior

Mapped to the flow-map Failure & Recovery table and the per-screen matrices:

Failure (flow-map)V2 board/detail behavior
Backend unreachable / WebSocket heartbeat timeoutS3 shows a top banner “Reconnecting to workspace…” with auto-retry (exponential backoff); cards freeze at last-known state; Start/Stop disabled until reconnect (S3 “Connection Lost” state). Detail route shows a “Reconnecting…” overlay; buffered output replays on reconnect (S5/S6 matrix).
Agent process crash (exit ≠ 0)The agent’s card flips to a red error badge with a one-line error summary and re-sorts to the top of the mosaic. Auto-restart for transient errors (≤3 retries) happens server-side; if it stays failed, the card offers Retry and Open (→ detail with full error output / stack trace, and a drop-to-terminal path for CLI tools). Maps to Path D (recovery).
Network tunnel dropSame banner pattern (“Tunnel reconnecting…”); cards unaffected in content (agents keep running on managed infra), just marked stale until the tunnel returns.
Model key / subscription expiry or rate limitAffected agent pauses; card badge → waiting/blocked with a “credential” reason; an inline link routes to S10 Settings to update the credential; agent resumes post-update without losing session state.
GitHub OAuth token expired/revokedAffected agents pause; a prompt to re-auth via GitHub OAuth appears (banner + card state); agents resume automatically once a new token is issued.
Git conflict during agent workAgent pauses and raises it as an approval with conflict details → handled in the folded detail (terminal for CLI tools so the user can resolve manually, or instruct the agent to resolve).
Disk exhaustion (>90%)Pre-critical warning banner on S3 with a cleanup suggestion; emergency pause of agents to prevent corruption, reflected as paused card states.
Browser/PWA crashService worker detects unclean shutdown; PWA reloads to S3 at current server state — no data loss (agents are server-side).
Concurrent device conflict (same agent, two devices)Last-write-wins for approvals; the mosaic on each device reconciles via WebSocket to eventual consistency; terminal sessions are separate tmux windows so they don’t collide.

S4 modal failures (launch fails / tool unavailable): inline error in the modal with a Retry; if the selected tool isn’t installed in the workspace, a warning with an Install tool link to S10 and Start disabled until a valid tool is chosen (S4 matrix).

Recovery design principle for V2: failures are expressed on the card (badge + summary + re-sort to top) so the board itself is the recovery queue — no separate error console. Deep remediation happens by clicking through to the folded detail.

11. Page & Flow Changes vs. the Flow-Map Baseline

V2 is intentionally the minimal-divergence variation. Tracking:

Net: two screens (S3, S4) match the baseline; the detail folds two baseline screens into one tool-adaptive route. This should propagate to the flow-map as an amendment, not a silent edit (per the brief).

12. Navigation Model

Place-based, route-driven — the canonical SaaS-dashboard nav:

Routing surface (illustrative): / board · /agents/:id folded detail · S4 and S8 as overlays/interstitials over whatever route is active · /settings (S10).

13. Screen-by-Screen Layout (Regions & Proportions)

S3 — Agent Board (mosaic)

S4 — Start Agent (modal overlay)

S5/S6 — Folded Agent Detail (full-page route)

Proportions summary: board = full-bleed grid; modal = centered ~520px card; detail = top bar + dominant primary region + a slim right rail.

14. Key Components & Controls

Sort/filter: default sort error > waiting > running > idle. Filtering is deliberately light in V2 (a simple status filter / repo filter chip row is the most it should add) — heavy filtering is V1’s territory; V2 leans on big legible cards instead.

15. Button & Link Behavior

All destructive/irreversible actions (stop, reject) get a confirmation affordance; launching and approving do not (they’re the happy path and reversible enough).

16. Spatial Density, Sizing, Hierarchy

17. Responsive Behavior at 3 Breakpoints

Mobile ≤ 640px (1 column):

Tablet ≤ 1024px (2 columns):

Desktop > 1024px (3 columns):

Co-equal posture means no breakpoint is the “real” one — the same components reflow; the phone gets a usable read-and-approve experience, the desktop gets comfortable multi-agent monitoring, neither is the privileged target. (Contrast: V1 privileges desktop, V5 privileges phone.)

18. Visual Tone

Familiar, trustworthy, boringly competent — on purpose. It should read like a polished deploy/CI dashboard (Vercel/Railway/Render lineage) so a new user feels instant fluency. Dark house palette; calm grayscale chrome with status color as the only loud element; soft card elevation; restrained motion (cards fade in on launch, waiting cards pulse gently, stdout scrolls smoothly — no gratuitous animation). The terminal in detail reads as a real console (monospace, darker bg). The overall message: “this is the safe, legible way to watch your fleet” — which is exactly the control’s job.

19. Strengths

20. Risks & Failure Modes

21. Implementation Complexity

Lowest of the five (baseline). Rough estimate: low.

Why it’s the floor:

Largest single risk to the estimate is external, not structural: stdout-parsing reliability and terminal latency (both shared open questions), not V2’s own UI.

22. Prototype Scope

For the serial build → /uat --variant-evaluation comparison, prototype:

Out of prototype scope (linked, not built here): full S9 diff review, S8 recap timeline beyond a stub, S10 settings, real push delivery, real tunnel/auth, and Mode C/IDE (excluded).

23. Readiness Signal for /ui-interview

This spec is ready to hand to /ui-interview [v2-card-mosaic] once V2 is selected or built. It fixes, for the visual-mockup stage:

Open items the UI interview should resolve (not blockers, just the live questions): exact card dimensions / preview line-count, the precise side-rail width and collapse behavior in detail, the light-filtering affordance (status/repo chips) if any, and mobile terminal ergonomics copy (“open on desktop” handoff). Bounding open questions (terminal latency, stdout-parsing reliability) are flagged as risks, not resolved here, and carry forward as evaluation criteria.