UX Variations — monitoring-start

Alignment Status: review · Tier: prototype · Category: product-design · Skill: ux-variations · Date: 2026-06-18

Summary. This is a review (pre-approval) alignment page for five contrasting UX progression-mode variations of the gblock-party personal-workstation monitoring-start surface — the board where a solo “Agent Conductor” (5–12 concurrent agents across 1–3 repos) monitors their fleet (S3 Agent Board), launches new agents one-click (S4 Start Agent), and drills into a single agent in a folded, tool-adaptive detail surface (merged S5/S6).

The five variations deliberately spread four axes — V1 Operator Console (dense table + split-pane), V2 Card Mosaic (control), V3 Kanban Pipeline (status columns + drawer), V4 Command-First (palette + focus mode), V5 Glanceable Stream (mobile-first feed). Every variation’s full 23-section build-grade spec is rendered below with no context loss. The gates ask you to confirm the surfaced assumptions, the fixed-vs-variable scope, the concept set, the evaluation method, the proposed flow-map amendment, the artifact paths, coverage, and the post-approval route, before any /ui-interview routing. Awaiting your review.

2. Scope & Decision Surface

This plan varies the monitoring-start surface of the gblock-party personal workstation: the board where a solo “Agent Conductor” (5–12 concurrent agents across 1–3 repos) monitors their fleet and launches new agents one-click, then drills into a single agent to work it.

What is being varied

Two-layer surface. Every variation must serve both an orchestration layer (S3/S4 — fleet-level: scan, fan out parallel agents, launch fast) and a deep-work layer (folded S5/S6 — a single agent). What the variations contrast is how the orchestration layer connects to the deep-work layer.

Run mode. These are default progression-mode UX variations (how the user advances through the flow), not --layout-mode variations. Layout/density/nav differences fall out of each variation’s progression thesis rather than being the primary object of variation.

Parent flow branch: monitoring-start of design/user-flow-personal-workstation.md (Approved) — Happy Path steps 3–6 + Path C “Parallel Sprint”. Screens in scope: S3 Agent Board · S4 Start Agent · S5/S6 folded tool-adaptive Detail. Mode: Flat (single-product). Topic: monitoring-start.

3. Confirmed Reframe (applies to ALL variations)

Three load-bearing decisions from the brief are carried by every variation:

  1. Two-layer surface, not one job. S3/S4 = orchestration (fleet-level: scan, fan out, launch fast); folded S5/S6 = deep-work (single agent). Every variation serves both layers; they contrast only in how the layers connect.
  2. Folded + tool-adaptive detail. S5 (Structured) and S6 (Terminal) collapse into one detail surface whose primary rendering is determined by the tool’s integration pattern, not a user choice: terminal-primary (xterm.js over WebSocket) for headless-CLI tools (Claude Code CLI, Codex CLI, Cursor worker, Aider, Continue, Cline, Copilot CLI, Augment), structured for API-native tools (Claude Agent SDK). The old “Mode A vs Mode B user toggle” framing is dropped. Terminal latency still matters for CLI tools (it is the live interface) but it is intrinsic, not an optional feature toggle.
  3. All web PWA. S3–S6 all render in the browser PWA; even the terminal is xterm.js in the browser over WebSocket. PWA installable on mobile + desktop.

Proposed flow-map amendment (NOT silently edited)

The approved flow-map (design/user-flow-personal-workstation.md) lists S5 and S6 as two separate screens and frames Mode A/B as a user-selectable access mode (decision D2). The confirmed reframe contradicts that on two points:

Per the brief, variations carry the merged tool-adaptive model, and this canonical plan proposes it back to the flow-map as an amendment pending the flow-map owner’s acceptance. The approved flow-map document is not edited by this plan. (IDE / Mode C / code-server remains out of scope unless a variation makes a deliberate case; none does.)

4. The Four Variant Axes

The five concepts deliberately spread four axes so nothing is pre-decided:

AxisWhat it variesSpread across the set
a) Board representationhow the fleet is showndense table (V1) · card grid (V2) · status columns (V3) · quiet list (V4) · vertical feed (V5)
b) Orchestrate↔deep-work connectionhow drill-in relates to the boardpersistent split-pane (V1) · full-page route (V2) · side drawer over board (V3) · focus mode + back-stack (V4) · full-screen sheet (V5)
c) Launch affordancehow a new agent is started (S4)inline top-of-table row + ⌘K (V1) · modal overlay (V2) · “+” add-card per column (V3) · command palette ⌘K (V4) · duplicate-from-template sheet (V5)
d) Device posture / densitydesktop-dense ↔ mobile-firstdesktop-dense (V1) · co-equal responsive (V2) · desktop reflow→stacked (V3) · keyboard-desktop / mobile bottom-bar (V4) · mobile-first (V5)

Device posture is a deliberate axis, not an accident. “Desktop-primary” was explicitly rejected as a global default: instead, the set spans from V1’s desktop-dense extreme to V5’s mobile-first extreme, with V2 co-equal in the middle, so the right posture is tested empirically across the comparison rather than assumed up front.

5. Variation Comparison Matrix

#VariationBoard representationConnection modelLaunch affordanceDevice postureComplexityBest-fit userPrimary tradeoff
V1Operator Consoledense sortable data tablepersistent split-pane (board L / detail R)inline top-of-table row + ⌘Kdesktop-densemediumkeyboard power user at max density (8–12 agents, wide monitor)intimidating / weak on mobile; one-line cell truncates rich stdout
V2Card Mosaic controlresponsive card grid w/ live stdoutboard → full-page detail routemodal overlayco-equal responsivelow (baseline)balanced default; cross-device; recognition-over-recallcards waste space at 12 agents; route-switch drops board context
V3Kanban Pipelinecolumns = lifecycle statusboard → side drawer (board stays behind)“+” add-card in Queued columndesktop reflow → stacked sectionsmedium-hightriage-first / kanban mental modelcard motion/instability; weak deep-work continuity; slower pure launch
V4Command-Firstquiet minimal listpalette nav; detail = focus mode + back-stackcommand palette (⌘K)keyboard-desktop / mobile bottom-barmediumkeyboard-native dev (vim/tmux/Raycast)low discoverability; weaker fleet glanceability; command-typing awkward on phone
V5Glanceable Streamprioritized vertical feedfeed → full-screen sheetduplicate-from-template quick-start sheetmobile-firstmedium“closed laptop, checked from phone” conductor; launches mostly known recipesweak dense-desktop multitasking; feed reorders; template-launch less flexible for novel prompts

Axis Explorer interactive preview

Pick a value on each of the four axes; the variation(s) that best match your selection highlight. This is a preview — an exploration aid over the same data as the comparison matrix, not production code. No data is stored. Reset restores the initial state. Requires JavaScript; with JS disabled, the comparison matrix above and the “View as table” fallback below carry the same data.

a) Board representation

b) Connection

c) Launch

d) Device posture

V1 Operator Console
V2 Card Mosaic
V3 Kanban Pipeline
V4 Command-First
V5 Glanceable Stream

§6 below embeds the complete text of each intermediate spec, unsummarized. Each variation’s internal heading hierarchy nests under its variation heading. Source files: design/ux-variations-monitoring-start/{v1-operator-console,v2-card-mosaic,v3-kanban-pipeline,v4-command-first,v5-glanceable-stream}.md.

V1 — Operator Console (v1-operator-console)

Branch: monitoring-start · Variation id: v1-operator-console · Run mode: default progression-mode (NOT layout-mode). Parent flow-map: design/user-flow-personal-workstation.md (Approved) · Shared brief: design/_working/ux-variations-monitoring-start-brief.md. Screens in scope: S3 Agent Board · S4 Start Agent · S5/S6 folded tool-adaptive Detail. Status: Draft spec (build-grade) — not yet UI-interviewed.

1. Name & Thesis

Operator Console. The agent fleet is a NOC (Network Operations Center — the always-on monitoring room where operators watch a wall of live system status and act the moment something turns red). One dense, sortable, filterable data table — one row per agent, 5–12 rows — with a persistent split-pane: the table holds the left column, the selected agent’s tool-adaptive detail holds the right. The fleet never leaves view. You triage like an SRE reading a monitoring dashboard: scan the status column, sort errors to the top, act inline (approve / unblock / stop) without leaving the row, and keep one agent open in the detail pane while the other eleven keep streaming beside it.

One-line thesis: Maximum-density, keyboard-driven fleet command — every agent on one screen, every action one keystroke away, drill-in without ever losing the fleet.

The design bet: a solo conductor running 12 agents does not want twelve cards to scroll through; they want a spreadsheet of their fleet and a terminal stapled to its side. Density is the feature, not a regression to apologize for.

2. Parent User Flow & Branch Relationship

This variation is one of five sibling treatments of the monitoring-start branch (Happy Path steps 3–6 + Path C “Parallel Sprint”). All five serve the same two-layer surface and contrast only in how the orchestration layer (S3/S4) connects to the deep-work layer (folded S5/S6).

idboard representationconnectionlaunchdevice posture
v1-operator-console (this)dense data tablepersistent split-paneinline top-of-table row + ⌘Kdesktop-dense
v2-card-mosaic (control)responsive card gridboard → full-page routemodal overlayco-equal responsive
v3-kanban-pipelinelifecycle status columnsboard → side drawer“+” add-card per columndesktop reflow → stacked
v4-command-firstquiet minimal listpalette nav + focus modecommand palette ⌘Kkeyboard-desktop / mobile bottom-bar
v5-glanceable-streamprioritized vertical feedfeed → full-screen sheetduplicate-from-template sheetmobile-first

Relationship to siblings: V1 sits at the explicit desktop-dense extreme of the device-posture axis (the brief’s deliberate counterweight to v5’s mobile-first). It shares v4’s keyboard-native instinct but diverges on glanceability: v4 hides the fleet behind a palette and shows a quiet list; V1 puts the entire fleet on screen at all times as structured tabular data. It is the antithesis of v2’s roomy cards. Mobile is V1’s deliberate degrade case, not its design center — that territory belongs to v5.

This spec carries the brief’s confirmed reframe (two-layer surface; folded tool-adaptive detail; all-web PWA) and does NOT silently edit the approved flow-map. Where it departs from flow-map structure, that is recorded in §11.

3. Target User Fit

Best fit: the keyboard-driven power user at peak load — the conductor running 8–12 concurrent agents across 2–3 repos from a wide desktop monitor, who lives in tmux/htop/k9s/Grafana and reads tables faster than cards. They want information per pixel, not whitespace; they sort and filter reflexively; they expect j/k to move and Enter to open.

Poor fit: a first-time user meeting agents for the first time (the table reads as intimidating cockpit chrome before there is anything to monitor); a phone-primary “quick-check” user (the table must collapse to survive small screens); a 1–3 agent light user (the density buys nothing — a card grid would serve them better, which is why v2 exists).

Fit signal: if the evaluator, during UAT, instinctively sorts the status column and uses keyboard nav to triage a 10-agent fleet faster than they did in v2/v5, this variation is earning its density.

4. Onboarding / Activation Model (first meeting + empty state)

V1 owns only the monitoring-start surface; first-run provisioning (S2 Setup Wizard, Path A) is upstream and out of scope. The activation question here is: how does a first-time user meet a NOC table with zero rows?

Empty state (zero agents): the table chrome is suppressed, not shown empty. A spreadsheet of nothing is hostile. Instead the split-pane collapses to a single centered column:

First-agent transition (activation moment): when the first agent launches, the empty column animates into the two-region split-pane. The first row appears with starting status and the detail pane auto-selects it (the only sensible selection), immediately showing live stdout. This is the “the console is now live” beat — the table header (sort/filter controls) fades in only once ≥1 row exists, so the user is never shown empty machinery.

Density ramp: columns are progressive. At 1–3 agents the table shows the full column set comfortably. The user is not asked to configure columns on day one; column customization is a power affordance discovered later (§14), never a setup gate.

5. Typical Workflow Sequence (orchestrate → drill-in loop)

The core loop. Numbered, keyboard-first, with mouse equivalents in parentheses.

  1. Open PWA → S3 console. Returning user lands on the populated table, default sort error > waiting > running > idle. The previously-selected agent is restored into the right pane (server-canonical; survives device switch).
  2. Scan the status column. Errors and approval-waits are sorted to the top and color-coded. A header chip shows counts: 2 waiting · 1 error · 7 running. Triage is a single visual sweep of the leftmost columns.
  3. Triage inline without drilling in. For a clear approval, press a on the focused row (or click the inline Approve affordance) — the row’s diff micro-preview expands one tier in place; confirm. For a blocked agent, u to unblock. For a runaway, x to stop (guarded, §15). The fleet stays put; no navigation.
  4. Drill in when a row needs attention. j/k (or arrow keys / click) to move the row cursor; Enter (or click the row) selects it into the right detail pane. The other 11 rows stay visible and keep streaming. This is the signature move: drill-in is selection, not navigation.
  5. Work the agent in the folded detail pane. The pane renders tool-adaptively (terminal-primary for CLI tools, structured for API-native — §13). Type into the terminal, read the structured step stream, inject a prompt, approve inline. The fleet table on the left keeps updating in your peripheral vision.
  6. Fan out a parallel sprint (Path C). To launch agent #2/#3 mid-flow: n (or click the inline + New Agent row pinned to the top of the table) — the top row becomes editable inputs (repo · tool · prompt). Fill and ⌘Enter to launch. New starting row drops into the table. Alternatively ⌘K → “New agent…” for a full quick-launch palette without leaving the keyboard.
  7. Round-robin approvals. As more agents hit decision points they sort to the top and the header count increments. Walk them: focus top waiting row → a → confirm → focus next. Or ⌘K → “Approve all clear” for the batch path (§15).
  8. Return to monitoring. Esc deselects the detail pane (collapses it wider on demand, §13) and returns focus to the table. Loop back to step 2.

The loop never leaves the single console route. Steps 2–3 are pure orchestration; steps 4–5 are deep-work; the split-pane is what fuses them.

6. Progression Model — and how it differs from the 4 siblings

How the user advances through the flow in V1: progression is lateral, not navigational. The user does not move between screens to advance; they move the selection cursor down the table and the detail pane updates underneath them. Advancement = re-sorting (state changes float rows up), selecting (binds a row to the pane), and acting inline (mutates state, which re-sorts). The flow’s S3→S5/S6 transition is collapsed into a pane-bind on one persistent route. There is no back-stack to manage because you never left the board.

Explicit contrast with each sibling:

Net: V1’s progression model is “sort + select + act-in-place on one persistent dense surface.” No other sibling keeps both the full fleet and a working agent on screen simultaneously — that simultaneity is V1’s defining progression property.

7. Sharing & Collaboration Model

N/A — single-user personal workstation. This is a solo dogfood environment; there is exactly one user (the Agent Conductor) and no second human to share with, hand off to, or collaborate alongside. There are no shared boards, comments, mentions, presence indicators, or co-editing. The only “handoff” is the same user across their own devices (cross-device continuity), which is server-canonical state, not collaboration — covered under §9 return-use. No sharing affordances are designed and none should be invented.

8. Permissions Model

N/A — single-user personal workstation. There are no roles, no per-row access control, no viewer/editor tiers, no team admin. The single user has full control of every agent in the table. The only authorization in the system is GitHub OAuth, which scopes what the agents may do to the user’s repos (pull/push/PR) — an upstream identity/repo-access concern owned by S1, not a board-level permissions model. The console exposes no permission UI and none should be invented. (The OAuth-expiry failure path — agents losing repo access mid-run — is a recovery concern, handled in §10, not a permissions model.)

9. Return-Use & Notification Model

Resume: on return, the console restores server-canonical state: same rows, same statuses, same sort, and the last-selected agent re-bound into the detail pane (its stdout buffer replayed on WebSocket attach). No reconnection ceremony — this is the flow-map’s core promise. If the user returns after a gap exceeding the recap threshold (flow-map D7), a dismissible recap strip docks above the table header — not a full-screen S8 takeover, because in the Operator Console the fleet itself is the recap. The strip summarizes “while you were away: 3 completed · 2 PRs · 1 error” with each item as a chip that, when clicked, selects the relevant row into the detail pane. Esc or “Got it” dismisses the strip and drops you onto the live table.

Approval surfacing within the board: approvals are first-class table citizens. A waiting agent sorts to the top with a waiting badge, its Approve / Reject affordances rendered inline in the row’s action cell, and a one-tier diff micro-preview expandable in place. The header carries a live approval-queue chip (2 waiting) that, when clicked or via ⌘K → “next approval”, jumps the cursor to the top waiting row. No approval ever requires leaving the console.

Error surfacing within the board: errors sort above everything, render a red error badge, and put the failure’s first line into the one-line stdout cell. The header error chip mirrors the count. Selecting the row binds the detail pane to the full error / stack trace with Retry / Restart / Open terminal actions (mapping flow-map S5 Error state).

Push notifications (PWA): web-push for approval-needed and error/completion events deep-links into the console with the relevant row pre-selected (not to a standalone approval screen — V1’s whole premise is that the board is the approval surface). On a desktop where the console is already open, the notification is redundant with the live header chips; on mobile it routes into the collapsed-list degrade view (§17) with the row expanded. Notification preferences themselves live in S10 Settings (out of scope here).

10. Failure Recovery Behavior (mapped to flow-map failure paths)

How the Operator Console surfaces and recovers each flow-map failure. The design principle: failures are table-and-pane states, surfaced where the user is already looking (header chips + row badges + detail pane), never as modal interrupts that hide the fleet.

Flow-map failureHow V1 surfaces itHow V1 recovers it
Backend unreachable (WS heartbeat >10s / health-check fail)A persistent top banner spans the full console width: “Reconnecting to workspace…” with a spinner. The table dims to ~70% opacity and renders last-known state with a “stale since HH:MM:SS” tag in the header. Inline mutate actions (approve/unblock/stop/launch) are disabled and visually greyed in every row.Auto-reconnect with exponential backoff. On reattach, banner clears, table un-dims, buffered WS deltas replay and re-sort the rows. Agents kept running on managed infra throughout — the dim signals “you’ve lost sight, not lost the fleet.”
Agent process crash (supervisor exit code ≠ 0)The agent’s row flips to red error badge, floats to the top via sort, and its one-line stdout cell shows the crash’s first line + a crashed tag. Header error chip increments.Supervisor auto-restarts transient crashes (≤3 retries) — the row shows a restarting (2/3) micro-state inline so the user sees the retry budget burn down. On persistent failure the row offers inline Retry / Restart; selecting the row binds the detail pane to the full crash output with Open terminal for manual intervention (flow-map Path D).
Network tunnel dropDistinct from backend-unreachable: a header pill reads “Tunnel reconnecting…” while the WebSocket itself may still be healthy. Table stays live but flagged.Tunnel auto-reconnects; pill clears. Agents unaffected (running on managed infra). Less severe than full backend loss → table is not dimmed, only flagged, to avoid crying wolf.
Model key / subscription expiry or rate-limitAffected agent’s row shows an amber paused badge with reason auth expired / rate-limited in the stdout cell. It does not sort as an error (it’s user-actionable config, not a crash) but sits above running.Inline “Update credential” affordance deep-links to S10 (Settings, out of scope). The agent’s session state is preserved; on credential update the row auto-resumes to running with no re-launch. Rate-limit case shows a cooldown countdown in the cell and self-resumes.
GitHub OAuth token expired / revoked (agent git op 401/403)Affected rows show amber paused — github auth badge and sort above running. Because all agents share the one OAuth identity, a header-level banner appears: “GitHub access expired — re-authorize to resume N agents.”One “Re-authorize GitHub” action in the banner re-runs OAuth; all paused-for-auth rows auto-resume on success. No per-agent action needed (single identity) — this is the one place a fleet-wide banner is warranted. No session state lost.
Git conflict during agent workThe agent pauses itself and sorts up as a waiting row (it is requesting approval, per flow-map). The stdout cell reads merge conflict — N files.Selecting the row binds the detail pane to the conflict detail. The user either resolves manually via the terminal in the folded detail pane (CLI tools) or injects an instruction (“resolve by taking theirs”) via the structured prompt box. Resumes inline.
Disk exhaustion (>90%)A header warning chip (disk 92%) appears before critical; at critical the backend pauses agents and their rows show paused — disk badges.Chip links to a cleanup action (old sessions / build artifacts). Pre-emptive, so the fleet rarely reaches forced-pause.
Browser / PWA crashNone visible to the running fleet (server-side). On reload the console rebuilds from server-canonical state.PWA service worker reloads to the console with current state and the prior selection restored. No data loss.
Concurrent device conflict (same agent acted on from two devices)Last-write-wins for approvals; the row’s action cell briefly flashes a “updated elsewhere” micro-toast so the user knows their stale click was superseded.Terminal sessions are separate tmux windows (independent). Console state is eventually consistent via WS — the table simply re-renders to truth.

Cross-cutting failure principle for V1: because the fleet is never hidden behind a route or modal, all of these failures remain visible in aggregate (header chips) and per-agent (row badges) at the same time. The split-pane means the user can be deep in one agent’s terminal and still see a second agent error appear at the top of the table — a property no route-switching sibling (v2) preserves.

11. Page & Flow Changes vs. the Flow-Map Baseline

The flow-map lists S3, S4, S5, S6 as four distinct screens/routes and frames Mode A (structured) vs Mode B (terminal) as a user-selectable access mode (flow-map D2). V1 carries the brief’s confirmed reframe and restructures as follows:

  1. S3 + S5 + S6 collapse onto ONE persistent route. There is no navigation from board to detail; the detail pane is a region of the console, bound by selection. The flow-map’s S3→S5/S6 “Downstream” handoff becomes an in-page pane-bind.
  2. S5/S6 fold into one tool-adaptive detail pane. The flow-map’s separate “S5 Structured” and “S6 Terminal” screens — and its Mode A/B user toggle (D2) — are dropped. The detail pane’s primary rendering is determined by the tool’s integration pattern, not a user choice: terminal-primary (xterm.js/WS) for headless-CLI tools, structured for API-native tools. There is no “switch to terminal” mode button as a primary affordance (an escape-hatch terminal toggle survives only for structured/API tools that also expose a shell, as a secondary control — §14).
  3. S4 changes from modal to inline. The flow-map’s centered modal-with-backdrop launch (S4) is replaced by an inline editable top-of-table row plus a ⌘K quick-launch palette. Launch happens in the table’s coordinate space, not over a dimmed board — preserving the never-lose-the-fleet principle even during launch.
  4. S8 Async Recap demotes from screen to strip. The full-screen recap becomes a dismissible docked strip above the table (§9), because the live table already is the recap surface in this layout.
  5. S7 Approval has no standalone screen. Approvals render inline in rows + in the detail pane; the flow-map’s dedicated S7 mobile-card is realized only in the mobile degrade view (§17), not on desktop.
  6. Default sort retained but extended. Flow-map’s error > waiting > running > idle is kept as the default, extended with starting/paused/stopped/done tiers and made user-resortable by any column (the flow-map sort was fixed; V1 makes it a table affordance).

These are variation-level proposals consistent with the brief’s amendment note; they do not edit the approved flow-map doc.

12. Navigation Model

Single-route, selection-driven, keyboard-first. The console is effectively a one-page app for the monitoring-start surface.

13. Screen-by-Screen Layout (regions & proportions)

All within one persistent route. Reference frame: desktop ≥1280px wide.

S3 — Console (the table) + persistent split-pane

┌──────────────────────────────────────────────────────────────────────────────┐
│ HEADER  gblock·console   [7 running ·2 waiting ·1 error]  [⌘K] [⟳tunnel] ⚙ │  ~52px
├──────────────────────────────────────────────────────────────────────────────┤
│ [filter: repo▾ tool▾ status▾]   sort: status▾            recap-strip(opt.) │  ~40px
├──────────────────────────────────────────┼────────────────────────────┤
│  + New Agent  (inline launch row)          │                              │
│ status·name·tool·repo·el·tok·act table       │   DETAIL PANE                │
│  ●E auth… aider api 12m 41k ⟳⏹          │   (folded, tool-adaptive)    │
│  ●W pay…  cc    web  4m 88k ✓✗         │   agent header / status      │
│  ●R index codex web  9m 120k ⏹  ◀───── binds selected row           │
│  ●R migr… sdk   api  2m 30k ⏹            │   PRIMARY RENDER:            │
│  ●I docs  cont… web 1h  9k ▶            │   terminal (xterm.js)        │
│  LEFT PANE ~55–62% width                     │   OR structured stream       │
│  table scrolls independently                 │   RIGHT ~38–45% width        │
│                                             │   footer: inject ▸ / stop    │
└──────────────────────────────────────────┼────────────────────────────┘

S4 — Start Agent (folded into the table, two ways in)

Folded S5/S6 Detail

Already described as the right pane above. Key point: there is no separate S5 vs S6 screen and no Mode A/B user toggle — one detail pane, rendering chosen by tool integration pattern. A secondary “Open full terminal” escape hatch exists for structured/API tools that happen to expose a shell, and a “Pop out” control can promote the pane to a focused full-width terminal for heavy interactive debugging (still the same route, the table temporarily yields width).

14. Key Components & Controls

15. Button & Link Behavior

All destructive or fleet-affecting actions are guarded or reversible; all resume actions are non-destructive and preserve session state.

16. Spatial Density, Sizing, Hierarchy

V1 is the explicit high-density end of the posture axis — density is the thesis, tuned, not careless.

17. Responsive Behavior at 3 Breakpoints

Mobile is V1’s deliberate degrade case (the design center is desktop-dense; v5 owns mobile-first).

18. Visual Tone

NOC / server-monitoring console. Dark, dense, instrument-panel calm-under-load. Monospace data, color-coded status as the dominant signal, tight gridlines, minimal chrome, no decorative illustration once the fleet is live (illustration appears only in the empty state). The aesthetic reference set: htop / k9s / btop / Grafana / a trading terminal — tools that pack maximum signal into a glance and reward a trained eye. Motion is restrained: rows re-sort with a quick settle, they don’t fly between regions (the deliberate anti-v3 choice); status badges may pulse subtly when waiting. The dark house palette is the baseline; status colors are the accent vocabulary. Confidence over friendliness — this screen should feel like a cockpit you’ve mastered, not a dashboard you’re being onboarded to.

19. Strengths

20. Risks & Failure Modes (of the design)

21. Implementation Complexity

Medium (matches the brief’s rating). Heavier than v2 (the card-grid control), lighter than v3 (kanban drag/animation/column state machine).

What drives the cost:

No multi-user, sharing, or permissions code (N/A) — that subtracts real complexity other product types would carry.

22. Prototype Scope (lightweight, evaluable build)

A lightweight prototype must include enough to let a solo evaluator run the §5 loop and feel the density tradeoff. Minimum:

  1. Populated fleet table with ~8–10 mock agents spanning every status (error / waiting / running / idle / paused / done / starting), default error > waiting > running > idle sort, at least sort-by-status and sort-by-elapsed working, and a status/repo filter.
  2. Live-ish one-line stdout per row (mock tail loop is fine) + status badges with color + text labels + status-tinted row edges.
  3. Keyboard navj/k/Enter row cursor + selection, Esc deselect, ⌘K palette (at least New-agent + Jump-to + Filter verbs).
  4. Persistent split-pane with a working selection bind: selecting a row renders the detail pane while the table stays visible and keeps updating. This simultaneity is the single most important thing to demonstrate — it’s V1’s whole thesis.
  5. Tool-adaptive detail — at least two mock agents: one CLI-type rendering a (mock or real) xterm.js terminal, one API-type rendering a structured step stream. Prove the auto-render-by-tool with no user toggle.
  6. Inline launch row producing a new starting row (mock launch) + ⌘K quick-launch equivalent.
  7. Inline actions — Approve/Reject on a waiting row (with one-tier diff peek), guarded Stop, Retry on an error row; action-cell clicks must not disturb the current pane selection.
  8. At least two failure surfaces rendered as table/pane states: backend-unreachable banner + table dim, and an agent-crash row with retry budget — to prove failures stay visible without hiding the fleet.
  9. Responsive proof — resize to ≤640px and show the table collapse to the stacked list + full-screen detail sheet (the honest degrade).

Mock WebSocket/stdout and mock approvals are acceptable; the evaluation target is the interaction model and density feel, not backend fidelity.

23. Readiness Signal — what makes this branch ready for /ui-interview

This branch is ready to advance to /ui-interview [v1-operator-console] when, during /uat --variant-evaluation against its siblings, the solo evaluator gives a clear signal such as:

If instead the evaluator wants V1’s split-pane fleet-visibility but v2’s card richness, or v5’s mobile posture, that is a consolidation signal (feed /consolidate-variations), not a ready-for-interview signal for V1 standalone. Readiness for /ui-interview specifically means: the Operator Console’s dense-desktop, split-pane, inline-action model is the chosen direction for the monitoring-start surface and is ready to be pinned down visually.

V2 — Card Mosaic (v2-card-mosaic) baseline / control

Parent flow branch: monitoring-start (“Board Monitoring & One-Click Start”). Run mode: default progression-mode UX variation (NOT layout-mode). Role in the set: BASELINE / CONTROL — closest to the approved flow-map’s low-fi wireframe notes; the other four variations are measured against this. Screens in scope: S3 Agent Board, S4 Start Agent, S5/S6 folded Agent Detail. Upstream: design/user-flow-personal-workstation.md, brief design/_working/ux-variations-monitoring-start-brief.md.

1. Name & Thesis

Card Mosaic. A familiar SaaS dashboard: the fleet is a responsive grid of rich agent cards, each a self-contained status tile with a live stdout preview. You scan the mosaic, launch from a modal, and drill into a single agent by navigating to a full-page detail route. The thesis is zero learning curve — anyone who has used Vercel, Railway, GitHub Actions, or a CI dashboard already knows how to read this board. It deliberately trades cleverness for legibility so it can serve as the control the other four concepts must beat.

Design center of gravity: glanceability per agent. Each card carries enough rich context (badge, tool, repo, elapsed, tokens, multi-line stdout) that the board answers “what is each agent doing right now?” without a drill-in. The cost — accepted on purpose — is fleet density: cards are large, so 12 agents scroll.

2. Parent User Flow & Branch Relationship

This variation realizes the monitoring-start branch: Happy Path steps 3–6 (S3 Board → S4 Start → S3 new card → S5/S6 Interact) plus Path C “Parallel Sprint” (launch agent #2, #3; monitor all; round-robin approvals).

It is the baseline/control. Where the brief’s four axes are spread across siblings, V2 takes the flow-map-native position on every axis:

AxisV2 positionFlow-map source
a) Board representationResponsive card grid w/ live stdoutS3 low-fi notes: “Responsive grid of agent cards (2–3 columns desktop, 1 column mobile)”
b) Orchestrate↔deep-work connectionBoard → full-page detail routeD3 “click existing → S5/S6” rendered as navigation
c) Launch affordanceModal overlayS4: “Access Mode A — modal overlay”; low-fi “Centered modal with backdrop”
d) Device posture / densityCo-equal responsive (balanced middle)S3 low-fi “2–3 columns desktop, 1 column mobile”

Because it tracks the flow-map most literally, V2 introduces the fewest net-new interaction concepts. The one deliberate divergence from the approved doc is the folded tool-adaptive detail (see §11) — a constraint carried by every variation per the brief’s confirmed reframe, not unique to V2.

3. Target User Fit

The balanced default Agent Conductor. Specifically:

Best-fit summary: the persona before they have strong opinions. V2 is what you reach for when you don’t yet know whether you’re a power-keyboard user (→ V1/V4), a triage-first thinker (→ V3), or a phone-up checker (→ V5). It degrades gracefully to 1-column mobile, so it is never wrong on any device, just never maximally optimized for one.

Weak fit: a max-density NOC operator babysitting 12 agents on a single 27" screen (cards waste vertical space — that’s V1’s job).

4. Onboarding / Activation Model

Entry into this surface is post-setup: Path A (onboarding) provisions the workspace and redirects to S3 with an empty board. V2 owns the empty → first-agent moment on the board.

Empty state (S3 Empty):

Activation = the first card appears. When the user completes the S4 modal, a card animates into the grid with a “starting” badge and flips to “running” within seconds, its stdout preview beginning to stream. That first live preview is the aha for this surface — the user sees the agent working without any SSH/tmux ceremony. The empty illustration is replaced by a real 1-card mosaic, which implicitly teaches the grid model.

Progressive disclosure: the modal’s “optional branch” field stays collapsed under a “More options” toggle on first run, so the first launch is the minimal repo + tool + prompt.

5. Typical Workflow Sequence

Parallel Sprint (Path C), the representative flow:

  1. Open PWA → lands on S3 Agent Board (returning session). Mosaic renders all current agents, sorted error > waiting > running > idle. Header shows approval-queue badge.
  2. Scan the mosaic. Each card’s badge + stdout preview answers “what’s happening” at a glance; no drill-in required for routine monitoring.
  3. Click Start Agent (header button) → S4 modal opens over the dimmed board.
  4. In the modal: pick repo (dropdown), pick tool (dropdown), type the prompt (textarea), optionally expand branch. Click Start Agent.
  5. Modal closes; a new “starting” card appears in the grid, transitions to “running,” stdout begins streaming. (Happy Path step 5.)
  6. Repeat 3–5 for agent #2 and #3 (different repos/branches). Board now shows 3+ live cards.
  7. An agent hits a decision point → its card flips to a waiting badge, pulsates subtly, and re-sorts toward the top; the header approval-queue count increments.
  8. Click the waiting card → navigate to the full-page detail route for that agent. The detail is folded + tool-adaptive: terminal-primary for CLI tools (xterm.js), or structured-primary for API-native tools. Handle the inline approval here.
  9. Click Back to Board → return to S3 (board re-reads server-canonical state; the card is now “running” again).
  10. Round-robin steps 7–9 across agents as approvals arrive.
  11. When agents finish, cards show a completion summary; a Review diffs CTA surfaces (handoff to S9, out of this branch’s deep scope but linked).

6. Progression Model — and How It Differs From the 4 Siblings

How the user advances in V2: by navigation and recognition. The orchestration layer (mosaic) and the deep-work layer (detail) are two distinct places. You progress by clicking a card to travel to its detail route, and you back out to return to the board. State is server-canonical, so the round trip is lossless — but it is a round trip: the board is replaced by the detail page, not held alongside it. New work enters via the modal, which is a temporary overlay on top of wherever you are.

This is the closest-to-baseline progression: a classic two-place dashboard → detail pattern with a modal for creation. Explicit contrast:

One-line positioning: every sibling deviates from the dashboard mental model on exactly one axis to test a hypothesis; V2 holds all four axes at the flow-map default so those deviations have something to be measured against.

7. Sharing & Collaboration Model

N/A — single-user personal workstation. Per the flow-map’s explicit non-goals (“Multi-tenant / team features (personal-workstation scope only)”) and the brief’s locked constraint, this product has exactly one user — the Agent Conductor who owns the managed VM and their own GitHub OAuth + model subscriptions. There are no shared boards, no invited collaborators, no co-viewing, no comments, and no handoff-to-another-person flow. The only “handoff” in scope is the same user moving between their own devices (cross-device continuity), which is server-canonical state, not collaboration. No sharing UI is designed.

8. Permissions Model

N/A — single-user personal workstation. There are no roles, no per-resource ACLs, no seat management, and no permission tiers, because there is exactly one principal. The only authorization in the system is external: GitHub OAuth grants the single user repo read/write + PR scopes so their agents can operate on their repos (flow-map S1), and the user’s own model keys/subscriptions authorize model traffic (BYO-client; the board never proxies). Those are credentials the user holds, not in-app permissions V2 grants to others. No permissions surface is designed; failures of those external credentials are handled as recovery paths (§10), not as a permissions UI.

9. Return-Use & Notification Model

V2 is a return-heavy surface (the Conductor checks in repeatedly across a day and across devices). Return mechanics:

10. Failure Recovery Behavior

Mapped to the flow-map Failure & Recovery table and the per-screen matrices:

Failure (flow-map)V2 board/detail behavior
Backend unreachable / WebSocket heartbeat timeoutS3 shows a top banner “Reconnecting to workspace…” with auto-retry (exponential backoff); cards freeze at last-known state; Start/Stop disabled until reconnect (S3 “Connection Lost” state). Detail route shows a “Reconnecting…” overlay; buffered output replays on reconnect (S5/S6 matrix).
Agent process crash (exit ≠ 0)The agent’s card flips to a red error badge with a one-line error summary and re-sorts to the top of the mosaic. Auto-restart for transient errors (≤3 retries) happens server-side; if it stays failed, the card offers Retry and Open (→ detail with full error output / stack trace, and a drop-to-terminal path for CLI tools). Maps to Path D (recovery).
Network tunnel dropSame banner pattern (“Tunnel reconnecting…”); cards unaffected in content (agents keep running on managed infra), just marked stale until the tunnel returns.
Model key / subscription expiry or rate limitAffected agent pauses; card badge → waiting/blocked with a “credential” reason; an inline link routes to S10 Settings to update the credential; agent resumes post-update without losing session state.
GitHub OAuth token expired/revokedAffected agents pause; a prompt to re-auth via GitHub OAuth appears (banner + card state); agents resume automatically once a new token is issued.
Git conflict during agent workAgent pauses and raises it as an approval with conflict details → handled in the folded detail (terminal for CLI tools so the user can resolve manually, or instruct the agent to resolve).
Disk exhaustion (>90%)Pre-critical warning banner on S3 with a cleanup suggestion; emergency pause of agents to prevent corruption, reflected as paused card states.
Browser/PWA crashService worker detects unclean shutdown; PWA reloads to S3 at current server state — no data loss (agents are server-side).
Concurrent device conflict (same agent, two devices)Last-write-wins for approvals; the mosaic on each device reconciles via WebSocket to eventual consistency; terminal sessions are separate tmux windows so they don’t collide.

S4 modal failures (launch fails / tool unavailable): inline error in the modal with a Retry; if the selected tool isn’t installed in the workspace, a warning with an Install tool link to S10 and Start disabled until a valid tool is chosen (S4 matrix).

Recovery design principle for V2: failures are expressed on the card (badge + summary + re-sort to top) so the board itself is the recovery queue — no separate error console. Deep remediation happens by clicking through to the folded detail.

11. Page & Flow Changes vs. the Flow-Map Baseline

V2 is intentionally the minimal-divergence variation. Tracking:

Net: two screens (S3, S4) match the baseline; the detail folds two baseline screens into one tool-adaptive route. This should propagate to the flow-map as an amendment, not a silent edit (per the brief).

12. Navigation Model

Place-based, route-driven — the canonical SaaS-dashboard nav:

Routing surface (illustrative): / board · /agents/:id folded detail · S4 and S8 as overlays/interstitials over whatever route is active · /settings (S10).

13. Screen-by-Screen Layout (Regions & Proportions)

S3 — Agent Board (mosaic)

S4 — Start Agent (modal overlay)

S5/S6 — Folded Agent Detail (full-page route)

Proportions summary: board = full-bleed grid; modal = centered ~520px card; detail = top bar + dominant primary region + a slim right rail.

14. Key Components & Controls

Sort/filter: default sort error > waiting > running > idle. Filtering is deliberately light in V2 (a simple status filter / repo filter chip row is the most it should add) — heavy filtering is V1’s territory; V2 leans on big legible cards instead.

15. Button & Link Behavior

All destructive/irreversible actions (stop, reject) get a confirmation affordance; launching and approving do not (they’re the happy path and reversible enough).

16. Spatial Density, Sizing, Hierarchy

17. Responsive Behavior at 3 Breakpoints

Mobile ≤ 640px (1 column):

Tablet ≤ 1024px (2 columns):

Desktop > 1024px (3 columns):

Co-equal posture means no breakpoint is the “real” one — the same components reflow; the phone gets a usable read-and-approve experience, the desktop gets comfortable multi-agent monitoring, neither is the privileged target. (Contrast: V1 privileges desktop, V5 privileges phone.)

18. Visual Tone

Familiar, trustworthy, boringly competent — on purpose. It should read like a polished deploy/CI dashboard (Vercel/Railway/Render lineage) so a new user feels instant fluency. Dark house palette; calm grayscale chrome with status color as the only loud element; soft card elevation; restrained motion (cards fade in on launch, waiting cards pulse gently, stdout scrolls smoothly — no gratuitous animation). The terminal in detail reads as a real console (monospace, darker bg). The overall message: “this is the safe, legible way to watch your fleet” — which is exactly the control’s job.

19. Strengths

20. Risks & Failure Modes

21. Implementation Complexity

Lowest of the five (baseline). Rough estimate: low.

Why it’s the floor:

Largest single risk to the estimate is external, not structural: stdout-parsing reliability and terminal latency (both shared open questions), not V2’s own UI.

22. Prototype Scope

For the serial build → /uat --variant-evaluation comparison, prototype:

Out of prototype scope (linked, not built here): full S9 diff review, S8 recap timeline beyond a stub, S10 settings, real push delivery, real tunnel/auth, and Mode C/IDE (excluded).

23. Readiness Signal for /ui-interview

This spec is ready to hand to /ui-interview [v2-card-mosaic] once V2 is selected or built. It fixes, for the visual-mockup stage:

Open items the UI interview should resolve (not blockers, just the live questions): exact card dimensions / preview line-count, the precise side-rail width and collapse behavior in detail, the light-filtering affordance (status/repo chips) if any, and mobile terminal ergonomics copy (“open on desktop” handoff). Bounding open questions (terminal latency, stdout-parsing reliability) are flagged as risks, not resolved here, and carry forward as evaluation criteria.

V3 — Kanban Pipeline (v3-kanban-pipeline)

Variation id: v3-kanban-pipeline · Parent flow branch: monitoring-start (Happy Path steps 3–6 + Path C “Parallel Sprint”) · Run mode: default progression-mode (NOT --layout-mode). Upstream (approved): design/user-flow-personal-workstation.md, manifest design/flow-tree-personal-workstation.yaml. Screens in scope: S3 Agent Board, S4 Start Agent, S5/S6 merged folded + tool-adaptive detail (one drawer, no Mode A/B user toggle).

1. Name & Thesis

Kanban Pipeline. The fleet is a status pipeline, not a list of things. Each agent is a card that lives in a lifecycle column — Queued → Running → Waiting-Approval → Error → Done — and moves between columns as its status changes. The conductor’s primary job here is triage, and triage is made spatial: a tall Waiting-Approval or Error column is an instantly visible bottleneck, readable in under a second from across the room or a glance at a phone. You don’t read rows and parse status badges; you read the shape of the board. Drill-in is a side drawer that slides over the board so the pipeline stays partially visible behind it — orchestration context is never fully lost when you go deep on one agent.

The wager: for a conductor running 5–12 concurrent agents, the most valuable on-screen signal is “where is the queue backing up?” and a status-keyed board answers that question with column height before the user reads a single word.

2. Parent User Flow & Branch Relationship

This variation is one of five sibling treatments of the monitoring-start branch. It covers:

Adjacent branches it touches but does not own: approval (S7) — here surfaced as the Waiting-Approval column + drawer-inline approval, not the full approval-routing experience; recovery (Path D) — surfaced as the Error column; device-handoff (Paths E/F) — surfaced as the persistent QR / “Open on Mobile” header control. Those branches get their own dedicated variation work; this spec only renders their entry points as columns/affordances on the monitoring-start surface.

Flow-map amendment carried (per brief): S5 + S6 are merged into one folded, tool-adaptive detail surface (the drawer). This contradicts the approved flow-map’s separate-screens + Mode A/B framing; the brief authorizes carrying the merged model in variations as a proposed amendment. The approved doc is not edited here.

3. Target User Fit

Best fit: the triage-first conductor — the user whose dominant loop is “scan the fleet, find what’s blocked, clear it, repeat.” This person runs 5–12 agents and experiences the board mostly as a queue-management problem: their pain is not “I can’t find agent X” but “which of my twelve agents need me right now and in what order?” The status-keyed board is purpose-built for that: it sorts the entire fleet into “needs-you” columns (Waiting-Approval, Error) vs. “leave-it-alone” columns (Running, Done) structurally, before any per-card reading.

Also fits: users with a project-management / kanban mental model (Trello, Linear board view, GitHub Projects) who already think in columns-as-status and will find the board immediately legible.

Poor fit: the deep-work-dominant user who spends most of their time inside ONE agent’s terminal for long stretches — the drawer is a good drill-in but the board’s motion behind it (cards moving columns) is a low-grade distraction for someone who wants the fleet to disappear. That user is better served by V4 Command-First (focus mode) or V1 Operator Console (persistent split-pane). Also poor for the pure-launch-velocity user who just wants to fire agents fast — the add-card is column-scoped and slightly slower than a global ⌘K (V1/V4) or a one-tap recipe (V5).

4. Onboarding / Activation Model (Empty Columns State)

This variation does not own first-time setup (that’s the onboarding branch, Path A → S2/S10). It owns the empty board a fresh-but-provisioned workspace lands on.

Empty-columns state. On a provisioned workspace with zero agents, the board renders all five column headers in their resting positions (Queued, Running, Waiting-Approval, Error, Done), each empty but structurally present — this teaches the lifecycle model at a glance: “agents are born in Queued and flow rightward.” The columns are not hidden; their visible emptiness is the onboarding.

Activation moment for this surface: the user clicks the Queued add-card, fills the inline launch form (§13/S4), and watches the new card appear in Queued and animate one column right into Running. That single observed transition — card moves on its own as status changes — is the “aha” specific to this variation: the board is alive and self-sorting. Time-to-first-agent target: under 30 seconds from empty board.

5. Typical Workflow Sequence

  1. Land on the board (S3). Returning user opens the PWA → routed to the column board. Instant read: scan column heights. Tall Waiting-Approval or Error = work to do; everything in Running/Done = relax.
  2. Triage by column, not by card. Eye goes first to Error (leftmost of the “needs-you” pair, highest priority), then Waiting-Approval. Within a column, cards are ordered oldest-blocked-first (stable; see §6/§20).
  3. Clear a blocker. Click the top card in Waiting-Approval → drawer slides in over the board (board dims but stays visible behind). Drawer shows the folded detail; the inline approval card is right there. Approve → card animates out of Waiting-Approval and back into Running (or into Done if the approval was the final step).
  4. Work the column down. Approve / unblock the next card. Because ordering is stable, “the next one” stays put — you’re not chasing a reflowing list.
  5. Launch new work (S4). Click “+” at the head of the Queued column → compact inline launch form expands in place (repo / tool / prompt). Submit → card lands in Queued → auto-advances to Running on spawn. Repeat for parallel sprint (#2, #3).
  6. Go deep on one agent (folded S5/S6). Click any Running card → drawer opens to the tool-adaptive detail: terminal-primary (xterm.js) for CLI tools, structured stream for API-native tools. The opened card is pinned in its column (§20) so it does not slide away under you while you work.
  7. Escalate / intervene. From the drawer, inject a prompt, stop/restart, or — for CLI tools — drop straight into the live terminal that the drawer already hosts. For an errored agent, the drawer opens with the error output + Retry/Restart/Terminal controls.
  8. Resolve & dismiss. Close the drawer (Esc, backdrop click, or X) → board returns to full view; any status changes that occurred while the drawer was open are reflected (the closed card may now be in a different column).
  9. Review completed work. Cards in Done carry a “Review diffs” affordance → routes to S9 (diff review, owned by the diff-review branch). Done is the staging area for shipping.
  10. Repeat. The loop is scan-shape → triage-column → clear → relaunch.

6. Progression Model — and Explicit Contrast vs. the 4 Siblings

How the user advances in V3: progression is the card’s literal motion through columns. “Advancing” a task = moving its card rightward (Queued → Running → Done) or unblocking it (Waiting-Approval/Error → Running). The user’s mental model of “where am I” is answered by which column the card is in. The system moves cards automatically on status change (server-canonical via WebSocket); the user moves them indirectly by acting on blockers. There is no manual drag in the default model — columns are status-derived, not user-assigned buckets (dragging a card to “Done” would lie about agent state). The progression metaphor is a pipeline you unblock, not a board you arrange.

Explicit contrast with siblings:

One-line summary of the axis V3 owns: progression is spatial and automatic (card motion across status columns), drill-in preserves the board (drawer), launch is column-anchored (+ in Queued).

7. Sharing & Collaboration Model

N/A — single-user personal workstation. This is a solo dogfood environment; there is no second user to share a board, column, card, or drawer with. No shared boards, no card assignment to teammates, no comments, no presence/cursors, no collaborative columns. (Kanban tools often imply assignment-to-people; that semantic is explicitly absent here — columns are keyed to agent lifecycle status, never to a person or owner.) Cross-device access by the same user is covered by the device-handoff branch (server-canonical state + QR), not by any sharing mechanism.

8. Permissions Model

N/A — single-user personal workstation. No roles, no per-column or per-card permissions, no multi-tenant access control, no viewer/editor distinction. The only identity/authorization in the system is the user’s own GitHub OAuth (identity + repo scopes so agents can pull/push/PR) and the user’s own model keys/subscriptions (BYO, never touched by the board) — both owned by the onboarding/settings branches, not surfaced as a permissions layer on this board. Every agent and every column belongs to the one user.

9. Return-Use & Notification Model

The two “needs-you” columns are the persistent triage queue — this is V3’s core return-use strength.

10. Failure Recovery Behavior

Mapped to the flow-map Failure & Recovery Paths table and Path D (recovery). The Error column is the spatial home of failure on this surface.

Flow-map failureHow V3 surfaces & recovers
Agent process crash (exit ≠ 0)Card animates into the Error column with a red badge + one-line error summary. Auto-restart for transient errors (≤3 retries) happens silently in place in Running with a small “retrying…” sub-state; only a persistent failure promotes the card to Error. Clicking the Error card opens the drawer to error output + stack trace + Retry / Restart / Open Terminal controls (folded S5 “Error” + S6 escalation). On successful restart the card animates back to Running.
Backend unreachable / WebSocket dropTop-of-board “Reconnecting to workspace…” banner with auto-retry (exponential backoff). Columns freeze at last-known state and are visually muted; launch (+) and card actions disabled until reconnect. Cards do NOT move while disconnected (no false transitions). Agents continue on managed infra.
Network tunnel drop“Tunnel reconnecting…” banner; columns dimmed but readable; no card motion. Agents unaffected.
Model key/subscription expiry or rate-limitAffected agent pauses → card enters Waiting-Approval (or a tinted sub-state of it) with a “credential needs attention” reason; drawer offers a deep link to S10 settings to update the credential, then resumes the agent in place without losing state.
GitHub OAuth token expired/revokedAffected agent’s card → Error (or Waiting-Approval with re-auth reason); drawer prompts re-auth via GitHub OAuth; on new token the card auto-returns to Running, no state lost.
Git conflict during agent workAgent pauses and requests approval with conflict details → card → Waiting-Approval; drawer shows the conflict; user resolves via the drawer’s terminal (CLI tools) or instructs the agent to resolve.
Disk exhaustionPre-critical warning banner across the board; emergency pause moves affected cards to Waiting-Approval with a cleanup prompt.
Browser/PWA crashService worker reloads to the board at current server state; columns re-render correctly (server-canonical). No data loss.
Concurrent device conflictLast-write-wins for approvals; the column position is server-derived so both devices converge to the same card placement via WebSocket.
S4 launch failure (tool not found / git error)The inline add-card form (still expanded in Queued) shows the error inline with Retry; no orphan card is created in Queued until spawn succeeds.

Error column’s distinct role: it is the only column that demands action AND implies something went wrong (vs. Waiting-Approval, where the agent is healthy and just wants a decision). Keeping Error as a separate, leftmost-priority column means a single glance distinguishes “broken” from “blocked,” which the flat-table (V1) and single-feed (V5) treatments blur together under a generic “needs attention” sort.

11. Page & Flow Changes vs. the Flow-Map Baseline

Flow-map baselineV3 changeRationale
S3 = responsive grid of cards, sorted error > waiting > running > idleS3 = status columns; the sort becomes spatial separation (one column per status). The known-good sort order is honored as column left-to-right priority (Error and Waiting-Approval placed to grab the eye).Triage-first thesis: make the sort a layout, not an ordering.
S4 = centered modal launch with backdropS4 = inline “+” add-card at the head of the Queued column; form expands in place, no backdrop, board stays visible.Launch is anchored where the new card will appear; preserves board context.
S5 + S6 = two separate screens; Mode A/B is a user toggleMerged into ONE folded, tool-adaptive drawer that slides over the board: terminal-primary for CLI tools, structured for API-native. No Mode A/B user choice.Carries the brief’s mandated reframe (folded + tool-adaptive).
Drill-in implied as a route/screen changeDrill-in is a drawer overlay; the board is never left, only dimmed behind.“Orchestrate↔deep-work connection” axis V3 owns.
Empty state = single centered illustration + CTAEmpty state = all columns visible & captioned, CTA in the Queued column.The empty board teaches the lifecycle model.

No screens are added or removed; S7/S8/S9/S10 are reached via the same handoffs (now from columns/cards/drawer instead of grid cards/modal). The flow-map document itself is not edited (amendment proposed, per brief).

12. Navigation Model

13. Screen-by-Screen Layout (Regions & Proportions)

S3 — Column Board (the home surface)

S4 — Inline Add-Card (launch)

Drawer — Folded S5/S6 Detail (tool-adaptive)

14. Key Components & Controls

15. Button & Link Behavior

16. Spatial Density, Sizing, Hierarchy

17. Responsive Behavior (3 Breakpoints) — Column Reflow Strategy

The defining responsive move: horizontal columns reflow to vertically stacked status sections.

Reflow rationale: the status-section stack on mobile preserves the exact triage semantics (priority order, count-per-status, expand-the-blockers) that the desktop columns provide — the mental model survives the reflow; only the axis (horizontal → vertical) changes.

18. Visual Tone

Dark, calm-operational, shape-led. Honors the house dark palette (--bg:#1a1a2e; --card:#16213e; --border:#2a2a4a; --accent:#4dabf7; --text:#e0e0e0; --muted:#8b8fa3). Status color is the one place chroma is spent and it’s spent meaningfully:

Motion is purposeful and restrained — card transitions between columns are short, eased, and rate-limited (§20) so the board reads as a settling pipeline, not a churning one. Typography compact and monospace-leaning in the terminal drawer; sans elsewhere. Overall feel: a quiet ops board where the only loud thing is a column that needs you.

19. Strengths

20. Risks & Failure Modes (with Mitigations)

The brief flags V3’s inherent risks explicitly. Addressing each:

21. Implementation Complexity

Estimate: medium-high — among the higher of the five (above V2 baseline, V4, V5; roughly peer to or slightly above V1).

Drivers:

Lower-cost relative to V1: no dense virtualized table; relative to V5: less novel gesture/swipe surface. The net premium over the V2 control is squarely in the column motion-control + reflow machinery.

22. Prototype Scope

For the build-grade prototype (feeds /uat --variant-evaluation and /ui-interview):

23. Readiness Signal for /ui-interview

This spec is ready to feed /ui-interview [v3-kanban-pipeline]. It locks the load-bearing decisions: status-column board representation, side-drawer drill-in with the board preserved, inline “+” add-card launch, desktop-columns→stacked-sections reflow, and the folded tool-adaptive detail (no Mode A/B toggle). It maps all four monitoring-start screens (S3/S4/S5/S6) and their state matrices, addresses the inherent card-motion risk with concrete mitigations (pin-open-card, stable ordering, rate-limited/coalesced transitions), and ships N/A statements for sharing/collaboration/permissions per single-user scope.

Open items deliberately deferred to /ui-interview (visual decisions, not flow decisions): exact column proportions and the empty-column collapse threshold; final status chroma values within the dark palette; drawer width ratios and the maximize-toggle breakpoint; card content density (which secondary line to show per status); the precise tablet behavior (5 narrow columns vs. 2.5-col rail) — to be settled empirically. None of these block the visual mockup; they are its agenda.

Bounding open questions carried forward (not resolved here): terminal latency (drawer terminal interactivity vs. mobile read-mostly handoff), stdout-parsing reliability (mitigated by coarse-lifecycle-driven column placement), and code-server/IDE (held out of scope — not argued in).

V4 — Command-First (v4-command-first)

Parent flow branch: monitoring-start (Happy Path steps 3–6 + Path C “Parallel Sprint”) · Upstream flow-map: design/user-flow-personal-workstation.md (Approved) · Shared context: design/_working/ux-variations-monitoring-start-brief.md. Screens in scope: S3 Agent Board, S4 Start Agent, S5 Agent Detail, S6 Agent Terminal. Run mode: default progression-mode (NOT --layout-mode).

1. Name & Thesis

Command-First. The fleet is a quiet, low-chrome list; the verb is a command palette (⌘K). Every action — launch, filter, jump, approve, escalate to terminal — is a command, not a hunted-for button. Opening an agent enters a full-attention focus mode with the board held as a back-stack (Esc / breadcrumb returns). The mental model is a terminal multiplexer or an editor command palette, ported to the web: a keyboard-native developer who resents pointing-and-clicking through UI gets a surface where typing is the navigation. Chrome stays out of the way until summoned; the orchestration layer (S3/S4) and the deep-work layer (S5/S6) are stitched together by one consistent typed grammar rather than by spatial layout.

2. Parent User Flow & Branch Relationship

This variation is one of five sibling expressions of the monitoring-start branch. It serves the two-layer surface the brief mandates:

Per the brief’s confirmed reframe, detail is folded + tool-adaptive: focus mode renders terminal-primary (xterm.js over WebSocket) for CLI tools and structured for API-native tools. There is no Mode A/B user toggle — rendering is determined by the agent’s tool integration pattern. The command-first ethos pairs naturally with CLI terminals (a keyboard user living in a terminal is the design’s home turf), and we lean into that, but the rendering choice remains tool-determined, not user-selected.

Upstream screens S7 (approval), S8 (recap), S9 (diff) are adjacent and reachable but owned by sibling branches; this spec treats them as palette/focus-mode destinations only where monitoring-start touches them (inline approval, error triage).

3. Target User Fit

The keyboard-native Agent Conductor: a solo developer running 5–12 concurrent agents across 1–3 repos who already lives in vim/tmux/VS Code-command-palette/Raycast and instinctively reaches for ⌘K before reaching for the mouse. This user:

Poor fit: a user who wants to see the whole fleet at a glance with rich live previews (better served by V2 Card Mosaic or V5 Glanceable Stream), or who triages spatially by status column (V3 Kanban). Command-First trades glanceability for velocity.

4. Onboarding / Activation Model

The risk of a command-first UI is invisible capability. We confront it at activation:

After first launch, the empty state is replaced by the quiet list (§13).

5. Typical Workflow Sequence

  1. Land on board (returning session) → quiet list of agents, sorted error > waiting > running > idle. Hint bar visible at bottom.
  2. Scan the list — status dots + one-line status give a fast read of who needs attention.
  3. ⌘K → start agent → palette walks a typed/guided launch: start agent on <repo> with "<prompt>" (tool inferred or appended using <tool>). Enter fires the launch.
  4. New agent appears in the list with a starting dot, flips to running within seconds (server-canonical via WebSocket).
  5. Fan out (Path C): ⌘K again → start agent on <repo-2>…, again → <repo-3>…. Each is a few keystrokes; no modal context-switch friction.
  6. Triage: ⌘K → filter status:waiting (or just / to filter the list inline) → list collapses to agents needing approval.
  7. Jump in: ⌘K → jump to <agent> (or arrow + Enter on the list) → focus mode opens; board pushed onto back-stack.
  8. Deep work: focus mode renders terminal-primary (CLI tool) or structured (API tool). Type into the terminal / inject a prompt / approve inline.
  9. Approve from focus mode: ⌘K → approve (acts on the focused agent) or click the inline approval card.
  10. Return: Esc (or click breadcrumb) pops focus mode, restoring the exact board scroll/filter state from the back-stack.
  11. Repeat triage → focus → return across the fleet. Round-robin approvals (Path C step 5) are ⌘K → approve next walking the waiting queue.

6. Progression Model — and Explicit Contrast vs. Siblings

How the user advances in V4: progression is verb-driven, not spatial. The user moves forward by issuing the next command, and moves deeper by jumping into focus mode (board → back-stack). The palette is the single throughline that unifies orchestration and deep-work: the same ⌘K surface starts an agent, filters the fleet, jumps to an agent, and acts on the focused agent. “Advancing” feels like a REPL — type, execute, observe, type again.

Explicit contrast with the other four siblings (per brief):

The shared progression invariant (per brief): all five serve both the orchestration and deep-work layers and use the folded tool-adaptive detail. V4’s distinctive claim is that the connection between layers is a typed command surface, not a spatial relationship.

7. Sharing & Collaboration Model

N/A — single-user personal workstation. Per the flow-map’s explicit non-goals and the brief’s single-user scope, there is no sharing, no shared boards, no co-presence, no comment/mention surface. The palette grammar deliberately contains no share/invite/assign verbs. (If multi-tenant ever arrives it is a separate product surface, out of scope here.)

8. Permissions Model

N/A — single-user personal workstation. There are no roles, no per-agent ACLs, no collaborator permissions. The only auth boundary is the single user’s GitHub OAuth session (identity + repo scopes), which is a session/connectivity concern owned upstream (S1), not a UX permission model on this surface. No permission UI appears in the board, palette, or focus mode.

9. Return-Use & Notification Model

The challenge: a quiet list must still make approvals and errors impossible to miss without becoming noisy. Mechanisms:

10. Failure Recovery Behavior (mapped to flow-map failure paths)

Each flow-map failure path gets a command-first expression:

Re-plan trigger (per CLAUDE.md): if a failure recovery would require new structural work (not a clear command path above), the implementation stops and re-plans rather than improvising.

11. Page & Flow Changes vs. Flow-Map Baseline

No upstream screen is removed; S7/S8/S9/S10 remain reachable as palette destinations. The flow-map’s tool-adaptive folding is carried as an amendment, not a silent edit to the approved doc.

12. Navigation Model (palette + back-stack + focus mode)

Three coordinated primitives:

  1. Command palette (⌘K): the universal verb surface. Available everywhere (board and focus mode). It is context-aware: on the board it ranks fleet/launch verbs; in focus mode it ranks single-agent verbs (approve, retry, inject, stop, open terminal) and implicitly targets the focused agent.
  2. Focus mode: opening an agent (jump to, Enter on a row, or push deep-link) covers the board with a full-attention single-agent view. This is the folded S5/S6 detail.
  3. Back-stack: the board is pushed onto a navigation stack when focus mode opens. Esc (or clicking the breadcrumb Board / <agent>) pops back to the exact prior board state (scroll position, active filter, selection). Nested jumps (agent → terminal sub-view, if any) stack and pop in order. The back-stack is what makes deep-work-then-return feel instant and lossless — the explicit contrast with V2’s context-dropping route.

Inline filter (/): a lightweight, palette-adjacent affordance — pressing / on the board focuses an inline filter on the list (fuzzy match on name/repo/status) without opening the full palette, for fast narrowing. filter status:error via ⌘K does the same with explicit grammar.

Breadcrumb: focus mode shows a single breadcrumb (Board / agent-name) — both an orientation cue and a clickable back affordance (the fallback for users who don’t reach for Esc).

13. Screen-by-Screen Layout (regions & proportions)

S3 — Quiet Minimal List (board)

S4 — Palette Launch (overlay on board)

S5/S6 — Focus Mode (folded, tool-adaptive detail)

14. Key Components & Controls

Command palette (the centerpiece). Grammar is verb-first, fuzzy-tolerant, and forgiving (partial matches resolve to ranked suggestions). Core grammar with examples:

VerbGrammarExample
Launchstart agent on <repo> with "<prompt>" [using <tool>]start agent on web-app with "fix the failing CI tests"
Jumpjump to <agent> / open <agent>jump to auth-refactor
Filterfilter status:<state> / filter repo:<name>filter status:error
Approveapprove [<agent>] / approve next / reject [<agent>]approve next
Controlstop <agent> · restart <agent> · retry [<agent>] · resume <agent>restart api-worker
Injectinstruct <agent> "<prompt>" / inject "<prompt>" (focused)inject "use pnpm not npm"
Escalateopen terminal <agent> (CLI)open terminal aider-2
Systemopen settings · open on mobile (QR) · show recap · re-authenticate github · ? helpopen on mobile

Palette behaviors: discovery mode (grouped suggestions when empty), context-aware ranking (pending approvals/errors surface first), inline errors with command pre-fill, segment-based guided launch, recent-command history (Up arrow recalls last commands).

15. Button & Link Behavior (and keyboard shortcuts)

The design is keyboard-first but every command has a fallback clickable affordance (anti-discoverability-trap principle):

Every keyboard action has a visible, pointer-accessible twin so the UI is never only discoverable by memory.

16. Spatial Density, Sizing, Hierarchy

17. Responsive Behavior at 3 Breakpoints

The device-posture axis for this variation is deliberately keyboard-desktop-first with a mobile bottom-bar fallback — distinct from V5’s mobile-first and V1’s desktop-dense-table.

18. Visual Tone

Terminal-native, quiet, confident, low-noise. Dark by default (house palette: --bg:#1a1a2e; --card:#16213e; --border:#2a2a4a; --accent:#4dabf7; --text:#e0e0e0; --muted:#8b8fa3). The aesthetic evokes a well-tuned tmux/Raycast/editor-command-palette: monospace accents, hairline separators, status as single colored dots, the palette as the one moment of focused brightness. Motion is near-absent except the subtle pulse on waiting rows. The feeling is fast and unobtrusive — the UI defers to the work.

19. Strengths

20. Risks & Failure Modes (with mitigations)

21. Implementation Complexity

Medium (matches the brief’s rating). Drivers:

22. Prototype Scope

For the variant-evaluation build (static-leaning HTML/CSS + light JS, dark house palette):

Out of prototype scope: real WebSocket backend, real tool processes, push delivery, S8/S9/S10 full screens (link/stub only), QR generation.

23. Readiness Signal for /ui-interview

This spec is build-grade and ready for prototype + /ui-interview on the v4-command-first branch. It fully specifies: the two-layer surface (S3/S4 orchestration + folded tool-adaptive S5/S6 deep-work), the palette/back-stack/focus-mode navigation spine, the command grammar with concrete examples, the discoverability mitigations (hint bar, discovery mode, help, clickable twins), the quiet-list notification model, failure-path mappings to the flow-map, three responsive breakpoints with the mobile bottom command bar, density/hierarchy/visual tone, and explicit contrast against all four siblings. Single-user scope (sharing/permissions N/A) and the bounding open questions (terminal latency, stdout-parsing reliability, IDE out-of-scope) are honored, not assumed away. Recommended next step per the brief: build serially → /uat --variant-evaluation/consolidate-variations, or /ui-interview v4-command-first if this branch is taken to visual mockup directly.

V5 — Glanceable Stream (v5-glanceable-stream)

Parent flow branch: monitoring-start (“Board Monitoring & One-Click Start”) · Variation id: v5-glanceable-stream · Run mode: default progression-mode (NOT --layout-mode). Screens in scope: S3 Agent Board, S4 Start Agent, S5 Agent Detail, S6 Agent Terminal (S5/S6 folded — see §13). Upstream (approved): design/user-flow-personal-workstation.md · Shared brief: design/_working/ux-variations-monitoring-start-brief.md.

1. Name & Thesis

Glanceable Stream. The fleet is a single prioritized vertical feed — the thing that needs you is at the top, everything calm sinks below the fold. The product is designed phone-up: the canonical posture is the conductor standing in line for coffee, thumb on the screen, glancing at “what needs me right now.” Approvals and errors pin to the top; waiting, running, and idle agents stack beneath in descending urgency. Interaction is big tap targets, swipe actions (swipe-to-approve, swipe-to-stop), and pull-to-refresh. Launching a new agent is not filling a blank form — it is duplicate-from-template: pick a saved recipe (repo + tool + prompt template) from a quick-start sheet and go. Drilling into an agent opens a full-screen sheet; on a phone its terminal is read-mostly with a one-tap “open on desktop” handoff, because a live terminal is genuinely hard to drive with a thumb.

The thesis in one line: the conductor whose laptop is closed and whose launches are mostly re-runs of known recipes should be able to run the whole monitor→approve→launch loop one-handed, in under a minute, from a feed that always shows the right thing first.

2. Parent User Flow & Branch Relationship

Position in the flow-map

V5 covers the monitoring-start branch: Happy Path steps 3–6 (S3 Agent Board → S4 Start Agent → S3 new card → S5/S6 Interact) plus Path C “Parallel Sprint” (launch agent #2/#3, monitor all, round-robin approvals). It is the monitor + launch surface of the product.

Two-layer obligation (hard constraint)

Per the brief’s confirmed reframe, this is a two-layer surface, not one job:

V5 serves both. What V5 contrasts against siblings is how the orchestration layer connects to the deep-work layer: here, a feed item taps open into a full-screen sheet that owns the viewport, then dismisses back to the feed (vs. V1’s persistent split-pane, V2’s full-page route, V3’s drawer-over-board, V4’s focus-mode back-stack).

Boundary vs. sibling branches (explicit — do not redesign these)

V5 lives next to two conceptually overlapping sibling branches. The boundary is deliberate:

If a behavior is deep-diff, queue-policy, QR-token, or recap-timeline logic, it belongs to a sibling branch. V5 owns the feed (monitor) and the template launch, and the thin connective tissue that surfaces sibling surfaces.

3. Target User Fit

Best fit: the “closed laptop, checked from phone, 3 PRs ready” conductor — the aha-moment persona at the deliberately phone-up end of the device-posture axis. This is the solo Agent Conductor (5–12 agents, 1–3 repos) whose day looks like:

Poor fit (state honestly): the keyboard-at-max-density power user who wants 12 agents in one dense viewport with split-pane drill-in and never leaves the desk — that user is V1 Operator Console. The command-native dev who lives in ⌘K is V4. V5 deliberately trades desktop multitasking density for one-handed glanceability. See §20 risks.

4. Onboarding / Activation Model

V5 inherits account/infra onboarding from the onboarding branch (Path A: S2 Setup Wizard → provision → health check → S10 summary → S3). V5 owns only the first-run state of the monitor+launch surface it lands on.

Empty feed (first arrival at S3)

The flow-map’s S3 “Empty” state (“Start your first agent” CTA) becomes a feed zero-state:

First recipe (the activation moment)

The activation event for V5 specifically is saving the first recipe. After the user’s first successful from-scratch launch, the launch-confirmation toast offers “Save as recipe” (repo + tool + prompt template, with the variable parts of the prompt optionally tokenized). Once one recipe exists, the quick-start sheet flips to recipe-first (duplicate-from-template primary, from-scratch secondary). This is the loop the persona is built for, so onboarding nudges toward it early.

PWA install prompt

5. Typical Workflow Sequence (the phone-up monitor→approve→launch loop)

The canonical V5 session, one-handed on a phone:

  1. Push fires — “Agent test-fix needs approval” (or “Agent deploy-prep errored”). User taps the notification. (Push + deep-link delivery is shared with approval/device-handoff; V5’s job starts at the feed.)
  2. Feed opens, urgent item already on top — PWA resolves to S3. The agent that fired is the top pinned feed item (approval or error), visually distinct. No scrolling, no hunting.
  3. Glance — the pinned item shows agent name, status, repo, the one-line “why it needs you” (e.g., “wants to run db:migrate” or “git merge conflict in auth.ts”), and a wait-duration timer.
  4. Decide by swipe — for a clear yes: swipe right → approve. For a clear stop: swipe left → stop/reject. Done; the item animates out and the next-most-urgent item rises. For anything needing thought, tap to open the full-screen sheet (→ hands into approval branch S7 for the real diff).
  5. Scan the rest — below the pinned zone, the user thumbs down the feed: waiting → running → idle. Live one-line stdout under each running item gives a calm pulse-check. Pull-to-refresh at top forces a state resync (also the manual escape hatch if the WebSocket looked stale).
  6. Launch a re-run — tap the persistent Launch action (bottom nav, §12). The template quick-start sheet slides up; the user taps a saved recipe (e.g., “lint-fix · repo-api · CLI”), optionally tweaks the one tokenized prompt field, taps Start. Sheet dismisses; a new “starting” feed item appears (transitions to running within seconds, live stdout begins).
  7. Parallel fan-out (Path C) — repeat step 6 for repo #2 and #3 from recipes; the feed now carries 3+ running items plus whatever pinned approvals arrive. Round-robin approvals are handled by repeating steps 2–4 as pushes land.
  8. Drop the phone — agents run on managed infra; the loop is closed. Deep review (real diffs, terminal debugging) is deferred to a desktop via the handoff (§13), not attempted on the phone.
Steps 1, 4-into-S7, and the desktop handoff are boundaries: V5 surfaces and routes; the approval and device-handoff branches own what happens past the boundary.

6. Progression Model — and Explicit Contrast With the Four Siblings

How the user advances in V5

Progression is gravity-driven, not navigation-driven. The user does not “go somewhere” to find work; the feed brings the work to the top. Advancement is:

  1. Top of feed (act on what’s pinned) →
  2. down the feed (scan calm agents) →
  3. launch (template sheet, return to feed) →
  4. drill (full-screen sheet only when an item demands it) →
  5. dismiss back to feed.

The feed is always the home position. Every drill-in is a temporary full-screen overlay that returns to the same feed scroll state. The mental model is a prioritized inbox: you work the top, things resolve and disappear, new things arrive at the top.

Explicit contrast with the other four variations

VariationBoard representationHow it advances / connectsV5’s contrast
V1 Operator Consoledense sortable data tablepersistent split-pane — fleet never leaves view; detail loads in right paneV5 has no persistent fleet view during detail — the sheet owns the whole phone viewport; you trade always-on fleet visibility for focus + one-handedness. V1 is for staying at the desk; V5 is for the phone.
V2 Card Mosaic controlresponsive card grid w/ live stdoutboard → full-page detail route; modal launchV5’s feed is one column, priority-sorted, not a grid — there’s never a “which card do I look at” scan; the answer is always “the top one.” V5 launch is template-first, not a blank modal.
V3 Kanban Pipelinecolumns = lifecycle statusboard → side drawer (board stays behind); “+” per columnV3 makes status spatial across columns; V5 makes status vertical priority in one stream. V3 cards move between columns as status changes (lateral motion); V5 items re-rank within one list (vertical reflow). V5 needs no horizontal scan — fatal on a phone, which is why V5 collapses Kanban’s axes into one.
V4 Command-Firstquiet minimal list + ⌘K palettepalette nav; detail = focus mode + back-stackV4’s verb is typing a command; V5’s verb is a swipe or a tap on a big target. Command-typing is awkward on a phone (V4’s own tradeoff); V5 is built for the exact posture where V4 is weakest. V4 advances by recall; V5 advances by recognition (it’s already on top).
V5 Glanceable Stream thisprioritized vertical feedfeed → full-screen sheet; template quick-start launch

The throughline: V5 is the only variation that assumes you are not looking at the screen continuously. It optimizes for the glance, not the session. The others assume an attentive operator at a desk; V5 assumes a distracted thumb and makes the right thing rise to it.

7. Sharing & Collaboration Model

N/A — single-user personal workstation. Per the locked scope constraint, gblock-party personal-workstation is a solo dogfood environment with no second human. There is no sharing of feeds, agents, recipes, or approvals; no shared views, no comments, no handoff-to-a-teammate. The only “handoff” in V5 is device-to-device for the same user (the device-handoff branch), which is continuity, not collaboration. No collaboration affordances are invented.

8. Permissions Model

N/A — single-user personal workstation. There are no roles, no per-agent ACLs, no shared-resource permissions, no multi-tenant boundaries. The single user owns everything. The only authorization in the system is GitHub OAuth identity + repo scopes (owned upstream by S1/onboarding), which is the user authenticating themselves to their own repos — not a permissions model over other users. V5 invents no permissions surface.

9. Return-Use & Notification Model (the strongest of the set)

This is V5’s signature strength — it is built for return-use, because the persona’s defining act is coming back to a closed-laptop fleet from a phone.

Pinned approvals & errors (the always-right-first feed)

The feed’s sort is V5’s reworking of the flow-map default (error > waiting > running > idle — explicitly reworkable per brief):

Pull-to-refresh

State is server-canonical over WebSocket and reflows automatically. Pull-to-refresh is the explicit user gesture for “resync now, I don’t trust what I’m seeing” — it forces a fresh state pull and reconnect, and is the natural mobile recovery action when the connection banner is showing (§10). It doubles as the “I just woke the phone” reassurance gesture.

Push notifications (installed PWA)

Why “strongest of the set”: V1/V4 optimize the attentive desktop session; V2/V3 are co-equal/desktop-reflow. Only V5 makes the return glance the primary designed-for moment — pinned-top + push + pull-to-refresh is a return-use loop the others treat as secondary.

10. Failure Recovery Behavior (mapped to flow-map failure paths)

V5’s core failure principle: anything that needs the user pins to the top of the feed, in the same pinned zone as approvals, with an error treatment. The user never has to go looking for a failure — failures come to them. Mapping the flow-map’s Failure & Recovery table:

Flow-map failureV5 surfacing in the feed
Backend unreachable (WebSocket heartbeat >10s)Top sticky banner above the pinned zone: “Reconnecting to workspace…” with auto-retry (exponential backoff). Last-known feed state stays visible (greyed live-stdout lines). Pull-to-refresh = manual reconnect. Launch/swipe actions disabled until reconnect (matches S3 “Connection Lost”).
Agent process crashAgent’s feed item pins to the top with an error badge + one-line cause. Auto-restart (up to 3, transient) shows inline as “restarting (2/3)…”. On persistent failure, tap → full-screen sheet (S5 Error state): error details/stack trace, Retry / Restart / Open terminal (terminal = desktop-handoff on phone, §13).
Network tunnel dropSame top banner treatment as backend-unreachable (“Tunnel reconnecting…”); agents unaffected (managed infra); feed shows last-known, resyncs on reconnect.
Model key / subscription expiry or rate limitAffected agent pins top with a distinct “needs credential” error item; tapping deep-links to S10 Settings to update the key, then the agent resumes without losing session state.
GitHub OAuth token expired/revokedAffected agents pin top as “needs re-auth”; tap → GitHub OAuth re-auth prompt; agents resume automatically once a new token issues.
Disk exhaustionPinned warning item (distinct from error) before critical, with a “free up space” suggestion; emergency pause surfaces as an error item.
Git conflict during agent workAgent pauses and pins top as an approval-style item with conflict details; resolving needs a real terminal/diff → full-screen sheet → “open on desktop” handoff (terminal resolution is desktop work; V5 routes there rather than faking it on a phone).
Browser/PWA crashService worker reloads to S3; agents unaffected (server-side); feed rebuilds from canonical state, pinned zone intact. No data loss.
Concurrent device conflictLast-write-wins for approvals (server-canonical); a swipe-approve from the phone that lost a race shows a brief “already handled on another device” toast and the item drops from the feed.

Recovery ergonomics: because failures share the pinned zone with approvals, the same swipe vocabulary applies where safe — e.g., swipe-to-retry on a crashed-agent item (with a confirm for destructive cases). Anything requiring real inspection (stack traces, conflict diffs, terminals) is a tap into the full-screen sheet, and on a phone routes to desktop handoff rather than cramming it onto the small screen.

11. Page & Flow Changes vs. the Flow-Map Baseline

Flow-map baselineV5 changeRationale
S3 Agent Board = responsive grid of cards (2–3 col desktop, 1 col mobile), sort error > waiting > running > idleSingle prioritized vertical feed with a sticky pinned zone (errors + approvals) above a scrollable stream zone (waiting > running > idle). Grid replaced by one column at all breakpoints.Phone-first; a grid forces a “which card?” scan, a feed answers “the top one.” Pinned zone is the reworked sort (allowed per brief).
S4 Start Agent = centered modal with blank repo/tool/prompt fieldsBottom-sheet quick-start that is recipe-first (duplicate-from-template); blank “new from scratch” demoted to secondary.Persona’s launches are mostly known re-runs; templating removes per-launch form-filling and is thumb-friendly as a sheet.
S5 Agent Detail + S6 Agent Terminal = two separate screens; Mode A/B is a user toggleOne folded full-screen sheet, tool-adaptive (terminal-primary for CLI tools, structured for API-native); no user Mode A/B toggle. Mobile terminal is read-mostly + “open on desktop” handoff.Brief’s confirmed reframe (folded + tool-adaptive; drop Mode A/B choice). A live terminal is impractical one-handed, so phone = read-mostly, desktop = interactive.
Drill-in handoff (S3 → S5/S6) “click card”Drill-in is a full-screen sheet that dismisses back to feed scroll state.Matches V5’s gravity-to-feed progression (§6).
Push tap → S7 or S3Push tap → feed with the item pinned top, or deep-link into the approval branch’s S7.V5 owns the landing feed; deep approval is the sibling branch.
S7 / S8 / S9 / S10Out of V5’s ownership — referenced, not redesigned. S7=approval, S8/handoff=device-handoff, S9=diff-review, S10=settings. V5 only links into them.Branch boundaries (§2).

No new screens are invented beyond folding S5/S6 and re-skinning S3/S4 — V5 is a re-presentation of the in-scope screens, not an expansion.

12. Navigation Model

Mobile-first navigation; three primitives only:

  1. The feed (home). Vertical scroll. Sticky pinned zone at top. Pull-to-refresh at the top edge. This is the always-return-to surface.
  2. Full-screen sheets. Two kinds — the template quick-start sheet (launch) and the agent detail sheet (drill-in). Both slide up over the feed, own the viewport, and dismiss (swipe-down on the sheet handle, or a top-left Back/Close) back to the exact feed scroll position. No nested deep stacks — the model is feed ↔ one sheet, never sheet-on-sheet (the only exception: detail sheet → into a sibling-branch surface like S7/S9, which is a branch boundary, not a V5 nav layer).
  3. Bottom nav (persistent, thumb-zone). Three to four targets, mobile-first ergonomics:
    • Feed (home / fleet)
    • Launch (opens template quick-start sheet) — center, emphasized
    • Approvals (jumps the feed to / filters the pinned approval items; badge count) — note the deep approval UX is the approval branch; this is just the fast jump
    • Settings (→ S10, sibling settings branch)

On desktop (>1024) the bottom nav migrates to a slim left rail or top bar (same four targets), and sheets become centered overlays rather than full-bleed (§17). The navigation vocabulary stays identical across breakpoints — only the chrome position changes — so the mental model is portable when the user moves phone→desktop.

13. Screen-by-Screen Layout (regions & proportions)

Dark scheme per house style (--bg:#1a1a2e; --card:#16213e; --border:#2a2a4a; --accent:#4dabf7; --text:#e0e0e0; --muted:#8b8fa3). Proportions given for the primary mobile ≤640 viewport; desktop deltas in §17.

S3 — Agent Board (the feed)

+-------------------------------------------+
| [= workspace]          [approvals .3] [gear]|  Header  ~7% h
+-------------------------------------------+
| (pull-to-refresh affordance on drag)      |
+-------------------------------------------+
| ## PINNED ZONE (sticky) ##                |  Pinned
| +-------------------------------------+   |  zone:
| | [STOP] deploy-prep . repo-api       |   |  grows to
| | git conflict in auth.ts . 4m        |   |  fit, caps
| |   < swipe: retry      stop >         |   |  at ~45% h
| +-------------------------------------+   |  then
| | [PAUSE] test-fix . repo-web         |   |  scrolls
| | wants db:migrate . waiting 2m       |   |  internally
| |   < swipe: approve    reject >       |   |
| +-------------------------------------+   |
+-------------------------------------------+
| STREAM ZONE (scroll)                      |  Stream
| +-------------------------------------+   |  zone:
| | > lint-fix . repo-api . running 12m |   |  fills
| |   $ applying fix to eslintrc...     |   |  remainder
| +-------------------------------------+   |  (one-line
| | > docs-gen . repo-web . running 3m  |   |  live
| |   $ writing README section...       |   |  stdout per
| +-------------------------------------+   |  running
| | o refactor . repo-cli . idle        |   |  item)
| +-------------------------------------+   |
+-------------------------------------------+
|   [Feed]   [ + Launch ]   [Approvals] [gear]| Bottom nav ~9% h
+-------------------------------------------+

S4 — Template Quick-Start Sheet (launch)

Full-height bottom sheet sliding up over the feed (mobile); centered overlay (desktop).

+-------------------------------------------+
|  ___ (drag handle)                  [x]   |
|  Start an agent                           |
+-------------------------------------------+
|  YOUR RECIPES                             |  Recipe
|  +---------------+ +---------------+      |  grid/list:
|  | lint-fix      | | test-fix      |      |  primary,
|  | repo-api.CLI  | | repo-web.CLI  |      |  ~60% h
|  +---------------+ +---------------+      |
|  | dep-bump      | | docs-gen      |      |  tap a
|  | repo-cli.API  | | repo-web.CLI  |      |  recipe ->
|  +---------------+ +---------------+      |  prefilled
+-------------------------------------------+
|  -> selected: lint-fix                    |  Confirm
|  prompt: "fix lint in {{path}}"  [edit]   |  strip:
|  +-------------------------------------+  |  one
|  | {{path}} = src/                     |  |  tokenized
|  +-------------------------------------+  |  field
|           [  Start agent  ]               |  editable
+-------------------------------------------+
|  + New from scratch  (secondary)          |  Secondary
+-------------------------------------------+

Full-Screen Detail Sheet — folded S5/S6 (tool-adaptive)

One sheet, no Mode A/B toggle. Primary rendering is decided by the agent’s tool integration pattern:

+-------------------------------------------+
| [< Back]  test-fix . repo-web . > 12m [..]|  Sheet header
+-------------------------------------------+
|  STRUCTURED CHROME (always)               |  ~30-40%
|  - progress / current step                |  parsed
|  - file-change list                       |  status
|  - inline approval card (when waiting)     |  (Multica)
+-------------------------------------------+
|  PRIMARY RENDER (tool-adaptive)           |
|   CLI tool -> terminal pane                |  remainder
|     [phone: READ-MOSTLY, scrollback only] |
|     +---------------------------------+   |
|     | $ npm test                      |   |
|     | + 42 passing  x 1 failing       |   |
|     | ...live xterm.js (read) ...     |   |
|     +---------------------------------+   |
|     [ Open on desktop to type ]  <-handoff|
|   API-native tool -> structured detail     |
|     (no terminal; rich parsed view)        |
+-------------------------------------------+
|  [ Approve ] [ Inject prompt ] [ Stop ]   |  Action bar
+-------------------------------------------+
Open-question alignment: mobile read-mostly terminal directly hedges the terminal-latency open question (flow-map Open Q #1) — we don’t bet the phone experience on sub-100ms keystroke echo; we bet it on reading, and route typing to desktop. stdout-parsing reliability (Open Q #2) bounds how much the structured chrome can promise per tool; where parsing is thin, the chrome degrades to “raw stdout tail” rather than over-promising structure. code-server/IDE (Open Q #4) is out of scope for V5 — V5 makes no deliberate case for an in-feed IDE; the “go deeper than a terminal” answer is the desktop handoff, not a browser IDE.

14. Key Components & Controls

Referenced-but-not-owned components (sibling branches): the deep diff viewer / risk-rated approval card (approval), the QR / pre-authed deep-link generator (device-handoff), the recap timeline S8 (device-handoff), S9 diff/merge (diff-review), S10 settings (settings).

15. Button & Link Behavior (swipe gestures + tap targets)

Swipe gestures (per feed item)

Item classSwipe rightSwipe leftTap (whole row)
Pinned approvalApprove (one-tap, optimistic; resolves & item exits)Reject (opens reason field if required → hands to approval S7)Open full-screen sheet / approval S7
Pinned errorRetry (with confirm if destructive)Stop agentOpen detail sheet (S5 Error: stack trace, retry/restart/terminal)
Stream — running(none / dismiss-to-idle n/a)Stop (confirm)Open detail sheet
Stream — waiting/idleStop / archiveOpen detail sheet

Tap targets & links

16. Spatial Density, Sizing, Hierarchy

17. Responsive Behavior at Three Breakpoints

Mobile ≤640 (primary, the design center):

Tablet ≤1024:

Desktop >1024 (centered feed + optional second pane):

The responsive story is a posture gradient, not a layout swap: phone = glance & route, desktop = read & act deeply. Same feed, same nav vocabulary, progressively more capable as the screen and input grow.

18. Visual Tone

Calm, focused, inbox-like. Dark house palette throughout (--bg/--card/--border/--accent/--text/--muted). The feeling target is “a quiet notification feed that lights up only when it matters” — most of the screen is calm --muted running/idle items; the pinned zone is where color and weight concentrate (error reds, approval --accent). Motion is purposeful and minimal: items animate out on resolve and rise on re-rank, swipe reveals are spring-eased, pull-to-refresh has a satisfying snap. Nothing pulses or demands attention except genuine pinned urgency. Generous whitespace, big legible type, thumb-friendly — it should feel like a well-made mobile app, not a dashboard crammed onto a phone.

19. Strengths

20. Risks & Failure Modes (with mitigations)

21. Implementation Complexity

Rough estimate: medium (consistent with the brief’s complexity column).

22. Prototype Scope

For a tangible, evaluable prototype (feeding /uat --variant-evaluation and /ui-interview):

In scope (build):

Out of scope (stub / reference only): real diff viewer & deep approval logic (approval branch), QR/deep-link token plumbing (device-handoff), recap S8, diff/merge S9, settings S10 internals, real backend/WebSocket (mock the stream), real auth, code-server/IDE (explicitly excluded). Push can be simulated rather than wired to a real push service.

Fidelity: static-HTML/CSS interactive prototype (per /prototype for UI projects) with mocked state transitions sufficient to feel the glance→swipe→launch loop — interaction feel is the thing being evaluated, so the swipe/refresh/re-rank gestures should be genuinely interactive even if data is faked.

23. Readiness Signal for /ui-interview

This spec is build-grade and ready for /ui-interview [v5-glanceable-stream] (or for serial prototyping ahead of /uat --variant-evaluation/consolidate-variations). It is decided and unambiguous on:

Open items intentionally deferred to /ui-interview (not blockers): exact type ramp / spacing tokens, precise swipe-threshold and animation-timing values, final glyph/color semantics within the dark palette, the exact pinned-zone height cap, and the precise bottom-nav target set (3 vs 4) — all visual/interaction-tuning decisions appropriate to the UI interview, not the UX variation.

Cross-variation note for /consolidate-variations: V5’s distinctive, cherry-pickable elements are (1) the pinned-zone return-use loop, (2) swipe-to-approve/stop, (3) the duplicate-from-template launch, and (4) the read-mostly-mobile-terminal + desktop-handoff pattern — any of which could be lifted into a consolidated direction even if the full feed model isn’t the chosen spine.

7. Experimentation Plan

Build all five variations serially, in full — this is the user’s stated preference. We do not build a subset first and gate the rest; each variation gets a complete, evaluable build before moving to UAT. (Per §6 each spec carries its own §22 Prototype Scope defining a build-grade, evaluable prototype.)

Route-based experiments (one per variation)

Each variation is built as its own route experiment under a dedicated path, so the five live side by side in one app shell without forking the repo:

VariationExperiment route
V1 Operator Console/experiments/operator-console
V2 Card Mosaic/experiments/card-mosaic
V3 Kanban Pipeline/experiments/kanban-pipeline
V4 Command-First/experiments/command-first
V5 Glanceable Stream/experiments/glanceable-stream

Rules for the experiments:

8. Comparison Criteria (defined BEFORE selecting a winner)

The criteria a winner is judged on are fixed now, before evaluation, so selection is evidence-driven rather than post-hoc rationalized. Each variation is scored against the same set during /uat --variant-evaluation:

CriterionWhat it measuresWhy it matters / which variations it pressures
Time-to-launch an agentwall-clock + steps to start a new agent (cold + parallel-sprint #2/#3)Launch is a core orchestration verb. Pressures V3 (column-anchored “+”) and V2 (modal re-entry) vs V5 (template), V4 (⌘K), V1 (inline row).
Glanceability of fleet status at 8–12 agentscan the user read “what needs me / what’s each agent doing” in one glance at scaleThe core monitor job at realistic load. Pressures V2/V5 (cards/feed scroll) vs V1 (dense table), V3 (column shape).
Drill-in continuityis fleet/board context preserved while working one agent; is return losslessTwo-layer fusion quality. Pressures V2 (route drops board) vs V1 (split-pane), V3 (drawer), V4 (back-stack).
Approval / error triage speedtime + friction to find and clear the next blocker; round-robin throughputReturn-use core loop. Pressures all; favors V3 (spatial queues), V5 (pinned + swipe), V1 (sort + inline).
Cross-device / posture fithow well the variation works across phone ↔ desktop without a model breakDevice posture is tested empirically, not pre-decided. Pressures V1 (mobile degrade) vs V5 (mobile-first), V2 (co-equal).
Implementation costbuild + maintenance complexity (relative, per each §21)Budget realism. V2 lowest; V3 highest; V1/V4/V5 medium.
Discoverabilitycan a user find capability without memorization/trainingFirst-meeting + low-frequency use. Pressures V4 (palette-hidden), V1 (NOC density) vs V2 (familiar dashboard).
Device posture is not a pre-decided winner: it is one of the criteria above, scored empirically across the set. A variation that wins on glanceability but loses on cross-device fit is a consolidation signal, not an automatic disqualifier.

9. Lock-In Checklist (decision record)

When a direction is chosen, record it as a decision (not a vague preference) by completing this checklist. The output becomes a dated decision entry referenced by /consolidate-variations and the eventual /ui-interview.

10. UAT Handoff Checklist (/uat --variant-evaluation)

Pack prerequisite

/uat --variant-evaluation requires the product-testing pack, which is NOT in .agents/project.json.enabled_packs. Run npx skillpacks install product-testing first, then start a new session before invoking /uat --variant-evaluation.

Per the variant-evaluation skill, prepare the following before running UAT against the five built route experiments:

11. Branch Routing

Sibling flow / variation dependencies

These siblings own deep stories that the monitoring-start variations only border — they are referenced, not redesigned here:

Next step (gated until approval)

Per the serial-build plan, all five variations are built before any /ui-interview, then /uat --variant-evaluation/consolidate-variations. When a /ui-interview candidate is taken to visual mockup:

Recommended first /ui-interview candidate (pending approval): v2-card-mosaic. As the baseline/control closest to the approved flow-map, it is the lowest-risk anchor for the comparison — interviewing it first establishes the control reference against which the other four are measured. (This recommendation does not change the serial-build order: all five are built first.) Downstream routing language stays gated until this review page is approved via compiled response YAML.

Alignment Gates

Each gate below is a required review decision placed in a visually distinct question block. Answer the required radio question; the two standing options (“Other / None of the above” and “Need clarification”) are always available. When you pick a non-standing option, an Additional notes box appears to qualify your choice. You do not have to answer every gate before compiling partial responses or sending section feedback.

Gate A — Surfaced Assumptions required

The plan rests on an assumptions manifest. Confirm it is correct before the variation specs are treated as canonical.

Is the surfaced assumptions manifest correct? (1) two-layer surface — every variation serves orchestration S3/S4 + deep-work folded S5/S6; (2) desktop-not-primary — device posture is a deliberate empirical axis, not a fixed desktop default; (3) tool-adaptive detail — rendering is determined by the tool integration pattern, not a user Mode A/B toggle; (4) single-user N/A — sharing/collaboration/permissions are N/A for this personal-workstation scope.

Gate B — Variation Manifest / Scope (fixed-vs-variable) required

The plan fixes some things across all variations and varies others. Confirm the fixed/variable split.

Is the fixed-vs-variable split right? Fixed (carried by all 5): the §3 reframe — two-layer surface, folded tool-adaptive detail (no Mode A/B), all-web PWA. Variable (the 4 axes): (a) board representation, (b) orchestrate↔deep-work connection, (c) launch affordance, (d) device posture/density.

Gate C — Concept Selection required

Pre-build, the available moves are: keep all 5, make one bolder, or add another. Removing or merging concepts is deliberately not offered before the build (that is the post-UAT /consolidate-variations job).

How should the concept set proceed into the serial build?

Gate D — Evaluation Method required

The plan proposes a specific evaluation pipeline after the builds.

Confirm the evaluation method: serial build of all 5 (each a complete, evaluable route experiment) → /uat --variant-evaluation (needs the product-testing pack installed first) → /consolidate-variations.

Gate E — Proposed Flow-Map Amendment required

Every variation carries the merged tool-adaptive S5/S6 model and drops the Mode A/B user toggle. The approved flow-map still lists S5/S6 as separate screens with Mode A/B (D2). This plan does not silently edit the flow-map; it proposes the amendment for the flow-map owner’s acceptance.

Do you approve carrying the merged tool-adaptive S5/S6 (drop Mode A/B) as a proposed amendment back to the approved flow-map?

Gate F — Artifact Destination & Proposed File Changes required

Destination and the proposed file-write set ask the same path-destination question here, so they are combined into one gate per the de-duplication rule.

Approve the canonical artifact paths for this UX-variations cycle?

Gate G — Coverage Checkpoint required

Before finalizing, confirm nothing decision-relevant is missing.

Are any decision criteria, risks, validation steps, or implementation constraints missing before finalizing the plan?

Gate H — Post-Approval Route required

This only takes effect after approval. While in review, the next action is review of this page, not downstream routing.

After approval, confirm the next step is /ui-interview v2-card-mosaic (with the serial-build-all-5 plan intact, and downstream routing language staying gated until approval).

Compile Responses

Aggregates all answered gate questions (with notes) and every selected section-feedback entry into one response YAML payload. Enabled as soon as one gate is answered or one section feedback is selected — you do not have to answer every gate to send partial responses.