Idea Scope Brief — Task ROI

PM-facing token-spend planning & model-performance audit • 2026-06-04 • Product path slug: task-roi (working name)
alignment_status: review  |  This is a pre-approval preview. Nothing is written to research/ yet. Review each section, then either send feedback-only YAML (👍/👎/❓ on any section) for revisions, or answer the required gate questions and use Compile Answers at the bottom to approve. Coverage-checkpoint confirmation alone does not authorize canonical writes — only final compiled YAML does.

1. Scope Resolution & Branch Relationship

from repo The repo is currently in flat single-product mode: all of research/*.md describe one concept — the cost-intelligence platform aimed at the Platform Engineering Lead ICP (locked 2026-05-24, hidden-multiplier cost modeling).

from user This brief introduces a second, sibling product path aimed at a different beneficiary — project / product managers — with a different core job: deciding which tasks justify agentic token spend vs. hand-coding, plus auditing model performance.

Decision recorded this session
QuestionYour answer
Relationship to cost-intelligence pathSibling product path — both stay active; separate research tracks
Meaning of "auditing model performance"Keep all three angles open (quality-vs-cost, model comparison, spend efficiency); research narrows to the most promising pain point
Source of "task" unit & quality signalNot sure yet — flagged as a top unknown to test in research

Because two concepts now coexist, this run will (on approval) create research/.progress.yaml recording both as product paths, rather than merging this into the existing brief.

Gatescope/non-goalsrequired
Confirm the structural relationship between this path and the existing cost-intelligence path.

2. Concept Identity & Slug

inferred Proposed working slug: task-roi — it captures both pillars (return on a unit of agentic task spend, both forward-planning and backward-audit). Output paths derive from this slug, so it's worth locking now (it can be renamed later, but renaming moves files).

Candidate slugReads asTrade-off
task-roiReturn on per-task agentic spendCovers plan + audit; slightly abstract
agentic-roiROI of agentic codingClear, but emphasizes coding over PM planning
spend-planningPlanning token spendMisses the model-audit pillar
velocity-auditAuditing AI-assisted velocityEmphasizes audit over forward planning
Gateidea identityrequired
Which slug should scope this product path (drives research/<slug>/ output paths)?

3. Idea Assumptions Manifest

Each assumption is tagged by source. These are hypotheses, not findings — downstream research (starting with /icp) owns evidence.

DimensionAssumptionSource
Concept summaryA PM-facing tool answering two linked questions: which tasks are worth agentic token spend vs. hand-coding (forward planning) and did the model perform well on what we spent on (backward audit)from prompt
ProblemPMs have no principled way to decide where agentic spend pays off; can't see which tasks burned tokens for low-quality output; planning is gut-feel, cost is a post-hoc surpriseinferred
BeneficiaryProject / product / delivery managers — the people planning work and defending budget — distinct from the Platform Eng Lead of the existing pathfrom prompt
CategoryAI delivery/velocity planning + model-output QA; adjacent to FinOps-for-AI but framed around task ROI and output quality, not infra attributioninferred
Value wedgePer-task "agent it vs. hand-code it" decision support combined with model-performance auditing — spend × quality on one task ledgerinferred
ConstraintsSolo dev, ship-fast; read-only provider data preferred — but quality auditing may need signals billing APIs don't expose; multi-providerfrom repo
Non-goalsNot a prompt-engineering tool; not a request-path proxy; not a replacement for the cost-intelligence pathinferred
Riskiest unknownWhere the "task" unit and quality signal come from — likely needs task/outcome telemetry (PR merged/reverted, review churn) beyond read-only billinginferred
Gateassumptions/confidencerequired
Are these core assumptions right enough to proceed to research?

Sections 4–12 below are the proposed content of research/task-roi/idea-brief.md, rendered in full for review. Nothing is written until approval.

4. Summary

A planning-and-audit tool for project/product managers that turns LLM spend from an infrastructure line item into a per-task ROI decision. It answers two linked questions:

  1. Forward (plan): Which upcoming tasks are worth spending agentic tokens on, and which should be hand-coded?
  2. Backward (audit): Did the model actually perform well on the tasks we spent on — was the spend worth it?

This is a sibling product path to the existing CalcLLM cost-intelligence platform. Where cost-intelligence serves Platform Engineering Leads with hidden-multiplier cost attribution, this path serves delivery managers with task-level spend-vs-quality decision support.

5. Problem Hypothesis

Hypothesis Product/delivery managers are now accountable for AI-assisted velocity but have no principled way to decide where agentic coding spend pays off. Decisions about "should we let the agent do this or hand-code it?" are made on gut feel, and there's no feedback loop telling them whether past spend produced acceptable work.

Specific pains (to validate)

6. Beneficiary Hypothesis

from prompt Primary beneficiary: project / product / delivery managers who plan work and defend budget. Distinct from the existing path's Platform Engineering Lead — these are people who think in tasks, tickets, sprints, and deliverables, not gateways and attribution.

Candidate personaWhy they might careConfidence
Product / project managerOwns roadmap & estimates; needs to price agentic work and justify itMedium
Eng manager / team leadOwns delivery quality; wants to know if agentic output is trustworthy per task typeMedium
Tech lead / staff engSets the "agent vs. hand-code" norms the team followsLow

Which of these is the true buyer/user is exactly what /icp resolves — this brief does not select it.

7. Product Category Guess

inferred AI delivery / velocity planning + model-output QA. Sits adjacent to FinOps-for-AI (the cost-intelligence path's category) but is framed around task ROI and output quality rather than infrastructure cost attribution. Think "the planning & scorecard layer for agentic work" rather than "the billing microscope."

8. Value Wedge

Core wedge

Spend × quality on one task ledger. Pure cost tools show what you spent; pure eval tools show output quality. Neither tells a PM "this kind of task is worth agentic spend and this kind isn't." The wedge is joining the two at the task unit for a planning audience.

The "auditing model performance" pillar spans three angles kept deliberately open — research will pick the most promising one:

AngleQuestion it answersData difficulty
Output quality vs. costDid the tokens we spent produce acceptable work (merged, low-rework)?High — needs outcome telemetry
Model / provider comparisonWhich model is best/cheapest for which task type?High — needs labeled outcomes per model
Spend efficiencyWhere is waste (retries, oversized context, redundant runs) at task granularity?Medium — partly from usage data

Note: the comparison angle crosses the existing path's stated non-goal ("Not a model evaluation or benchmarking platform"). That's intentional for this branch — see Non-Goals gate.

Gatescope/non-goalsrequired
How should the three audit angles be scoped for now?

9. Constraints

10. Non-Goals

Gatescope/non-goalsrequired
Confirm crossing the prior "no model evaluation/benchmarking" non-goal for this branch.

11. Assumptions & Unknowns

Assumptions

Riskiest unknowns

#UnknownImpact
1Task unit & quality-signal source. Where does the "task" boundary and the quality signal come from — VCS/PRs, agentic-tool session logs, or manual PM tagging? You marked this "not sure yet"Blocks core value
2Read-only data sufficiency. Can quality be inferred without instrumenting the request path or dev pipeline?Shapes architecture
3Buyer identity. Is the PM the buyer, or does eng leadership own this and the PM just consumes it?Shapes GTM
4Overlap with cost-intelligence. Where do the two paths share a spine vs. diverge — does this cannibalize or complement?Shapes portfolio
5Which audit angle wins. Quality-vs-cost, comparison, or efficiency — which is the sharpest pain?Shapes wedge
Gateassumptions/confidencerequired
How should the task-unit / quality-signal source (unknown #1) be handled?

12. ICP Readiness

Status: Ready for /icp as a new product path, scoped to research/task-roi/.

Inputs /icp should use:

Assumptions to test first:

  1. The task-unit / quality-signal source (unknown #1) — the architecture hinges on it.
  2. Whether the PM is the buyer or eng leadership owns it.
  3. Which of the three audit angles is the sharpest pain.
GateICP readinessrequired
Is the concept ready to enter /icp for the task-roi path?

13. Artifact Destination & Proposed File Changes

On approval, these files would be created (the existing flat research/*.md cost-intelligence files are not touched):

PathActionContents
research/task-roi/idea-brief.mdcreateSections 4–14 above
research/task-roi/idea-brief-interview.mdcreateThis session's Q&A + decisions
research/.progress.yamlcreateTwo product paths: cost-intelligence (active) + task-roi (active, pipeline_stage: idea-scope-brief)
alignment/idea-scope-brief-task-roi.htmlconvert to confirmedThis page, post-approval

Slug shown as task-roi; if you pick a different slug in Gate 2, these paths update accordingly.

Gateartifact destination + file changesrequired
Approve these output paths and the file-creation scope?

14. Coverage Checkpoint

Concept is shaped enough to enter research. The single biggest open item is the task-unit / quality-signal source (unknown #1), deliberately left for research. No competitive analysis, journey mapping, UX, or implementation planning was done here — that's downstream.

Gatecoverage checkpointrequired
Is any core premise, constraint, or non-goal wrong before writing?

15. Proposed Next Steps

Proposed only — routing activates after approval, not now.

Primary (proposed): /icp research/task-roi — the business-discovery pack is already enabled, so ICP discovery is the next research step for this path.

Also useful: /product-line review once a third path appears (you'll have two after this).


Compile / Approve

Two paths:

8 required questions remaining.