Token Spend & Value Analysis review
Cross-platform analysis of Claude Code and Codex token usage, cost estimation, and value assessment across 6 months of development activity (Dec 2025 — Jun 2026).
Overview Stats
Source Comparison: Claude vs Codex
Monthly Activity Comparison
| Month | Claude Prompts | Claude Sessions | Codex Prompts | Codex Sessions |
|---|---|---|---|---|
| 2025-12 | 190 | 8 | — | — |
| 2026-01 | 1,309 | 289 | 89 | 13 |
| 2026-02 | 1,786 | 730 | — | — |
| 2026-03 | 2,544 | 1,039 | 176 | 36 |
| 2026-04 | 2,468 | 993 | 3,851 | 1,409 |
| 2026-05 | 1,553 | 680 | 2,292 | 1,035 |
| 2026-06 | 148 | 64 | 140 | 71 |
Key Observation: Platform Shift
Claude Code dominated Jan–Mar 2026 (sole primary tool). Codex adoption exploded in April — going from 176 prompts in March to 3,851 in April. By May, Codex surpassed Claude in raw prompt volume. This coincides with Codex becoming production-ready and your shift to parallel autonomous workflows.
Claude Code usage dropped ~37% from April to May, suggesting active migration of certain workflow types to Codex rather than net-new capacity.
Claude Code Activity Density Over Time
| Month | Messages | Sessions | Tool Calls | Msgs/Session | Tools/Session |
|---|---|---|---|---|---|
| 2026-01 | 75,467 | 278 | 14,872 | 271 | 53 |
| 2026-02 | 220,660 | 736 | 34,114 | 300 | 46 |
| 2026-03 | 130,571 | 850 | 83,018 | 154 | 98 |
| 2026-04 | 139,118 | 1,175 | 68,889 | 118 | 59 |
Trend: Sessions Got Shorter, Tool Use Got Denser
Messages per session dropped from 300 (Feb) to 118 (Apr) while tools per message increased from 0.15 to 0.50. This means sessions became more focused and tool-heavy — less conversational, more execution-oriented. Likely driven by skill adoption and shipping conventions.
Codex Token Deep Dive
Token Breakdown
| Category | Tokens | % of Total |
|---|---|---|
| Total input | 1,052,000,000 | 99.5% |
| — Cached input | 992,873,000 | 94.4% of input |
| — Non-cached input | 59,124,000 | 5.6% of input |
| Total output | 4,887,000 | 0.5% |
| — Reasoning output | 883,000 | 18.1% of output |
| — Non-reasoning output | 4,005,000 | 81.9% of output |
Top Projects by Codex Token Spend
Costliest Individual Sessions
| Date | Project | Total Tokens | Output Tokens | Tok/Output Ratio |
|---|---|---|---|---|
| 2026-05-22 | agentic-skills | 23.6M | 39.7K | 595:1 |
| 2026-06-01 | mobile-ideas | 23.0M | 55.5K | 415:1 |
| 2026-05-21 | agentic-skills | 22.3M | 46.0K | 485:1 |
| 2026-05-21 | agentic-skills | 21.6M | 38.8K | 556:1 |
| 2026-06-01 | agentic-skills | 17.6M | 43.6K | 404:1 |
Claude Code Activity Deep Dive
Top Projects by Claude Prompts
Skill Usage (Claude Code)
| Skill | Count | % of Total Skill Calls | Assessment |
|---|---|---|---|
| /ship | 1,972 | 79.2% | High-value shipping automation |
| /run | 212 | 8.5% | Verification & testing |
| /investigate | 120 | 4.8% | Bug triage & root cause |
| /sync | 97 | 3.9% | Repo sync utility |
| /pack | 51 | 2.0% | Skill pack management |
| /analyze-sessions | 35 | 1.4% | Meta-analysis |
| /expert-review | 13 | 0.5% | Code review |
| /code-review | 10 | 0.4% | Code review |
Cost Estimation
Codex Cost Breakdown (Actual Token Data)
| Month | Est. Cost | Sessions | Input Cost | Output Cost |
|---|---|---|---|---|
| 2026-03 | $31.59 | 28 | $29.52 | $2.08 |
| 2026-04 | $0.99 | 1 | $0.94 | $0.04 |
| 2026-05 | $222.71 | 228 | $209.98 | $12.74 |
| 2026-06 | $104.29 | 70 | $97.64 | $6.65 |
| TOTAL | $359.58 | 327 | $338.08 | $21.51 |
Claude Code Cost Estimate (No Local Token Data)
| Month | Est. Cost | Messages | Sessions | Tool Calls |
|---|---|---|---|---|
| 2026-01 | ~$817 | 75,467 | 278 | 14,872 |
| 2026-02 | ~$2,310 | 220,660 | 736 | 34,114 |
| 2026-03 | ~$1,891 | 130,571 | 850 | 83,018 |
| 2026-04 | ~$1,852 | 139,118 | 1,175 | 68,889 |
| TOTAL | ~$6,870 | 565,816 | 3,039 | 200,893 |
• Model tier: Opus sessions cost ~5× more than Sonnet
• Session length: longer sessions accumulate larger context windows
• Caching: actual cache rate varies per session
• May–Jun data: stats-cache only covers Jan–Apr; May–Jun is extrapolated from history.jsonl
Check your Anthropic billing dashboard for authoritative numbers.
Value Analysis
Where the Tokens Go
Claude Code: Interactive Development Hub
Claude Code is your primary interactive development tool. The heaviest spend is on:
- bismarck-v0.4 (1,228 prompts) — game project, heavy /ship usage
- loadoutworks.com (962 prompts) — web project, active shipping
- metternich-engine (906 prompts) — game engine, active development
- lexcorp-war-room (694 prompts) — active project with /investigate usage
- /ship is 79% of all skill calls — the shipping automation is by far the most-used workflow
Codex: Autonomous Batch Processor
Codex handles autonomous, parallelizable work. Top spend:
- mobile-ideas (106 sessions, 321M tokens) — largest Codex consumer by far
- agentic-skills (47 sessions, 282M tokens) — skill development and meta-work
- pitwall-monorepo (68 sessions, 190M tokens) — monorepo operations
- Average cost per session: $1.10 (very cheap)
- Average cost per prompt: $0.05 (negligible)
Value per Dollar
| Metric | Claude Code | Codex | Winner |
|---|---|---|---|
| Cost per user prompt | ~$0.69 | $0.05 | Codex (14× cheaper) |
| Cost per session | ~$2.26 | $1.10 | Codex (2× cheaper) |
| Interactivity | High (chat-driven) | Low (fire-and-forget) | Claude (by design) |
| Skill integration | Deep (/ship, /investigate, /run) | Prompt-based | Claude (richer) |
| Parallelism | 1 session at a time | Many concurrent | Codex |
| Total projects active | 15+ | 15+ | Tie |
Value Verdict
Claude Code (~$6.9K) delivers value through interactive development, debugging, and shipping automation. The /ship skill alone (1,972 calls) automates commit+push+docs workflows that would take 5–10 min each manually — conservatively saving ~165 hours of shipping ceremony. At a $100/hr developer rate, that's ~$16,500 in avoided manual work from /ship alone.
Codex (~$360) delivers exceptional token efficiency. At $1.10/session and $0.05/prompt, it's a bargain for autonomous batch work. The high cache rate (94.4%) means you're barely paying for repeated context reads. The value proposition is clear: Codex handles the long-tail of parallel, autonomous tasks that don't need real-time interaction.
Combined, the tools complement each other well — Claude for interactive/creative/debugging work, Codex for autonomous/batch/parallel work. The shift from Claude-heavy (Jan–Mar) to Codex-augmented (Apr–Jun) suggests you've found a productive division of labor.
Efficiency & Optimization Opportunities
1. Codex Input/Output Ratio (215:1)
Observation: Codex reads 215 tokens for every 1 it writes. The costliest sessions hit 400–600:1.
Implication: Even with 94.4% caching, the sheer volume (1.05B input tokens) costs $338 in input alone. If the non-cached fraction could be reduced from 5.6% to 3%, monthly Codex spend drops by ~40%.
Action: For agentic-skills (the project with the highest tokens-per-prompt at 1,369K), consider whether Codex sessions are re-reading full repo context unnecessarily. Scoping base instructions or using smaller working sets could cut input tokens significantly.
2. Claude Code Feb Spike ($2,310)
Observation: February was the costliest month — 220,660 messages, 300 messages/session (longest average sessions).
Implication: Long sessions accumulate context that pushes up per-message cost. The Feb spike is likely from early-stage project development with lots of back-and-forth.
Action: The trend already improved — Apr dropped to 118 msgs/session. The natural shift to shorter, tool-heavy sessions is the right trajectory.
3. Skill Coverage Gap
Observation: /ship (79%) dominates skill usage. /code-review and /expert-review together are only 23 calls — surprisingly low for the volume of code being produced.
Action: Consider whether more systematic code review before /ship would catch issues earlier. The /investigate count (120) suggests bugs are being found after the fact.
Evidence Matrix
| Claim | Source | Confidence | Notes |
|---|---|---|---|
| Codex total spend ~$360 | Token counts from ~/.codex/sessions/**/*.jsonl | High | Actual token data; pricing assumptions are the main variable |
| Codex cache hit rate 94.4% | Aggregated total_token_usage from session files | High | Direct measurement |
| Claude Code total spend ~$6.9K | Heuristic from stats-cache.json message/tool counts | Medium | No local token data; depends on model tier, session length, cache rate assumptions |
| /ship saves ~165 hours | 1,972 invocations × 5 min estimate | Medium | Time-per-ship is estimated; actual varies by complexity |
| Platform shift Apr–May | Monthly prompt counts from both history.jsonl files | High | Clear numerical trend |
| Sessions got shorter over time | stats-cache.json daily activity | High | msgs/session: 300→118 over Jan–Apr |
Assumptions & Confidence Register
| Assumption | Impact if Wrong | What Would Change It |
|---|---|---|
| Claude Code uses claude-sonnet-4 pricing ($3/$0.30/$15) | High — Opus sessions cost 5× more | Check billing dashboard for actual model mix |
| 80% cache hit rate for Claude Code | Medium — lower rate increases cost by ~2× | Real cache rate data from API logs |
| ~2K input tokens per message average | Medium — could be 3-5× higher for large contexts | Token data from billing API |
| Codex pricing: o4-mini rates | Low — Codex primarily uses o4-mini | Model field in session_meta confirms |
| stats-cache.json is complete for Jan–Apr | Low — it's the official local cache | Gap in dates would indicate missed data |
Approval Gates
Gate: Evidence Accuracy
The Claude Code cost estimate (~$6.9K) is heuristic-based. Does this feel directionally correct based on your billing?
Gate: Value Assessment
Is the /ship time-savings framing (165 hours / ~$16.5K equivalent) a useful way to think about value, or do you measure value differently?
Gate: Optimization Priority
Which optimization area matters most to you right now?
Compile Section
Feedback YAML — send concerns or clarification requests before answering all gates.
Copied!Final Approval — answer all gates first.
Copied!