Your team spent $160.69 on agentic AI in the last 30 days.
We found
$50.90/month
in identified savings across 11 recommendations.
Burn & forecast
Plan caps are USD-equivalent estimates (Anthropic publishes caps in opaque units), user-adjustable.
Top projects
| Project | Spend | Sessions | Events |
|---|---|---|---|
| billing-service | $68.82 | 21 | 119 |
| acme-api | $49.82 | 51 | 342 |
| acme-web | $36.84 | 47 | 385 |
| infra-tools | $5.21 | 8 | 151 |
Spend by model
| Model | Spend | Events |
|---|---|---|
| claude-opus-4-8 | $105.50 | 295 |
| claude-sonnet-5 | $34.75 | 356 |
| gpt-5.4 | $9.97 | 94 |
| claude-haiku-4-5-20251001 | $7.63 | 213 |
| gemini-2.5-pro | $2.83 | 39 |
The edges — top savings
Route routine work off claude-opus-4-8 → claude-sonnet-5
$34.08 of claude-opus-4-8 spend came from routine-shaped sessions. claude-sonnet-5 costs 40% less (price ratio 0.60). Quality note: Sonnet 5 trails Opus 4.8 only slightly on SWE-bench Verified (79.6% vs 80.8%) and matches it on GDPval knowledge work (1618 vs 1615), but Opus leads on SWE-bench Pro (69.2% vs 63.2%) and long-horizon multi-file reliability; safe downgrade for most coding, keep Opus for long autonomous runs.
routineSpend=$34.08 of totalSpend=$105.50 on claude-opus-4-8 over 30.1d; priceRatio=0.60; savings=routineSpend×(1−0.60)×0.998/mo; source https://www.marktechpost.com/2026/07/13/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared/
Route routine work off claude-sonnet-5 → claude-haiku-4-5-20251001
$12.69 of claude-sonnet-5 spend came from routine-shaped sessions. claude-haiku-4-5-20251001 costs 67% less (price ratio 0.33). Quality note: Haiku 4.5 is Anthropic's fastest, near-frontier for high-volume simple work (file nav, small edits, linting, subagents) but weaker on complex multi-step refactors; downgrade only for scoped/mechanical tasks.
routineSpend=$12.69 of totalSpend=$34.75 on claude-sonnet-5 over 30.1d; priceRatio=0.33; savings=routineSpend×(1−0.33)×0.998/mo; source https://platform.claude.com/docs/en/about-claude/models/overview.md
Improve cache reuse on billing-service
Cache-read share is 64% vs the best project's 97%. Closing half that gap moves ~2,928,099 tokens to cache-read pricing (≈90% cheaper than fresh input).
share=64.1% best=97.1% gap=33.0%; movedTokens=2928099; avgCost/token=0.00000318; discount=0.9; savings/mo=$8.37
Improve cache reuse on acme-api
Cache-read share is 59% vs the best project's 97%. Closing half that gap moves ~3,318,996 tokens to cache-read pricing (≈90% cheaper than fresh input).
share=58.9% best=97.1% gap=38.1%; movedTokens=3318996; avgCost/token=0.00000254; discount=0.9; savings/mo=$7.56
Consolidate duplicate sessions on acme-api
5 overlapping same-project sessions ran concurrently, duplicating $4.44 of redundant spend — likely two agents on the same work.
overlappingPairs=5; duplicatedSpend=$4.44 (Σ smaller-session cost); savings/mo=$4.43
Eliminate off-hours idle spend on infra-tools
$2.22 was spent between 00:00–06:00 across 4 different days — a recurring off-hours cadence typical of unattended idle loops.
offHoursSpend=$2.22 across 4 days in window 0:00–6:00; savings/mo=$2.22
Route routine work off gemini-2.5-pro → gemini-2.5-flash
$2.21 of gemini-2.5-pro spend came from routine-shaped sessions. gemini-2.5-flash costs 75% less (price ratio 0.25). Quality note: Flash is far cheaper and faster and strong on straightforward tasks, but below 2.5 Pro on complex reasoning and long-context accuracy.
routineSpend=$2.21 of totalSpend=$2.83 on gemini-2.5-pro over 30.1d; priceRatio=0.25; savings=routineSpend×(1−0.25)×0.998/mo; source https://www.cloudzero.com/blog/gemini-pricing/
Consolidate duplicate sessions on acme-web
3 overlapping same-project sessions ran concurrently, duplicating $1.50 of redundant spend — likely two agents on the same work.
overlappingPairs=3; duplicatedSpend=$1.50 (Σ smaller-session cost); savings/mo=$1.50
Consolidate duplicate sessions on infra-tools
2 overlapping same-project sessions ran concurrently, duplicating $1.50 of redundant spend — likely two agents on the same work.
overlappingPairs=2; duplicatedSpend=$1.50 (Σ smaller-session cost); savings/mo=$1.49
Improve cache reuse on infra-tools
Cache-read share is 51% vs the best project's 97%. Closing half that gap moves ~1,162,621 tokens to cache-read pricing (≈90% cheaper than fresh input).
share=50.8% best=97.1% gap=46.3%; movedTokens=1162621; avgCost/token=0.00000094; discount=0.9; savings/mo=$0.98
Route routine work off gpt-5.4 → gpt-5.4-mini
$0.86 of gpt-5.4 spend came from routine-shaped sessions. gpt-5.4-mini costs 70% less (price ratio 0.30). Quality note: Same generation; mini trades reasoning/coding depth for ~3x lower blended cost and higher speed — fine for routine edits and tool calls, weaker on hard multi-step problems.
routineSpend=$0.86 of totalSpend=$9.97 on gpt-5.4 over 30.1d; priceRatio=0.30; savings=routineSpend×(1−0.30)×0.998/mo; source https://developers.openai.com/api/docs/pricing
Incidents — 1 runaway burst
- Runaway burn — billing-service, Jul 5, 03:00 UTC — $52.71 in a 15-min window vs $0.00 baseline (∞σ (flat baseline)).
CFO units approximate attribution
git log exited 1 in /home/dev/work/billing-service (not a git repo?); CFO PR attribution skipped.