How to run the $200 Max plan through the month at 2 to 10 active hours a day, which work to move to cheaper models, and where DeepSeek, Kimi, GLM, Qwen and the US hosts that serve them fit.
Every "$" figure for Claude usage in this report is what that usage would cost if bought through the API. It is a unit for measuring usage, not money charged. The Max plan covers usage up to its limits for the $200 monthly fee. Money beyond the fee is charged only when usage credits are on and usage passes the limit, and the account's monthly spend cap stops those charges (it did on 2026-08-26). Actual charges are on claude.ai under Settings → Usage. The dollar unit is used because Anthropic publishes limits in no other measurable form.
The usage figures below were measured on one principal's own AI subscription over two weeks in September 2026 and are reported as recorded, with names and internal systems removed. Source discipline: every vendor price, plan term, government action and standard was read on the publisher's own page on 2026-09-15 before it was used. Estimates are labelled estimate and carry their formula. Secondary figures that could not be verified were dropped; the list is in §10.
The $200 Max plan is not a flat price for unlimited work. It is a metered pool with two ceilings: a five-hour session limit and a weekly limit, shared by every Claude surface.12 From two limit events on this account, those ceilings sit near $487 of API-equivalent usage per five-hour window and $2,899 per week — roughly $12,600 a month (estimate: $2,899 × 30.44 ÷ 7).68 Those are single observations of an algorithmic system, not published constants.
Where the plan runs out today, and after optimization (15% of usage assumed to come from Claude chat, Design and desktop, which the transcripts do not capture; full sensitivity in §04):
| Active hours per day | Current habits — 5-day week | Current — 7-day week | Optimized — 5-day week | Optimized — 7-day week |
|---|---|---|---|---|
| 2 | Fits (52% of weekly) | Fits (73%) | Fits (25%) | Fits (35%) |
| 4 | Fits (80%), heavy 5-hour windows at 92% | Runs out ~day 7 (112%) | Fits (38%) | Fits (53%) |
| 6 | Runs out day 5 (108%); 5-hour limit hit on heavy days | Runs out day 5 (151%) | Fits (51%) | Fits (71%) |
| 8 | Runs out day 4 (136%) | Runs out day 4 (190%) | Fits (64%) | Fits (89%) — tight |
| 10 | Runs out day 4 (164%) | Runs out day 4 (230%) | Fits (77%) | Over by ~8% (108%, day 7) |
Optimized = Sonnet 5 as the default with Opus used for about 15% of cost-weighted work, Haiku for background subagents, conversations reset or compacted well before they reach hundreds of thousands of tokens, and no more than two background agents at once. Numbers are from the scenario model in §04 and are estimates.
What to move where.
Expected savings. At 6 active hours a day, 5 days a week, the combined levers take modeled weekly usage from $2,660 to $1,253–$1,429 (−46% to −53%), which moves the plan from "runs out on day 5" to "about half the weekly limit used." None of the levers costs money; each is a setting or a habit.
The principal's own telemetry covers every Claude Code session and subagent file on the principal's workstation: 1,502 files, 25,512 priced assistant messages since 2026-09-01, priced at Anthropic's published per-token rates.689
| Measure | Value | What it means |
|---|---|---|
| API-equivalent usage | $5,220.58 | What the same tokens would cost on the API; the plan meters against this kind of consumption |
| Cost by model | Opus 66.6% · Fable 28.0% · Sonnet 5.3% · Haiku 0.06% | Opus is the default on Max,13 and nearly nothing runs on Haiku |
| Cache-read tokens as share of cost | 72.6% | Long sessions re-send the whole conversation every turn; this outweighs model choice |
| Background agents (subagents) | 33.2% of cost | One third of usage is agents the principal did not type to directly |
| Rolling 5-hour windows | median $91 · p95 $324 · max $486 | The max window is the one that hit the session limit on 9/8 |
| Active human hours per day | 0.7 to 11.2 | Heaviest on 9/1 and 9/14 |
| Largest single session | $1,133, 1,926 messages | Open across ~264 calendar hours |
| Week 37 (Sep 7–13) | $2,123 for 20.5 active hours = $103.5 per active hour | The baseline week for the scenarios |
Two corrections to the input notes. The notes cite "97,361 unique priced messages (25,512 since 9/1)"; the 97,361 figure is every unique assistant message in the scanned files with no date filter, and 25,512 is the priced count since 2026-09-01 (it equals the sum of the per-day counts). The notes give the background share as 32.8%; the data file gives 33.2%.
Coverage gap. These transcripts see Claude Code only. Claude chat, Claude Design, the desktop chat and anything else signed in with the same account draw on the same pool, and Anthropic says so explicitly: "your usage of all different Claude product surfaces … counts towards the same usage limit."2 Every number above is therefore a floor. the principal's description of the situation — the same usage pattern in a different chat window — matches Anthropic's.
the firm's own telemetry for 2026-09-08 to 2026-09-15 found $4,158 of subscription usage at API rates in 8 days (37.8 active hours, about $110 per active hour) against ~$103 of real spend on the platform's API key.69 Its first pass reported $1,523 and missed 493 background-agent files; the corrected figure is the one used here.
Two items in that report need adjusting against the verified price sheet:
The telemetry report warned that this week's pace, if billed at API rates, implies a bill of $10,000 or more a month, and the principal caught the usage warning. The mechanism is real: once usage credits are enabled, "subsequent usage will be billed at standard API pricing rates."4 The size depends on how much usage exceeds the limit, not on total usage:
| Unmeasured share (chat, Design, desktop) | Weekly usage vs calibrated weekly limit | Estimated monthly overflow if uncapped |
|---|---|---|
| 0% | 126% | ~$3,200 |
| 15% | 148% | ~$6,000 |
| 30% | 179% | ~$10,000 |
Estimate: weekly pace = $4,158 ÷ 8 × 7 = $3,638; plan-wide = pace ÷ (1 − unmeasured share); overflow = (plan-wide − $2,899) × 30.44 ÷ 7. This ignores five-hour overflows on bursty days, so it is a floor. It is also not the $15,600 "per month at this pace" figure, which prices all usage rather than only the part past the limit.
Two safeguards already exist in Anthropic's product. A monthly spending cap can be set on usage credits (or set to unlimited),4 and the account hit that cap on 2026-08-26 about $158 past the weekly limit.68 Discounted usage bundles of $50, $250 and $1,000 carry 10%, 20% and 30% discounts, up to $2,000 a month on individual plans.5
The model bench ran eight graded suites across Anthropic and OpenAI models and effort levels for $2.23.70 "Sufficient" means the cheapest configuration within 95% of the best score (100% for the two engineering suites).
| Suite | Best configuration (pass) | Cheapest sufficient configuration | Opus 5 best effort, $/case ÷ sufficient $/case | How to read it |
|---|---|---|---|---|
| aisc (steel shape lookups) | Opus 5 low (94.4%) | Opus 5 low | 1.0× | Deterministic data — use a lookup table in code, not a model |
| units (unit conversion) | Haiku 4.5 (100%) | Haiku 4.5 | 26× | Deterministic — code first |
| classify | GPT-5 mini low, Haiku 4.5 and others (100%) | GPT-5 mini low | 19× | Small models suffice |
| memory-classify | Sonnet 5 low/medium (60%) | Sonnet 5 | 2.7× | Best is only 60%: a labelling or task-definition problem, not evidence any model is fine |
| oversight | Sonnet 5 and Opus 5 (100%) | Sonnet 5 low | 2.5× | Constructed cases, n=18 |
| grounded-qa | GPT-5 mini low (94.1%) | GPT-5 mini low | 12× | Constructed cases, n=17 |
| refuse-vs-guess | Several (100%) | GPT-5 mini low | 12× | Constructed cases, n=18 |
| code-fix | GPT-5 mini low, Haiku, Sonnet (100%) | GPT-5 mini low | 11× | n=8; Opus 5 scored 62.5–75% |
What the bench supports, read carefully:
An operational gotcha worth a paragraph. gpt-5 at high reasoning effort scored 0% on aisc and 5.6% on refuse-vs-guess. The model was not wrong; the calls ended with empty answers, because at a 300-token output ceiling its hidden reasoning consumed the entire budget before any visible text was written (the harness recorded output tokens at the cap with an empty answer string). GPT-5 mini at high effort showed the same failure less severely. Anthropic's adaptive-thinking models did not show it at the same ceilings. The practical rule: raising reasoning effort on a short-answer task can make results strictly worse unless the output limit rises with it, and raising the limit raises worst-case cost. Claude Code's own documentation makes the parallel point that thinking tokens are billed as output tokens and that lower effort is the cost lever for simpler tasks.12
| Term | Anthropic's published position (fetched 2026-09-15) |
|---|---|
| Price | Max 5x $100 a month; Max 20x $200 a month; monthly billing only1 |
| Multiple | "Max 20x provides 20 times more usage per session than the Pro plan"1 |
| Five-hour limit | "Your session-based usage limit will reset every five hours"1 |
| Weekly limit | A weekly limit across all models that "resets at a fixed time each week that is assigned to your account"1 |
| Discretionary limits | Anthropic "may limit your usage in other ways, such as weekly and monthly caps or model and feature usage"1 |
| Shared pool | claude.ai, Claude Code and Claude Desktop "count towards the same usage limit"; IDE usage too23 |
| What drives usage | Conversation length and complexity, features, model and effort level2 |
| Fable models | Included on Max; "up to 50% of your weekly usage limits on Fable models," drawn from the same weekly limit6 |
Agent SDK and claude -p | A June 2026 plan to move these to a separate monthly credit was paused; they still draw from subscription limits8 |
| At the limit | Wait, upgrade, or buy usage credits23 |
| Usage credits | Billed at standard API rates; monthly spend limit or unlimited; optional auto-reload; $2,000 daily redemption limit4 |
| Usage bundles | Prepay $50 / $250 / $1,000 at 10% / 20% / 30% off; up to $2,000 a month on Pro and Max5 |
| API key override | An ANTHROPIC_API_KEY environment variable makes Claude Code bill the API instead of the plan3 |
| Terms for automation | "Advertised usage limits for Pro and Max plans assume ordinary, individual usage of Claude Code and the Agent SDK"; plan credentials may not be used to route requests on behalf of other users10 |
No published page gives a number for either limit. The figures in this report come from this account's own limit events.
| Event (Central time) | What fired | Usage at that moment (API-equivalent) |
|---|---|---|
| 2026-08-25 08:39 | "hit your weekly limit" | Week-to-date $2,899 (week boundary Wednesday ~9 pm CT, inferred from the notice) |
| 2026-08-26 16:37 | "hit your monthly spend limit" | Week-to-date ~$3,057: about $158 of usage credits past the weekly limit before the credit cap stopped it |
| 2026-09-08 17:26 | "hit your session limit" | Trailing five hours $481–494, matching the independently computed $486.48 peak window that hour |
Source: The principal's own telemetry (limit-event scan of Claude Code transcripts; 16 account-limit notices and 5 subagent-concurrency notices after excluding false positives).68
Two observations on one account are the best available estimate, not guarantees. Anthropic describes limits as depending on conversation length, features, model and effort, and reserves the right to cap usage in other ways.12 A useful way to hold the numbers: the $200 plan bought roughly 63 times its price in API-equivalent usage at the weekly ceiling (estimate: $12,607 ÷ $200). The plan is excellent value inside its limits and bills at full API rates outside them.
These come from Claude Code's configuration documentation and explain much of the measured mix:
/autocompact, --autocompact or CLAUDE_CODE_AUTO_COMPACT_WINDOW; CLAUDE_CODE_DISABLE_1M_CONTEXT=1 holds sessions to 200K.13high on every model that supports effort; maxEffortLevel caps it.13model: field or CLAUDE_CODE_SUBAGENT_MODEL.14CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS changes it.14rate_limits.five_hour.used_percentage and rate_limits.seven_day.used_percentage, and /usage attributes plan usage to subagents, skills, plugins and MCP servers.1512Anthropic's help article for Claude Code states "Sonnet is the default" while the Claude Code configuration page lists Opus 5 as the Max default.713 The configuration page is the more specific source and matches the measured mix.
The scenario model scales the baseline week (Week 37: 20.5 active hours, 6 active days) linearly:68
Optimized applies: Opus reduced to 15% (or 30%) of cost-weighted main-session work with the rest at Sonnet prices (Sonnet costs 1/2.5 of Opus per token); context resets that remove 33.7% of main-session cache-read cost; Haiku for background agents (1/2 of Sonnet); and, for the five-hour figure only, background concurrency capped at 2 instead of the observed 20.914
| Hours/day × days | Current mix | Runs out | Optimized, Opus 30% | Optimized, Opus 15% | Runs out (optimized 15%) |
|---|---|---|---|---|---|
| 2 × 5 | 52% | — | 27% | 25% | — |
| 2 × 7 | 73% | — | 38% | 35% | — |
| 4 × 5 | 80% | — | 43% | 38% | — |
| 4 × 7 | 112% | day 7 | 60% | 53% | — |
| 6 × 5 | 108% | day 5 | 58% | 51% | — |
| 6 × 7 | 151% | day 5 | 81% | 71% | — |
| 8 × 5 | 136% | day 4 | 73% | 64% | — |
| 8 × 7 | 190% | day 4 | 103% (day 7) | 89% | — |
| 10 × 5 | 164% | day 4 | 89% | 77% | — |
| 10 × 7 | 230% | day 4 | 124% (day 6) | 108% | day 7 |
| Hours per day | Current mix | Optimized, Opus 30% | Optimized, Opus 15% |
|---|---|---|---|
| 2 | 46% ($223) | 18% | 15% |
| 4 | 92% ($446) | 35% | 30% |
| 6 | 137% ($669) | 53% | 45% |
| 8 | 183% ($892) | 70% | 60% |
| 10 | 229% ($1,115) | 88% | 75% |
The five-hour limit depends on hours per day, not days per week, so 5-day and 7-day rows are identical. At current habits, any day above about 4 active hours with parallel agents risks the session limit. After optimization, even a 10-hour day stays under it in the model, but only because background concurrency is capped.
| Hours/day × days | Current: 0% / 15% / 30% unmeasured | Optimized (Opus 15%): 0% / 15% / 30% |
|---|---|---|
| 2 × 7 | 62% / 73% / 88% | 30% / 35% / 42% |
| 4 × 5 | 68% / 80% / 97% | 32% / 38% / 46% |
| 4 × 7 | 95% / 112% / 136% | 45% / 53% / 64% |
| 6 × 5 | 92% / 108% / 131% | 43% / 51% / 62% |
| 6 × 7 | 128% / 151% / 184% | 61% / 71% / 86% |
| 8 × 5 | 116% / 136% / 165% | 54% / 64% / 78% |
| 8 × 7 | 162% / 190% / 231% | 76% / 89% / 109% |
| 10 × 5 | 139% / 164% / 199% | 65% / 77% / 93% |
| 10 × 7 | 195% / 230% / 279% | 91% / 108% / 131% |
Reading the grid. With current habits, the plan comfortably supports about 2 hours a day, and 4 hours a day on a 5-day week. With the optimized habits, it supports up to 10 hours a day on a 5-day week and up to 6–8 hours a day on a 7-day week, depending on how much usage happens outside Claude Code. Ten hours a day, seven days a week exceeds the weekly limit even when optimized — by about $219 a week at 15% unmeasured (estimate: $3,118 − $2,899), or roughly $950 a month of overflow at API rates before any bundle discount.
The model assumes weekly usage scales linearly with hours; that a smaller model finishes the same work in the same number of tokens (it may need more turns); and that context resets save a fixed fraction derived from the top-ten sessions' growth curves rather than a replay. It prices all non-Opus work at Sonnet rates, including the 28% that was Fable — which makes the savings from moving Fable work understated. The five-hour projections scale the shape of observed days. All of these are listed again in §10.
Baseline: 6 active hours × 5 days, current mix, modeled at $2,660 a week ($2,073 main sessions, $588 background agents). Savings are sequential where noted, so they do not simply add.68
| Rank | Lever | How (verified setting or habit) | Modeled weekly saving | Confidence |
|---|---|---|---|---|
| 1 | Sonnet 5 as default; Opus for judgment | /model sonnet or ANTHROPIC_DEFAULT_MODEL; opusplan plans with Opus and executes with Sonnet137 | $570 (Opus → 30%) to $803 (Opus → 15%) = 21–30% | Medium: assumes equal token volume |
| 2 | Compact early, clear between tasks | /autocompact 200k (100K–1M allowed); CLAUDE_CODE_AUTO_COMPACT_WINDOW for launched agents; /clear per task; CLAUDE_CODE_DISABLE_1M_CONTEXT=1 where long context is not needed137 | $507 alone (19%); $310–367 after lever 1 | Low–medium: rough estimate, likely understated given 72.6% cache-read share |
| 3 | Haiku (or Sonnet) for background agents | model: haiku in subagent frontmatter; CLAUDE_CODE_SUBAGENT_MODEL; Explore otherwise inherits the session model1412 | $294 (11%) | Medium for exploration and search; not for judgment tasks |
| 4 | Fable by explicit choice only | Fable is never the account default, but a /model choice persists to later sessions13 | Not in the scenario model. Moving the period's $1,464 of Fable usage to Opus saves ~$732 over the two weeks (estimate: × (1 − 5/10)); to Sonnet ~$1,171 | Medium on arithmetic; quality impact unknown |
| 5 | Cap concurrency | CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS (default 20); a cap in the agent launcher for launched agents14 | $0 weekly; keeps heavy days under the $487 five-hour limit (137% → 45–53% with levers 1–3) | Medium |
| 6 | Lower effort on routine work | /effort; maxEffortLevel managed setting; default is high13 | Not quantified; the bench showed no accuracy gain from higher effort on its task types | Low (small bench) |
| 7 | Schedule across windows | Start agent batches right after a five-hour reset; spread the same weekly hours over 7 days instead of 5 | $0 weekly; roughly halves the heaviest daily concentration | Medium |
| 8 | Usage monitor with alerts | Status line reading rate_limits.five_hour and seven_day percentages; /usage attribution by subagent and MCP server; alert at 70% and 90%1512 | Prevents overflow rather than reducing usage | High (feature is documented) |
| 9 | Overflow routing | Keep the usage-credit monthly cap; bundles at 10–30% off for planned overflow; the platform's own jobs on its API key with Batch at 50% off; bounded public-data jobs to US-hosted open weights after a bench run459 | Up to 30% off overflow dollars (bundles), 50% off batchable API work | Medium |
Combined: levers 1 (Opus → 30%), 2 and 3 take the week to $1,429 (−46%); with Opus at 15% the week is $1,253 (−53%). The input notes labelled the −46% combination as using the 15% Opus lever; recomputation shows it used the 30% lever (§10).
Why lever 2 matters more than it looks. Every turn re-sends the conversation. The per-turn cache-read cost at published prices:9
| Context carried per turn | Opus 5 ($0.50/M) | Sonnet 5 ($0.20/M) | Haiku 4.5 ($0.10/M) | Fable 5 ($1.00/M) |
|---|---|---|---|---|
| 150K tokens | $0.075 | $0.030 | $0.015 | $0.150 |
| 550K tokens | $0.275 | $0.110 | $0.055 | $0.550 |
| 967K tokens (default compaction point) | $0.483 | $0.193 | $0.097 | $0.967 |
A 1,900-message session at 550K tokens of Opus context carries roughly $520 of cache reads (estimate: 1,900 × $0.275) before any output is written. Anthropic's guidance calls /clear "the single most effective lever for both quality and cost."7
| Vendor | Plans and monthly price (verified) | Coding agent on the plan | Automation terms |
|---|---|---|---|
| Anthropic | Max 5x $100 · Max 20x $2001 | Claude Code, shared limits3 | Ordinary individual use; no automated access except via API key or explicit permission1011 |
| OpenAI | Plus $20; Pro $100 and Pro $200 (new sign-ups and upgrades to Pro $200 paused from 2026-09-10)19 | Codex on all ChatGPT plans; one allowance shared with ChatGPT Work and other agent features20 | Not reviewed in this pass (policy pages returned 403) |
| AI Plus $4.99 · AI Pro $19.99 · AI Ultra $99.99 (5× Pro limits) and $199.99 (20×)2223 | Gemini CLI: 1,000 requests/day (Code Assist individual), 1,500 (AI Pro), 2,000 (AI Ultra); API-key free tier 250/day24 | Not reviewed in this pass | |
| xAI | SuperGrok $30 · SuperGrok Plus $100; Heavy listed without a readable price26 | Grok Build, early beta for SuperGrok and X Premium Plus27 | Acceptable Use Policy prohibits "unauthorized automated or non-human means"28 |
| Z.ai (GLM) | GLM Coding Plan "starting at just 18 USD per month," with a five-hour and a weekly limit; works with Claude Code34 | Via Claude Code, Cline, OpenCode | Singapore-based controller (see 6.5)35 |
Where the fetched pages describe limits, they are metered allowances rather than unlimited use: five-hour and weekly limits at Anthropic and Z.ai, daily request quotas for Gemini CLI, and a shared allowance with paid credits for Codex. A second subscription would add a second pool to manage, not remove the need to manage the first.
| Model | Input | Output | Cached input | Notes |
|---|---|---|---|---|
| Claude Opus 5 | 5.00 | 25.00 | 0.50 | Opus 4.5–4.8 same price; Batch 50% off; US-only inference 1.1×9 |
| Claude Sonnet 5 | 2.00 | 10.00 | 0.20 | Introductory price made standard; planned rise to $3/$15 cancelled9 |
| Claude Haiku 4.5 | 1.00 | 5.00 | 0.10 | 9 |
| Claude Fable 5 / Fable 5.1 | 10.00 | 50.00 | 1.00 / 0.25 | Newer tokenizer on 4.7+ models yields ~30% more tokens for the same text9 |
| GPT-6 Astra | 10.00 | 50.00 | 1.00 | Short context; Batch/Flex half16 |
| GPT-5.6 Sol | 4.00 | 20.00 | 0.40 | Promotional pricing at least through 2026-11-2116 |
| GPT-5.6 Terra | 2.00 | 12.00 | 0.20 | 16 |
| GPT-5.6 Luna | 0.20 | 1.20 | 0.02 | 16 |
| GPT-5 / GPT-5 mini | 1.25 / 0.25 | 10.00 / 2.00 | — | Models used in the bench1718 |
| Gemini 3.1 Pro Preview | 2.00 | 12.00 | 0.20 | Prompts ≤200K tokens21 |
| Gemini 3.8 Flash | 0.75 | 3.75 | 0.075 | Through 2026-12-31; $1.50 / $7.50 from 2027-01-0121 |
| Gemini 3.5 Flash | 1.50 | 9.00 | 0.15 | Model used in the bench21 |
| Grok 4.6 | 2.00 | 6.00 | — | 500K context25 |
| Model | Vendor | Input (cache miss) | Output | Cache hit | Notes |
|---|---|---|---|---|---|
DeepSeek V4.1 Flash (deepseek-flash) | DeepSeek | 0.30 peak / 0.15 off-peak | 1.20 / 0.60 | 0.006 / 0.003 | Peak 01:00–04:00 and 06:00–10:00 UTC weekdays; 1M context; Anthropic-format endpoint29 |
| DeepSeek V4 Pro | DeepSeek | 1.32 / 0.66 | 3.96 / 1.98 | 0.044 / 0.022 | Service continued past 2026-09-1429 |
| Kimi K3 | Moonshot AI | 3.00 | 15.00 | 0.30 | ~1M context32 |
| Kimi K2.7 Code | Moonshot AI | 0.95 | 4.00 | 0.19 | 32 |
| Kimi K2.6 | Moonshot AI | 0.95 | 4.00 | 0.16 | 32 |
| GLM-5.3 | Z.ai (Zhipu) | 1.40 | 4.40 | 0.26 | 33 |
| GLM-5.3-Flash | Z.ai | 0.15 | 0.50 | 0.03 | GLM-4.7-Flash listed free33 |
| Qwen3.8-Max | Alibaba Cloud Model Studio | 2.00 | 6.00 | — | Singapore deployment, "International" scope; the page lists Global, US and Chinese-mainland scopes for other models36 |
| MiniMax-M3 | MiniMax | 0.30 | 1.20 | 0.06 | After a listed "permanent 50% off" ($0.60 / $2.40 list), ≤512K input37 |
| Doubao / Seed | ByteDance (BytePlus ModelArk) | — | — | — | Prices not readable in the fetched pages; not used38 |
Against Sonnet 5 ($2 / $10), DeepSeek V4.1 Flash at peak is 15% of the input price and 12% of the output price; GLM-5.3 is 70% and 44%; Kimi K3 is above Sonnet 5 on both. The large savings are in the Flash-class models, not the Chinese flagships.
Figures come from each host's own price page where readable, and otherwise from OpenRouter's public endpoints API, which lists the price and quantization of every upstream serving a model on OpenRouter.40
| Model | Fireworks44 | DeepInfra46 | Other listed routes (OpenRouter endpoints API)40 |
|---|---|---|---|
| DeepSeek V4 Pro | 1.32 / 3.96 | 1.30 / 2.60 | Azure (US) 1.91 / 3.83 · Baseten 1.74 / 3.48 (fp4) · lowest listed: StreamLake 0.95 / 1.89 and Baidu 0.95 / 1.90 (fp8) |
| DeepSeek V4.1 Flash | 0.22 / 0.66 | 0.20 / 0.60 (fp8, via OpenRouter) | Together 0.30 / 1.20 · DeepSeek first-party 0.15 / 0.60 |
| Kimi K2.6 | 0.95 / 4.00 | 0.75 / 3.50 | Crusoe 0.70 / 3.50 (bf16) · lowest: Inceptron 0.46 / 2.61 (int4) |
| Kimi K3 | 3.00 / 15.00; "K3 US" 3.30 / 16.50 | 2.85 / 14.25 | Together 3.00 / 15.00 · Alibaba 3.45 / 17.25 |
| GLM-5.3 | 1.40 / 4.40 | 0.90 / 3.00 (fp4, via OpenRouter) | Together 1.40 / 4.40 · Z.AI 1.40 / 4.40 |
| gpt-oss-120b | — | 0.037 / 0.17 (bf16, via OpenRouter) | Groq, Together, Amazon Bedrock 0.15 / 0.60 · Cerebras 0.35 / 0.75 |
| MiniMax-M3 | — | 0.28 / 1.10 (fp8, via OpenRouter) | Together 0.30 / 1.20 · MiniMax 0.30 / 1.20 |
Hyperscalers also carry these weights: Amazon Bedrock lists DeepSeek (R1, V3.1, V3.2), MiniMax and Qwen models48; Microsoft lists DeepSeek-V4-Pro and DeepSeek-V4-Flash among models sold directly by Azure49; Vercel AI Gateway, on which Salud's apps already deploy, passes provider list prices through "no markup and no platform fee"50 (DeepSeek V4 Flash from $0.06 / $0.18).
Three observations from the endpoint data:
First-party Chinese APIs. DeepSeek's privacy policy names Hangzhou DeepSeek Artificial Intelligence Co., Ltd. as controller, states that it collects, processes and stores personal data in the People's Republic of China, and lists training and improving its models among the uses (with an opt-out right).30 Its Open Platform terms are governed by the laws of mainland China.31 Z.ai's privacy policy names a Singapore company as controller and says data is "generally processed in Singapore."35 The input notes' claim that Z.ai excludes API content from training by default was not found in that policy and is not repeated here.
PRC law, from the National People's Congress texts. National Intelligence Law Article 7 (2017 text): all organizations and citizens shall support, assist and cooperate with national intelligence work in accordance with law, and keep its secrets (our translation).66 Data Security Law Article 35: when public security or state security organs lawfully obtain data for national security or criminal investigation, the organizations and individuals concerned shall cooperate (our translation).65 The Personal Information Protection Law, in force since 2021-11-01, sets conditions for sending personal information outside China (Article 38).64 Authoritative text is Chinese; the NPC hosts an English PIPL page.67 This is a description of statutes, not legal advice; counsel should read it before any decision that depends on it.
Routers. OpenRouter can exclude providers that may store data (data_collection: "deny"), restrict routing to zero-data-retention endpoints (zdr), and allow or ignore named providers per request; an account setting excludes providers that train on prompts; enterprise accounts can pin processing to US or EU regions.4142 Its documentation also says it "does not have routing rules that change based on data retention policies of providers" — the zero-data-retention flag is the control, not a retention-window filter as the input notes described.42 Per-key credit limits with a configurable reset cap spend on a key; an exhausted limit returns an error rather than a bill.43
Hosts. Fireworks states it "does not log or store prompt or generation data for open models, without explicit user opt-in."45 DeepInfra's privacy notice states its security measures "comply with SOC 2 and ISO 27001 standards."47 Together AI, Groq and Baseten publish trust centers whose content did not render for our fetcher.515253 No SOC 2 report or ISO certificate was inspected for any vendor in this report; a trust-center page is a claim, and the report behind it is the evidence. Novita's jurisdiction (a Singapore entity with a San Francisco address, per the input notes) is unresolved; it appears as an upstream on OpenRouter routes above and should be excluded until resolved.
Anthropic. US-only inference on the API is available at a 1.1× price multiplier.9
| Body | Instrument | Relevance |
|---|---|---|
| U.S. House of Representatives | H.R. 1121, "No DeepSeek on Government Devices Act," introduced in the 119th Congress and referred to the Committee on Oversight and Government Reform58 | Signals federal posture; later status not verified (congress.gov returned 403) |
| State of New York | Governor Hochul's 2025-02-10 ban on DeepSeek on ITS-managed government devices and networks59 | State-level precedent relevant to public-sector clients |
| NIST / CAISI | DeepSeek V4 Pro evaluation (May 2026) and DeepSeek models evaluation (Sept 2025)5657 | Capability and security evidence |
| NIST | AI Risk Management Framework 1.0, under revision per the White House AI Action Plan54 | Voluntary governance frame for vendor selection |
| NIST | AI 600-1, Generative AI Profile (2024-07-26)55 | GenAI-specific risk actions |
| ISO/IEC | ISO/IEC 42001, AI management systems60 | Certification to ask vendors about (page returned 403; link only) |
| AICPA | SOC suite, including SOC 261 | The report to request, not the badge |
| GSA FedRAMP | FedRAMP and its Marketplace62 | Relevant only if work touches federal systems |
| European Union | Regulation (EU) 2024/1689, the AI Act, of 13 June 202463 | Relevant to EU-routed processing |
| PRC National People's Congress | PIPL, Data Security Law, National Intelligence Law646566 | Legal basis of the jurisdiction concern |
| Tier | Where | Task classes | Guardrails |
|---|---|---|---|
| 0 — Code, no model | Deterministic functions and tables | domain reference shape weights and properties, unit conversions, arithmetic roll-ups | The bench's aisc and units suites are deterministic; a model is the wrong tool |
| 1 — Claude Max plan, interactive | Claude Code, chat, Design on the principal's own login | Architecture, cross-cutting changes, hard debugging, review of drafts, anything needing judgment; all never-offload work below | Sonnet 5 default; Opus by choice; Fable by explicit choice; compaction window set; monitor at 70% / 90% |
| 2 — Claude API on the platform's key | internal agent services with a Console spend limit | Cockpit chat (Sonnet 5), memory classifier (Haiku 4.5), oversight (Sonnet 5 low), scheduled research, batchable jobs via the Batch API at 50% off9 | Separate key per service; no ANTHROPIC_API_KEY on the principal's interactive shell, where it would silently switch Claude Code to API billing3 |
| 3 — US-hosted open weights | Fireworks or DeepInfra direct, or OpenRouter with ignore on non-US and unresolved providers, data_collection: "deny", zdr: true, and a per-key credit limit4143 | Public or non-confidential bulk work only: tagging public catalogue data, summarizing public product literature, drafting public marketing copy, chores on open-source code | Only after the bench's second round clears the task type; nothing from the list below |
| Never offload | Stays on the current Anthropic arrangement; never to a first-party Chinese API, a free tier, or an unvetted host | Financial records · legal and counsel files · client documents and client documents · pricing, proposals and costs · credentials, keys, .env values · patent and trade-secret material | Existing internal policies already restrict where these travel; this list adds no exceptions |
| # | Recommendation | Owner | Cost | Expected effect |
|---|---|---|---|---|
| 1 | Set Claude Code's default model to Sonnet 5; use opusplan or pick Opus deliberately for design and hard debugging13 | Principal (user settings); platform (launcher defaults) | $0 | −21% to −30% of weekly usage (modeled) |
| 2 | Set /autocompact 200k (or 300K) in user settings and CLAUDE_CODE_AUTO_COMPACT_WINDOW for launched agents; /clear between unrelated tasks13 | Principal; platform | $0 | −12% to −19% further (modeled, rough) |
| 3 | Give every custom subagent an explicit model: — haiku for search and exploration, sonnet for building; set CLAUDE_CODE_SUBAGENT_MODEL in launched sessions14 | Platform | $0 | −11% (modeled) |
| 4 | Set CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS=4 normally and 2 on heavy days; add the same cap to the cockpit launcher14 | Platform | $0 | Keeps heavy days under the five-hour limit |
| 5 | Make Fable opt-in per task; review why it carried 28% of usage6 | Principal | $0 | Up to ~14% of total usage if that work moves to Opus (estimate) |
| 6 | Build a status-line and internal alert on rate_limits.five_hour and seven_day at 70% and 90%; add plan-usage percentages to the weekly telemetry report15 | Platform | A few build hours | Early warning before overflow |
| 7 | Keep the usage-credit monthly cap at a figure the principal chooses; buy bundles only for planned overflow (10–30% off)45 | Principal | Capped by choice | Bounds the extra-usage bill |
| 8 | Fix the internal pricing table: Opus 4.8 at $5/$25, add Fable 5/5.1 and Sonnet 5 at published prices; reconcile Fable attribution between the telemetry report and the usage pass971 | Platform | $0 | Correct spend reporting (the internal agent platform key ~$84, not $103, for 9/8–9/15) |
| 9 | Move cockpit chat off Opus 4.8 to Sonnet 5; keep the classifier on Haiku 4.5; run oversight on Sonnet 5 low (the bench's sufficient configuration)70 | Platform | $0 | Lower API-key spend |
| 10 | Replace model calls for domain reference lookups and unit conversions with deterministic code | Estimating | Build hours | Removes a 94%-accurate model from a 100%-knowable task |
| 11 | Fix memory-classify labels or task definition before choosing its model | Platform | Build hours | A valid bench for the classifier |
| 12 | Bench round two under the same $35 cap: add OpenRouter (PRC-company and unresolved providers ignored, zdr on), a paid Google key and xAI credit, plus one agentic multi-step suite | the platform team runs; the principal approves and creates any accounts | ≤$35 plus prepaid credits | Evidence for Tier 3 before any production move |
| 13 | If OpenRouter is adopted: per-key credit limit with a reset period, data_collection: "deny", zdr: true, provider ignore list, no business data4143 | Principal / security | $0 | No surprise bill; no unintended jurisdiction |
| 14 | Do not add ChatGPT, Google AI or xAI subscriptions now; revisit only if the optimized habits still exceed the weekly limit at the hours the principal chooses | Principal | $0 now | Avoids a second pool to manage |
| 15 | No first-party Chinese API for any business data; counsel review before any exception | Principal / counsel | Counsel time if raised | Keeps the never-offload list intact |
| 16 | If many concurrent agents remain the primary workflow, ask Anthropic whether Max remains the right arrangement or whether Team Premium, Enterprise or API billing fits better, since advertised Max limits assume ordinary individual use10 | Principal | A call | Removes ambiguity on plan fit |
high effort and twenty concurrent subagents are each reasonable alone. Together they turn a $200 plan with roughly $12,600 of monthly capacity into one that runs out on day 4 or 5 at 6–10 active hours a day.~/.claude/projects/**/*.jsonl, including subagent files), deduplicated by message, with token counts priced at Anthropic's published per-token rates (Opus $5/$25, Sonnet $2/$10, Haiku $1/$5, Fable $10/$50; cache reads at 0.1× input and writes at 1.25×).968 Active hours cluster the principal's own messages. Limit events were found by matching Claude Code's own limit notices and excluding 1,588 generic word matches and 12 self-referential research matches. The internal telemetry report (2026-09-08 to 2026-09-15) and the model bench are separate internal sources.6970| Claim in the input notes | Disposition |
|---|---|
| Third-party estimates of Sonnet-hours and Opus-hours per week for Max tiers; Opus-to-Sonnet auto-fallback at 20% / 50% | Dropped — not from Anthropic |
| "97,361 unique priced messages (25,512 since 9/1)" | Corrected: 25,512 priced since 9/1; 97,361 is undated |
| Background share 32.8% | Corrected to 33.2% |
| "Combined (1+3+4) −46%" described with the Opus → 15% lever | Corrected: −46% uses Opus → 30%; Opus → 15% gives −53% |
| "Current (all-Opus)" scenario label | Corrected: current mix is 66.6% Opus, 28.0% Fable |
| Fable price UNVERIFIED | Verified $10 / $50 (Fable 5 cache reads $1.00; Fable 5.1 $0.25) |
| OpenAI flagship ~$10/$50, mid-tier ranges, gpt-5-mini (secondary) | Replaced by verified OpenAI pages |
| ChatGPT Go $8 and Business seat prices; Codex headless flag | Dropped — not verified |
| Gemini CLI free tier "cut June 2026" | Dropped — current docs list a free tier |
| xAI Heavy $300; Grok 4.3 $1.25 / $2.50 | Dropped — not in page text |
| Artificial Analysis index scores, Kimi K3 93.4% SWE-bench, "highest-ranked open-weights" claims | Dropped — aggregator sources |
| Qwen Coder-Plus tier pricing, Qwen Flash via aggregator | Dropped; Qwen3.8-Max verified instead |
| GLM Coding Plan Pro ~$72 and Max ~$160; MiniMax ¥9.9 plan | Dropped — only "starting at 18 USD" verified |
| Doubao pricing; PRC phone verification requirement | Dropped — not readable / not verified |
| Model licences (DeepSeek MIT, Qwen Apache 2.0, Kimi "Modified MIT" and revenue threshold, GLM licence) | Unverified — licence files not pulled |
| Z.ai excludes API content from training by default | Dropped — not in the privacy policy fetched |
| Federal Commerce and Navy DeepSeek bans; Texas, Virginia (order number conflicts between notes and search), Iowa, Kansas, South Dakota, Nebraska bans | Dropped — primary text not readable or not found; New York verified |
| H.R. 1121 current status | Unverified — introduced text verified only |
| BIS export-control rules (Jan 2026, May 31 2026 guidance); OMB M-24-10 | Dropped — secondary sources only |
| OpenRouter FedRAMP and HIPAA claims; Together, Groq, Baseten, Cerebras, SambaNova, Hyperbolic, Nvidia NIM, Cloudflare compliance claims | Dropped or link-only — trust pages did not render; certificates not inspected |
| OpenRouter "data-retention filter" | Contradicted by OpenRouter docs; zdr flag described instead |
| "Quantization rarely disclosed" | Contradicted for OpenRouter endpoints, which label it |
| Fireworks spend-gate tiers; Together budget alerts | Dropped — not verified |
| AWS Bedrock and Vertex per-token prices; Cloudflare neurons pricing; NIM free tier | Dropped — not readable |
| Novita jurisdiction | Unresolved |
| Local 8 GB VRAM throughput figures | Dropped — secondary source |
| an internal notes file (named as an input) | Not present in the input folder; not used |
All external sources accessed 2026-09-15.