Skip to content

feat(pricing): support OpenAI two-stage pricing and add the gpt-5.6 family - #1414

Merged
ryoppippi merged 4 commits into
mainfrom
feat/gpt-5-6-two-stage-pricing
Jul 9, 2026
Merged

feat(pricing): support OpenAI two-stage pricing and add the gpt-5.6 family#1414
ryoppippi merged 4 commits into
mainfrom
feat/gpt-5-6-two-stage-pricing

Conversation

@ryoppippi

@ryoppippi ryoppippi commented Jul 9, 2026

Copy link
Copy Markdown
Member

Summary

OpenAI introduced two-stage (short/long context) pricing with the gpt-5.6 family: requests with more than 272K input tokens are billed at higher long-context rates. The same tier structure also applies to gpt-5.5, gpt-5.5-pro, gpt-5.4, and gpt-5.4-pro on the current pricing page. This PR adds the gpt-5.6-sol/terra/luna models to the built-in pricing table and makes Codex reports bill long-context requests at the correct rates.

What Changed

  • Per-model tier threshold: Pricing gains long_context_threshold; the previous hardcoded 200K boundary stays the default for LiteLLM *_above_200k_tokens data, while the OpenAI models switch tiers above 272K input tokens.
  • New built-in entries: gpt-5.6-sol ($5/$30, long $10/$45), gpt-5.6-terra ($2.50/$15, long $5/$22.50), and gpt-5.6-luna ($1/$6, long $2/$9) per 1M input/output tokens, including cached-input and cache-write rates for both tiers.
  • Long-context overlay: tier rates that upstream sources do not publish are re-applied after every pricing load via builtin_long_context_rates. A live LiteLLM refresh replaces whole entries (LiteLLM currently lists gpt-5.5/gpt-5.4 with flat rates only), so tier rates set directly on built-in entries would be silently dropped once a refresh succeeds. Upstream tier data wins over the overlay when it exists, and date-pinned keys such as gpt-5.5-2026-04-23 share their base model's rates.
  • Codex per-request tiering: OpenAI decides the tier per request and bills the whole request (input, cached input, and output) at long-context rates. Codex costs are computed from per-model sums, so CodexModelUsage now tracks the long-context portion while events are aggregated (each token_count event is one request), and calculate_codex_model_cost prices the short and long buckets independently. Models without tier rates are priced exactly as before.

Notes

  • Report JSON shape and table output are unchanged; only costUSD for long-context usage differs.
  • The generic cost path (Claude Code, pi, and other non-Codex adapters) now mirrors the Codex per-request tiering: models with a per-model long_context_threshold bill the whole request at the long-context rates once input exceeds the threshold, so it is no longer a marginal breakpoint. LiteLLM *_above_200k_tokens data (no per-model threshold) keeps its marginal above-threshold semantics at the 200K default.
  • Codex aggregation derives each request's tier from the matched model's threshold (long_context_split_threshold) instead of a single hardcoded 272K constant, so any model whose long-context boundary is not 272K is split at its own threshold.
  • gpt-5.6 context limits mirror the 1,050,000-token window of gpt-5.5/gpt-5.4 until upstream metadata lands.
  • pricingOverrides cannot override the tier threshold yet; that can be added later if someone needs it.

Testing

  • just test (Rust workspace + Node) passes: 343 + 61 + 16 + 11 crate tests, 0 failures.
  • New tests: embedded gpt-5.6 rates and 272K threshold, overlay survival across a simulated LiteLLM refresh (and deferring to upstream tier data), date-suffix stripping, per-request long-context split during Codex aggregation (fixture-backed), two-stage Codex cost math, and flat-model equivalence.
  • cargo clippy -p ccusage --tests and just fmt are clean.

View with Codesmith
Need help on this PR? Tag /codesmith with what you need. Autofix is enabled.


Summary by cubic

Adds support for OpenAI two-stage pricing (short vs. long context) and the gpt-5.6 family. Codex now bills any request with more than 272K input tokens at long-context rates while keeping report shapes unchanged.

  • New Features
    • Per-model tier threshold in pricing. Default stays 200K; OpenAI models use 272K.
    • Added gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna with embedded short-context and cache rates.
    • Filled long-context tier rates via a built-in overlay applied after every pricing load. Upstream tier data wins, and date-pinned keys inherit their base model’s rates.
    • Codex aggregation tracks long-context usage per request and prices short and long buckets separately. Models without tier rates are unchanged.
    • Set a 1,050,000-token context window for the gpt-5.6 family. Only costUSD changes for long-context usage; reports remain the same.

Written for commit 7d66785. Summary will update on new commits.

Review in cubic

Summary by CodeRabbit

  • New Features

    • Added long-context token tracking and pricing support, with separate accounting for long-context input/cached input/output.
    • Cost calculations now apply long-context/two-stage rates using model-specific thresholds.
  • Bug Fixes

    • Improved aggregation to correctly preserve long-context usage when combining multiple requests.
    • Updated billing logic to better match tiered long-context behavior across supported model pricing sources.
  • Tests

    • Expanded unit coverage for long-context split/threshold handling, tier switching, and aggregation correctness.

ryoppippi added 3 commits July 9, 2026 19:10
OpenAI introduced two-stage (short/long context) pricing with gpt-5.6:
requests with more than 272K input tokens are billed at higher
long-context rates. The same tier also applies to gpt-5.5, gpt-5.5-pro,
gpt-5.4, and gpt-5.4-pro on the current pricing page.

The existing tier support hardcoded the LiteLLM 200K boundary, so
Pricing gains a per-model long_context_threshold (defaulting to 200K
for LiteLLM *_above_200k_tokens data) and tiered_cost takes the
threshold as a parameter.

New built-in entries cover gpt-5.6-sol, gpt-5.6-terra, and
gpt-5.6-luna, including their cache-write rates. Long-context tier
rates live in a builtin_long_context_rates overlay that is re-applied
after every pricing load: a live LiteLLM refresh replaces whole
entries, and LiteLLM currently publishes these models with flat rates
only, so tier rates set directly on built-in entries would be silently
dropped whenever a refresh succeeds. Entries that already carry tier
rates are left untouched so upstream data wins once it exists.
Date-pinned keys such as gpt-5.5-2026-04-23 share their base model's
overlay rates.

The gpt-5.6 context limits mirror the 1,050,000-token window of the
other long-context GPT-5 flagship models until upstream data lands.
OpenAI decides the pricing tier per request: once a request's input
exceeds 272K tokens, every token of that request (input, cached input,
and output) is billed at the long-context rates. Codex cost calculation
runs on per-model sums aggregated across many requests, so the tier
cannot be recovered from the totals afterwards.

CodexModelUsage now tracks the portion of tokens that came from
long-context requests. The split is recorded while token_count events
are aggregated, where each event still represents a single request,
and merged across parallel shards like the other counters.

calculate_codex_model_cost prices the aggregated usage as two
independent buckets: the short bucket at the flat rates and the long
bucket at the *_above_200k rates, falling back to the flat rates for
models without a long-context tier so their costs are unchanged. The
existing fast-speed multiplier applies to both buckets.

Report JSON and table output are unchanged; only costUSD values for
long-context requests differ.
Codex review suggested filling missing tier fields independently when a
refreshed LiteLLM entry carries partial *_above_200k_tokens data. That
would mix rates that assume the 200K LiteLLM boundary with built-in
rates that assume the OpenAI 272K boundary under a single per-model
threshold, mispricing both tiers, so the overlay defers to upstream
entirely once any tier rate exists. Record that rationale next to the
check.
Copilot AI review requested due to automatic review settings July 9, 2026 18:19
@cursor

cursor Bot commented Jul 9, 2026

Copy link
Copy Markdown

Bugbot is not enabled for this team, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Jul 9, 2026

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
ccusage-guide b65043a Commit Preview URL

Branch Preview URL
Jul 09 2026, 06:41 PM

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Jul 9, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

This PR adds long-context billing support for Codex. Pricing now carries a per-model long-context threshold and overlayed tier rates, cost calculation splits long and short buckets using that threshold, and Codex aggregation tracks long-context token totals per request.

Changes

Codex long-context billing

Layer / File(s) Summary
Data shape for long-context token tracking
rust/crates/ccusage/src/types.rs
Adds long_context_input_tokens, long_context_cached_input_tokens, and long_context_output_tokens to CodexModelUsage.
Pricing schema, thresholds, and built-in long-context overlay
rust/crates/ccusage/src/pricing.rs
Adds long_context_threshold and threshold constants, overlays built-in long-context rates, preserves thresholds through overrides, and updates related tests.
Threshold-aware tiered cost calculation
rust/crates/ccusage/src/cost.rs, rust/crates/ccusage/src/main.rs
Adds a threshold parameter to tiered_cost, threads long-context threshold selection through pricing-based cost calculation, and updates tiering tests.
Codex model cost split into long/short buckets
rust/crates/ccusage/src/adapter/codex/report.rs, rust/crates/ccusage/src/adapter/codex/mod.rs
Splits Codex usage into long and short token portions for billing and expands long-context cost tests.
Per-request long-context aggregation
rust/crates/ccusage/src/adapter/codex/aggregate.rs
Accumulates and merges long-context token counters per model, with a test covering above-threshold vs below-threshold requests.

Estimated code review effort: 4 (Complex) | ~55 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CodexEvent
  participant accumulate_codex_event_into_group
  participant CodexModelUsage
  participant PricingMap
  participant calculate_codex_model_cost

  CodexEvent->>accumulate_codex_event_into_group: token usage event
  accumulate_codex_event_into_group->>CodexModelUsage: update totals and long_context_* counters
  calculate_codex_model_cost->>PricingMap: read pricing + long-context threshold
  calculate_codex_model_cost->>calculate_codex_model_cost: split usage into long and short buckets
  calculate_codex_model_cost->>calculate_codex_model_cost: compute weighted cost and apply speed multiplier
Loading

Possibly related PRs

  • ccusage/ccusage#651: Both PRs implement two-bucket long-context pricing by splitting token counts around a threshold and applying separate above-threshold rates.
  • ccusage/ccusage#1313: Both PRs modify Codex per-model token aggregation in adapter/codex/aggregate.rs.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly captures the main change: OpenAI two-stage pricing support and adding the gpt-5.6 family.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/gpt-5-6-two-stage-pricing

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ryoppippi

Copy link
Copy Markdown
Member Author

Pre-PR Codex review (codex exec review --base main, gpt-5.5) reported two P2 findings; both are deliberate decisions:

  1. Generic cost path keeps marginal above-threshold tiering. Only the Codex adapter has verified per-request token semantics (input_tokens includes cached input), which the whole-request tier decision needs. Other adapters' token mappings were not audited in this PR, so their behavior is unchanged apart from the per-model threshold. Called out in the PR notes.
  2. The long-context overlay defers to upstream entirely once any tier rate exists. Filling missing fields independently would mix upstream rates that assume the 200K LiteLLM boundary with built-in rates that assume the OpenAI 272K boundary under one per-model threshold, mispricing both tiers. Rationale recorded in the code comment (7d66785).

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
rust/crates/ccusage/src/adapter/codex/report.rs (1)

128-178: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Correct implementation of long/short bucket billing, verified against test expectations.

Optionally, consider extracting the long/short rate derivation and bucket sizing into a small helper (e.g. fn long_context_buckets(usage, pricing) -> (u64, u64, u64)) to keep calculate_codex_model_cost focused on the final cost formula, since the function now does rate derivation, clamping, and cost aggregation all in one place.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust/crates/ccusage/src/adapter/codex/report.rs` around lines 128 - 178, The
long/short billing logic in calculate_codex_model_cost is correct, but the
function is doing rate derivation, clamping, and aggregation all at once, making
it hard to follow and maintain. Extract the long-context bucket sizing and rate
selection into a small helper (for example, a helper near
calculate_codex_model_cost that returns the long/short bucket values or derived
rates) and keep calculate_codex_model_cost focused on the final pricing formula.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@rust/crates/ccusage/src/adapter/codex/report.rs`:
- Around line 128-178: The long/short billing logic in
calculate_codex_model_cost is correct, but the function is doing rate
derivation, clamping, and aggregation all at once, making it hard to follow and
maintain. Extract the long-context bucket sizing and rate selection into a small
helper (for example, a helper near calculate_codex_model_cost that returns the
long/short bucket values or derived rates) and keep calculate_codex_model_cost
focused on the final pricing formula.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: aa14520e-ebc0-4421-8ab6-c2604c4eb214

📥 Commits

Reviewing files that changed from the base of the PR and between 0c1c658 and 7d66785.

📒 Files selected for processing (7)
  • rust/crates/ccusage/src/adapter/codex/aggregate.rs
  • rust/crates/ccusage/src/adapter/codex/mod.rs
  • rust/crates/ccusage/src/adapter/codex/report.rs
  • rust/crates/ccusage/src/cost.rs
  • rust/crates/ccusage/src/main.rs
  • rust/crates/ccusage/src/pricing.rs
  • rust/crates/ccusage/src/types.rs

@pkg-pr-new

pkg-pr-new Bot commented Jul 9, 2026

Copy link
Copy Markdown

Open in StackBlitz

ccusage

npx https://pkg.pr.new/ccusage@1414

@ccusage/ccusage-darwin-arm64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-arm64@1414

@ccusage/ccusage-darwin-x64

npx https://pkg.pr.new/@ccusage/ccusage-darwin-x64@1414

@ccusage/ccusage-linux-arm64

npx https://pkg.pr.new/@ccusage/ccusage-linux-arm64@1414

@ccusage/ccusage-linux-x64

npx https://pkg.pr.new/@ccusage/ccusage-linux-x64@1414

@ccusage/ccusage-win32-x64

npx https://pkg.pr.new/@ccusage/ccusage-win32-x64@1414

commit: 7d66785

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 7 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread rust/crates/ccusage/src/cost.rs
Comment thread rust/crates/ccusage/src/adapter/codex/aggregate.rs Outdated
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 7d66785e5a0d
Base SHA: 0c1c65814a08

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 346.1ms 2.91 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 308.3ms 3.27 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 102.4ms 9.83 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 80.4ms 12.52 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 27.5ms 4.3ms 6.41x 54.00 MiB 10.20 MiB 0.19x 0.06 MiB/s 0.36 MiB/s
claude session --offline --json 0.00 MiB 24.2ms 2.5ms 9.85x 54.00 MiB 10.19 MiB 0.19x 0.06 MiB/s 0.63 MiB/s
codex daily --offline --json 0.00 MiB 24.5ms 2.2ms 10.99x 53.75 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.38 MiB/s
codex session --offline --json 0.00 MiB 22.9ms 2.2ms 10.43x 53.75 MiB 8.18 MiB 0.15x 0.04 MiB/s 0.39 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 330.8ms 312.6ms 1.06x 950.33 MiB 962.33 MiB 1.01x 3.04 GiB/s 3.22 GiB/s
codex --offline --json 1.01 GiB 105.7ms 82.3ms 1.28x 425.03 MiB 433.04 MiB 1.02x 9.53 GiB/s 12.23 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.54 KiB 18.54 KiB +0.00 KiB 1.00x
installed native package binary 4061.50 KiB 4064.31 KiB +2.81 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: 7d66785e5a0d
Base SHA: 0c1c65814a08

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 330.0ms 3.05 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 337.1ms 2.99 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 106.9ms 9.42 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 82.1ms 12.26 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 28.7ms 25.4ms 1.13x 53.50 MiB 53.75 MiB 1.00x 0.05 MiB/s 0.06 MiB/s
claude session --offline --json 0.00 MiB 26.4ms 26.1ms 1.01x 53.75 MiB 53.75 MiB 1.00x 0.06 MiB/s 0.06 MiB/s
codex daily --offline --json 0.00 MiB 23.0ms 25.0ms 0.92x 53.50 MiB 53.50 MiB 1.00x 0.04 MiB/s 0.03 MiB/s
codex session --offline --json 0.00 MiB 24.1ms 23.0ms 1.05x 53.75 MiB 53.75 MiB 1.00x 0.04 MiB/s 0.04 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 347.3ms 333.7ms 1.04x 962.34 MiB 940.33 MiB 0.98x 2.90 GiB/s 3.02 GiB/s
codex --offline --json 1.01 GiB 103.6ms 106.7ms 0.97x 421.03 MiB 411.03 MiB 0.98x 9.72 GiB/s 9.44 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.54 KiB 18.54 KiB +0.00 KiB 1.00x
installed native package binary 4061.50 KiB 4064.31 KiB +2.81 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 3 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="rust/crates/ccusage/src/pricing.rs">

<violation number="1" location="rust/crates/ccusage/src/pricing.rs:1356">
P2: Codex costs can undercount requests between 200K and 272K input tokens when upstream publishes `*_above_200k_tokens` for a built-in OpenAI model, because `long_context_split_threshold` returns the built-in 272K boundary without checking whether the loaded `Pricing` entry actually deferred to upstream tier data. Consider deriving the split threshold from the resolved `Pricing` entry (`long_context_threshold.unwrap_or(DEFAULT_LONG_CONTEXT_THRESHOLD_TOKENS)`) at aggregation time instead of from `builtin_long_context_rates` alone.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

/// `Pricing::long_context_threshold`, and falls back to the default 200K
/// boundary used for LiteLLM `*_above_200k_tokens` data.
pub(crate) fn long_context_split_threshold(model: &str) -> u64 {
builtin_long_context_rates(model_without_date_suffix(model))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Codex costs can undercount requests between 200K and 272K input tokens when upstream publishes *_above_200k_tokens for a built-in OpenAI model, because long_context_split_threshold returns the built-in 272K boundary without checking whether the loaded Pricing entry actually deferred to upstream tier data. Consider deriving the split threshold from the resolved Pricing entry (long_context_threshold.unwrap_or(DEFAULT_LONG_CONTEXT_THRESHOLD_TOKENS)) at aggregation time instead of from builtin_long_context_rates alone.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At rust/crates/ccusage/src/pricing.rs, line 1356:

<comment>Codex costs can undercount requests between 200K and 272K input tokens when upstream publishes `*_above_200k_tokens` for a built-in OpenAI model, because `long_context_split_threshold` returns the built-in 272K boundary without checking whether the loaded `Pricing` entry actually deferred to upstream tier data. Consider deriving the split threshold from the resolved `Pricing` entry (`long_context_threshold.unwrap_or(DEFAULT_LONG_CONTEXT_THRESHOLD_TOKENS)`) at aggregation time instead of from `builtin_long_context_rates` alone.</comment>

<file context>
@@ -1345,6 +1345,19 @@ fn builtin_long_context_rates(base_model: &str) -> Option<LongContextRates> {
+/// `Pricing::long_context_threshold`, and falls back to the default 200K
+/// boundary used for LiteLLM `*_above_200k_tokens` data.
+pub(crate) fn long_context_split_threshold(model: &str) -> u64 {
+    builtin_long_context_rates(model_without_date_suffix(model))
+        .map(|rates| rates.threshold)
+        .unwrap_or(DEFAULT_LONG_CONTEXT_THRESHOLD_TOKENS)
</file context>

@ryoppippi
ryoppippi merged commit 726ecb3 into main Jul 9, 2026
22 of 23 checks passed
@ryoppippi
ryoppippi deleted the feat/gpt-5-6-two-stage-pricing branch July 9, 2026 18:53
@github-actions github-actions Bot mentioned this pull request Jul 9, 2026
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: b65043a02979
Base SHA: 0c1c65814a08

This compares the PR package against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 343.0ms 2.94 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 333.5ms 3.02 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 110.5ms 9.11 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 83.2ms 12.11 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 26.7ms 26.7ms 1.00x 53.75 MiB 53.75 MiB 1.00x 0.06 MiB/s 0.06 MiB/s
claude session --offline --json 0.00 MiB 23.6ms 26.9ms 0.88x 53.50 MiB 53.75 MiB 1.00x 0.07 MiB/s 0.06 MiB/s
codex daily --offline --json 0.00 MiB 25.9ms 23.5ms 1.10x 53.75 MiB 54.00 MiB 1.00x 0.03 MiB/s 0.04 MiB/s
codex session --offline --json 0.00 MiB 23.0ms 24.5ms 0.94x 53.75 MiB 53.75 MiB 1.00x 0.04 MiB/s 0.03 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published ccusage package from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 341.9ms 351.9ms 0.97x 956.33 MiB 950.33 MiB 0.99x 2.94 GiB/s 2.86 GiB/s
codex --offline --json 1.01 GiB 105.9ms 111.0ms 0.95x 415.03 MiB 425.04 MiB 1.02x 9.51 GiB/s 9.07 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.54 KiB 18.54 KiB -0.00 KiB 1.00x
installed native package binary 4061.50 KiB 4064.56 KiB +3.06 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

ccusage performance comparison

PR SHA: b65043a02979
Base SHA: 0c1c65814a08

This compares the Rust PR release binary against the configured base package on the same CI runner.

Package runtime diagnostics

Compares the PR package wrapper, the installed native optional dependency binary, and the workspace release binary on the same large fixture. This identifies whether slow package results come from JavaScript wrapper overhead, the published native binary build, or the Rust core itself.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
All rows run --offline --json, measured by hyperfine with 0 warmups and 1 runs. This isolates wrapper overhead from the installed native optional dependency and the workspace release binary built on the runner.

Command Runtime Input Median Throughput Samples
claude --offline --json Package wrapper 1.01 GiB 312.6ms 3.22 GiB/s 1
claude --offline --json Installed native binary 1.01 GiB 282.5ms 3.56 GiB/s 1
codex --offline --json Package wrapper 1.01 GiB 118.8ms 8.47 GiB/s 1
codex --offline --json Installed native binary 1.01 GiB 89.0ms 11.31 GiB/s 1

Committed fixture performance

Committed small fixtures for stable PR-to-PR feedback and explicit Claude/Codex command coverage.

Fixtures: Claude apps/ccusage/test/fixtures/claude (0.00 MiB, 2 files), Codex apps/ccusage/test/fixtures/codex (0.00 MiB, 1 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 2 warmups and 7 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude daily --offline --json 0.00 MiB 29.4ms 4.7ms 6.27x 54.25 MiB 10.20 MiB 0.19x 0.05 MiB/s 0.33 MiB/s
claude session --offline --json 0.00 MiB 26.8ms 4.2ms 6.45x 53.75 MiB 10.20 MiB 0.19x 0.06 MiB/s 0.37 MiB/s
codex daily --offline --json 0.00 MiB 26.0ms 2.8ms 9.26x 53.50 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.31 MiB/s
codex session --offline --json 0.00 MiB 26.3ms 2.5ms 10.52x 53.75 MiB 8.18 MiB 0.15x 0.03 MiB/s 0.34 MiB/s

Large real-world-shaped fixture performance

Generated fixtures shaped from aggregate local log statistics: thousands of JSONL files, many small sessions, and a long tail of larger sessions. No real prompts, paths, or outputs are stored in the fixtures.

Fixtures: Claude /home/runner/_work/_temp/ccusage-large-fixture (1.01 GiB, 2597 files), Codex /home/runner/_work/_temp/ccusage-large-codex-fixture (1.01 GiB, 2597 files)
Base runs the published ccusage package from pkg.pr.new, installed before measurement; PR runs the published native ccusage binary from pkg.pr.new, installed before measurement. Both run --offline --json, measured by hyperfine with 0 warmups and 1 runs.
Peak RSS is measured separately with /usr/bin/time using 1 runs. Lower RSS ratios are better.

Command Input Base median PR median PR vs base Base peak RSS PR peak RSS PR/base RSS Base throughput PR throughput
claude --offline --json 1.01 GiB 334.3ms 297.9ms 1.12x 944.33 MiB 946.32 MiB 1.00x 3.01 GiB/s 3.38 GiB/s
codex --offline --json 1.01 GiB 116.5ms 85.7ms 1.36x 433.04 MiB 425.04 MiB 0.98x 8.64 GiB/s 11.75 GiB/s

Artifact size

Artifact Base PR Delta Ratio
packed ccusage-*.tgz 18.54 KiB 18.54 KiB -0.00 KiB 1.00x
installed native package binary 4061.50 KiB 4064.56 KiB +3.06 KiB 1.00x

Lower medians and smaller artifacts are better. CI runner noise still applies; use same-run ratios as directional PR feedback, not release guarantees.

axisrow added a commit to axisrow/ccusage that referenced this pull request Aug 12, 2026
* chore(ci): remove pullfrog because they dont serve free tokens anymore

* Restore `pullfrog.yml` workflow

* feat(statusline): show reasoning effort level next to model name (ccusage#1405)

* feat(statusline): show reasoning effort level next to model name

Claude Code 2.1.119+ includes an optional top-level effort.level field
(low, medium, high, xhigh, or max) in the statusline hook JSON,
reflecting the live /effort setting. Parse it from the hook input and
append it to the model segment, e.g. '🤖 Fable 5 (high)'.

The field is absent for models without the effort parameter and on
older Claude Code versions, in which case the statusline keeps showing
just the model label as before.

* test(statusline): add Fable 5 fixture with effort level

Adds a manual statusline fixture for the latest model shape, including
the effort.level field, plus a test-statusline-fable5 recipe wired into
test-statusline-all so the effort display can be smoke-tested from the
CLI.

* docs(statusline): document effort level next to the model name

Updates the statusline guide examples to the current model display
('Fable 5 (high)') and explains that the reasoning effort level comes
from Claude Code 2.1.119+, with a fallback example for models or
versions that do not report it.

* Revert "Restore `pullfrog.yml` workflow"

This reverts commit 04f45b0.

* fix(codex): skip forked session replay history (ccusage#1369)

* feat(json): emit modelBreakdowns in per-agent JSON reports (ccusage#1395)

Per-agent subcommands (ccusage pi|opencode|amp|hermes|... daily/weekly/
monthly/session --json) compute per-model cost breakdowns — the table
view renders them with --breakdown — but the shared per-agent JSON
serializer never emitted them, forcing JSON consumers to re-derive
model costs they cannot actually reconstruct.

Add "modelBreakdowns" to agent_summary_json, mirroring the unified
serializers (summary_json / session_summary_json). Purely additive:
every pre-existing key and value is unchanged; the codex-native
serializer (models object) is deliberately untouched.

Tests: shared-shape insta snapshot now shows populated breakdowns for
all four report kinds; pi daily JSON asserts a full single-element
breakdown array with non-zero cost (the motivating case); fixture-
driven copilot (real pricing via read_otel_file) and qwen (real JSONL
fixture line) assertions pin their entire breakdown arrays.

* fix(pi): align unified session date filtering (ccusage#1394)

* fix(pi): align unified session date filtering

Filter default pi unified session entries by date before summarizing, matching
`ccusage pi session --pi-path` behavior for inclusive `--until` days.

* style: apply treefmt formatting

---------

Co-authored-by: ryoppippi <[email protected]>

* feat(unified): --sections and --by-agent for single-invocation reporting (ccusage#1396)

* feat(unified): --sections and --by-agent for single-invocation reporting

Dashboards polling ccusage today need one unified invocation per section
plus one per-agent invocation per agent — every call re-scanning all
stores. Two additive flags on the unified commands (and the bare root
invocation) collapse that to a single call:

  ccusage daily --json --sections daily,monthly,session --by-agent

--sections <csv> emits each requested grouping in one envelope from at
most TWO store scans (daily/weekly/monthly share one Daily-kind base
load; session adds one Session-kind load), one process, one pricing
load. Every section is produced by exactly the code path its standalone
command uses — load_sections delegates to the same load_rows machinery,
so section output is identical to a standalone invocation by
construction (covered by fixture equivalence tests including claude
agent-progress usage lines and codex cross-session/model-alias dedupe
cases). Envelope order is deterministic via a local ordered serializer:
invoked section first, remaining sections in canonical order, totals
last; single-section envelopes use the unchanged existing path.

--by-agent adds an "agents" array to daily/weekly/monthly rows (the
internal per-agent breakdowns, now serialized: tokens, cost, and
modelBreakdowns per agent). Session rows are already per-agent, so the
flag is a no-op there. Per-agent costs sum exactly to the combined row.

Backward compatibility: without the new flags, JSON and table output
are byte-identical to before (verified against a 20-invocation golden
matrix on real stores). Tables render requested sections sequentially;
--by-agent is JSON-only.

* refactor(unified): address review feedback on duplication and detected agents

- row_json now composes agent_json and layers on the row-level fields
  (period, metadata, agents), so the shared row shape has a single
  serialization path; output unchanged.
- Extract parse_unified_report_arg so the root, unified-command, and
  top-level-session parse sites share one --all/--sections/--by-agent
  block; the root site keeps its mark_used bookkeeping.
- Carry daily-load and session-load detected agents separately so each
  --sections table header shows the same detected list as the
  equivalent standalone invocation.

* style: apply treefmt formatting

---------

Co-authored-by: ryoppippi <[email protected]>

* feat(pi): named pi-format stores as config-declared agents (ccusage#1397)

* feat(pi): named pi-format stores as config-declared agents

Tools built on pi (oh-my-pi and other forks) keep pi-format session
stores at their own paths. ccusage could only read one pi path universe
and labeled everything it found as agent "pi". Declare named extra
stores in the config file:

  { "pi": { "stores": [ { "name": "omp", "path": "~/.omp/agent/sessions" } ] } }

Each named store loads through the existing pi parser and surfaces as
its OWN agent in the unified reports: rows tagged in metadata.agents,
sessions with projectPath/lastActivity like pi, model labels prefixed
"[<name>] ". Named stores are additive to the default pi store and use
the same path-list semantics (comma-separated, ~-expansion, dedupe) and
the same date-window filtering as `ccusage pi ... --pi-path`.

Costs are computed from the unprefixed model name — the configurable
store name never participates in pricing lookup (a store named "o3"
cannot fabricate o3 pricing; regression-tested), while prefixed
pricingOverrides keys are consulted first and keep working.

Config validation: names match ^[a-z][a-z0-9_-]{0,31}$, reject
collisions with built-in agents (single source of truth asserted
against the unified loader's registry), duplicates, empty paths, and
stores whose resolved paths overlap the default pi store or another
store (silent double-counting is never possible). Invalid stores error
through the same config-error path as other invalid config content.
Absent store paths yield clean empty results, like default pi.

Backward compatibility: without pi.stores configured, all output is
byte-identical to before (verified against a golden matrix on real
stores, including a known pre-existing until-day session-window quirk
in the default pi unified path, deliberately preserved here and fixed
in a separate patch). Committed config schema regenerated.

No CLI surface changes: per-agent subcommands remain a closed set;
named stores appear in unified reports only.

* fix(pi): reject nested/partial named-store path overlaps, dedupe path parsing

- Session files are collected recursively, so a named store rooted at an
  ancestor or descendant of the default pi store (or another named
  store) would ingest the same files twice under different dedupe
  identities. The resolver now rejects any overlap — equal, ancestor,
  or descendant — and partial collisions error instead of silently
  dropping the colliding path, matching the documented contract.
  Regression tests for a store nested inside the default pi path and a
  partial overlap across two stores.
- Extract a shared existing_paths helper in pi/paths.rs; the default
  and named-store variants now differ only in their path mapper, with
  the deliberate ~-expansion difference documented.
- Update config/pi docs for the stricter overlap wording.

* docs(pi): move trailing space out of code spans (markdownlint MD038)

---------

Co-authored-by: ryoppippi <[email protected]>

* fix(kimi): support Kimi Code new wire format (ccusage#1362)

Kimi Code (`~/.kimi-code`) emits a new `wire.jsonl` schema that the old
adapter could not parse, so its usage was silently dropped (ccusage#1261).

- Detect `~/.kimi-code` and the deeper layout
  `sessions/<ws>/<session>/agents/<agent>/wire.jsonl` (5 path components)
  alongside the legacy 3-component layout.
- Parse top-level `type == "usage.record"` lines: camelCase token fields
  (`inputOther`, `inputCacheRead`, `inputCacheCreation`), `time` in
  milliseconds, and `model` prefixed with `kimi-code/` (stripped for
  pricing lookup). Skip cumulative `usageScope == "session"` records.
- Deserialize `time` leniently so a float- or string-encoded timestamp
  degrades to the file-mtime fallback instead of dropping the whole line.
- Walk the correct number of parents in `kimi_root_from_wire_path` for the
  deeper layout so config resolution looks at the right root.
- Keep full backward compatibility with the old StatusUpdate format.
- Update the Kimi guide and data-source docs for `~/.kimi-code`.

Fixes ccusage#1261

Co-authored-by: Claude Opus 4.8 <[email protected]>

* perf: cache PricingMap::find() results and skip redundant opencode pricing checks (ccusage#1407)

* perf(pricing): cache PricingMap::find() results to avoid repeated fuzzy matching

PricingMap::find() does an exact HashMap lookup followed by expensive
fuzzy matching through all ~2,200 pricing entries when the exact model
name is not in the map. When adapters repeatedly query the same model
names, a large fraction of lookups miss the HashMap and trigger a full
scan of the pricing table for every call.

Add a OnceLock&lt;Mutex&lt;FxHashMap&gt;&gt; cache that memoizes find()
results by model name (including None for models not found in pricing).
Once a model name has been resolved, future lookups complete in O(1)
instead of O(n) over the pricing table.

Also add clear_find_cache() called from load_json_with_overrides(),
load_models_dev_models(), and apply_overrides() so the cache stays
consistent when the pricing table is mutated.

* perf(opencode): skip redundant missing-pricing check when cost is known

calculate_open_code_cost and missing_open_code_pricing independently
iterate through the same model candidates. When the cost calculation
already found a valid positive cost (either from a stored cost_usd
field or from pricing lookup), skip the missing-pricing check entirely
since pricing was already resolved.

---------

Co-authored-by: turtton <[email protected]>

* ci(release): migrate from bumpp to tagpr (ccusage#1406)

* ci(release): add tagpr release PR automation

Introduce Songmu/tagpr to manage releases via an auto-generated
release PR: every push to main creates or updates a PR that bumps all
nine workspace package.json versions (tagpr versionFile) and syncs the
Rust workspace via the new `just sync-rust-version` recipe run as
postVersionCommand. Merging the PR tags the merge commit.

GitHub Release creation and CHANGELOG.md generation are disabled in
.tagpr because changelogithub keeps generating the release notes in
the existing style.

Tags pushed with GITHUB_TOKEN do not trigger `on: push: tags`
workflows, so tagpr.yaml dispatches release.yaml explicitly with
`gh workflow run --ref <tag>`; release.yaml gains a workflow_dispatch
trigger for that purpose.

* chore(release): drop bumpp local release flow

Releases are now driven by tagpr in CI, so the local `just release`
recipe and the bumpp dependency are no longer needed. bump.config.ts
is deleted because its cargo set-version hook moved to the
`just sync-rust-version` recipe that tagpr runs as postVersionCommand.

* docs(skills): document tagpr release flow

Replace the removed `just release` recipe in the development skill
command list with a note on the tagpr release PR flow and the
minor/major bump labels.

* ci(release): gate release jobs to tag refs and isolate actions:write

Gate release.yaml build/publish/release jobs behind startsWith(github.ref, 'refs/tags/') so a workflow_dispatch from a branch cannot bypass the tag-only release flow. tagpr dispatches with --ref <tag>, so the intended path is unaffected.

Move the release dispatch out of the tagpr job into a dependent dispatch-release job that alone holds actions: write, keeping tagpr on its documented least-privilege scopes (contents/pull-requests/issues).

Co-authored-by: Codesmith <[email protected]>

* ci(release): consolidate release pipeline into tagpr workflow

Move the build/publish/release jobs from release.yaml into tagpr.yaml,
gated on the tagpr job's tag output, and delete release.yaml. Running
everything in one workflow removes the workflow_dispatch chaining that
worked around GITHUB_TOKEN-pushed tags not triggering `push: tags`
workflows, along with the dispatch-release job and its `actions: write`
grant.

The release jobs check out the freshly created tag explicitly.
changelogithub resolves the release tag with `git tag --points-at
HEAD`, not GITHUB_REF, so it picks the right release even though the
run's ref is refs/heads/main.

A failed release is retried with "Re-run failed jobs"; a full re-run
finds no new tag and skips the release jobs.

---------

Co-authored-by: Codesmith <[email protected]>

* ci(release): use Conventional Commits title for tagpr release PRs (ccusage#1409)

tagpr titles its release PRs "Release for vX.Y.Z", which fails the
check-pr-title workflow because it has no Conventional Commits type
prefix (seen on PR ccusage#1408). tagpr takes the first line of the rendered
pull request template as the PR title, so point .tagpr at a custom
template whose first line is "chore: release {{.NextVersion}}". The
rest of the template mirrors tagpr's default body, minus the unused
tag-prefix placeholder.

The template is a Go text/template, and oxfmt's markdown rewrites
break its <details> block and nested list structure, so exclude it
from treefmt.

This also restores the title style used by the previous bumpp-based
release flow ("chore: release v20.0.14").

* feat(pricing): support OpenAI two-stage pricing and add the gpt-5.6 family (ccusage#1414)

* feat(pricing): add gpt-5.6 family and OpenAI long-context tier rates

OpenAI introduced two-stage (short/long context) pricing with gpt-5.6:
requests with more than 272K input tokens are billed at higher
long-context rates. The same tier also applies to gpt-5.5, gpt-5.5-pro,
gpt-5.4, and gpt-5.4-pro on the current pricing page.

The existing tier support hardcoded the LiteLLM 200K boundary, so
Pricing gains a per-model long_context_threshold (defaulting to 200K
for LiteLLM *_above_200k_tokens data) and tiered_cost takes the
threshold as a parameter.

New built-in entries cover gpt-5.6-sol, gpt-5.6-terra, and
gpt-5.6-luna, including their cache-write rates. Long-context tier
rates live in a builtin_long_context_rates overlay that is re-applied
after every pricing load: a live LiteLLM refresh replaces whole
entries, and LiteLLM currently publishes these models with flat rates
only, so tier rates set directly on built-in entries would be silently
dropped whenever a refresh succeeds. Entries that already carry tier
rates are left untouched so upstream data wins once it exists.
Date-pinned keys such as gpt-5.5-2026-04-23 share their base model's
overlay rates.

The gpt-5.6 context limits mirror the 1,050,000-token window of the
other long-context GPT-5 flagship models until upstream data lands.

* feat(codex): bill long-context requests at OpenAI two-stage rates

OpenAI decides the pricing tier per request: once a request's input
exceeds 272K tokens, every token of that request (input, cached input,
and output) is billed at the long-context rates. Codex cost calculation
runs on per-model sums aggregated across many requests, so the tier
cannot be recovered from the totals afterwards.

CodexModelUsage now tracks the portion of tokens that came from
long-context requests. The split is recorded while token_count events
are aggregated, where each event still represents a single request,
and merged across parallel shards like the other counters.

calculate_codex_model_cost prices the aggregated usage as two
independent buckets: the short bucket at the flat rates and the long
bucket at the *_above_200k rates, falling back to the flat rates for
models without a long-context tier so their costs are unchanged. The
existing fast-speed multiplier applies to both buckets.

Report JSON and table output are unchanged; only costUSD values for
long-context requests differ.

* docs(pricing): explain all-or-nothing long-context overlay check

Codex review suggested filling missing tier fields independently when a
refreshed LiteLLM entry carries partial *_above_200k_tokens data. That
would mix rates that assume the 200K LiteLLM boundary with built-in
rates that assume the OpenAI 272K boundary under a single per-model
threshold, mispricing both tiers, so the overlay defers to upstream
entirely once any tier rate exists. Record that rationale next to the
check.

* fix(pricing): apply two-stage rates to whole request and per-model split

Co-authored-by: Codesmith <[email protected]>

---------

Co-authored-by: Codesmith <[email protected]>

* chore: use black smith more

* chore: release v20.0.15 (ccusage#1408)

[tagpr] prepare for the next release

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* chore(ci): rename it back to releaese.yaml

* chore: release v20.0.16 (ccusage#1416)

[tagpr] prepare for the next release

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* docs: update Lineman affiliate links to CCUsage landing page (ccusage#1417)

Point GitHub README and docs site sponsor links at the dedicated
LinkJolt redirect for CCUsage traffic (free tier + voucher funnel).

Co-authored-by: Cursor Agent <[email protected]>
Co-authored-by: ryoppippi <[email protected]>

* docs: update Star History chart (ccusage#1419)

* docs: update Star History chart

Switch the README and sponsorship guide to the current Star History chart endpoint, including light and dark variants. Allowlist the public read-only sealed chart token so secret scanning does not report a false positive.

Co-authored-by: ryoppippi <[email protected]>

* chore: exclude sealed token from spellcheck

Mark the exact public Star History token allowlist line as a spellcheck exclusion so its random character sequence does not fail the documentation preflight.

Co-authored-by: ryoppippi <[email protected]>

* chore: format sealed token allowlist

Use the repository's TOML formatting and bracket the random token with the supported spellchecker block directives.

Co-authored-by: ryoppippi <[email protected]>

* style: align Gitleaks TOML indentation

Match the repository formatter's tab indentation for the multiline allowlist entry.

Co-authored-by: ryoppippi <[email protected]>

* fix: match full Star History token URL

Configure the global Gitleaks allowlist to evaluate the full finding match so the narrowly scoped sealed_token pattern suppresses the six intentional chart URLs.

Co-authored-by: ryoppippi <[email protected]>

---------

Co-authored-by: Cursor Agent <[email protected]>
Co-authored-by: ryoppippi <[email protected]>

* fix(claude): count advisor model usage (ccusage#1423)

* fix(claude): count advisor model usage

Expand advisor_message iterations into distinct usage entries so their tokens and model-specific costs are included in every report path. Keep main-model iteration totals unchanged and cover both standard and daily loaders.

Co-authored-by: ryoppippi <[email protected]>

* docs(claude): clarify advisor cost modes

Co-authored-by: ryoppippi <[email protected]>

---------

Co-authored-by: Cursor Agent <[email protected]>
Co-authored-by: ryoppippi <[email protected]>

* chore: release v20.0.17 (ccusage#1418)

[tagpr] prepare for the next release

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* perf(nix): keep dependency cache across releases (ccusage#1424)

* build(perf): migrate benchmark harness to Babashka (ccusage#1432)

* build(perf): migrate benchmark harness to Babashka

Replace the large Nushell PR benchmark script with a Babashka implementation split by data, system, benchmark, report, and orchestration responsibilities. The new process boundary keeps argv, environment, and working-directory data explicit while preserving hyperfine, package installation, memory, size, and Markdown behavior.

Move the CI caller and profiling guidance to the executable Babashka entry point. Add focused tests behind their own Nix shebang so contributors can run the harness suite without adding Babashka to the full development shell.

* docs(agents): document implementation language choices

Route small command-oriented automation to Nushell and data-heavy, testable automation to Babashka. Keep production binaries in Rust and npm-integrated APIs in TypeScript so future tooling changes follow the same criteria used by the benchmark migration.

* test(ci): run Babashka harness tests

Execute the self-contained benchmark harness test entry point in the CI test job so changes to CLI parsing, normalization, fallback decisions, and report rendering cannot bypass pull request validation.

* fix(perf): harden platform and tarball paths

Normalize version-qualified Windows os.name values to win32 so native executable and package paths use the expected suffixes. Resolve relative pnpm pack filenames against the temporary destination while preserving the absolute paths emitted by current pnpm versions.

Add regression coverage for both platform normalization and relative or absolute tarball filenames.

* fix(perf): size local package fallbacks

Use remote tarball sizing only after the corresponding preview package was installed successfully. When either package URL times out, benchmark and size the available local checkout so fallback runs can still produce a complete report.

Cover base and head source selection and verify both unavailable URLs through a committed-fixture smoke run.

* fix(perf): bound harness child processes and skip RSS on unsupported platforms

Add a cancellable timeout to run-process and thread --package-runner-timeout-ms through the package URL probe, install, pnpm pack, and git rev-parse flows so a stalled child cannot outlive the deadline; give the curl probe and download explicit connect and read limits.

measure-memory now warns once and skips gracefully when /usr/bin/time is unavailable (unsupported platforms) instead of throwing and aborting the entire benchmark run.

Co-authored-by: Codesmith <[email protected]>

* Revert "fix(perf): bound harness child processes and skip RSS on unsupported platforms"

This reverts commit 9140a99.

---------

Co-authored-by: Codesmith <[email protected]>

* build(perf): migrate fixture generator to Bun (ccusage#1433)

* build(perf): migrate fixture generator to Bun

Replace the Nushell fixture generator with a dependency-free Bun script.\n\nKeep the generated Claude and Codex fixture layouts and command-line\ninterface while using Bun file writers and Bun Shell for file operations.

* build(perf): type Bun fixture script

Add Bun development types so the fixture generator is checked alongside the package tooling.\n\nAwait file writer operations to preserve ordered writes and satisfy the\nrepository promise lint rule.

* build(perf): avoid Bun type dependency

Keep the fixture generator dependency-free by declaring its small Bun API surface locally.\n\nRemove the Bun type package and restore the package TypeScript configuration so\npublishing the fixture generator does not expand package dependencies.

* chroe(ci): fix nix cache

* ci: add GitHub-hosted runner fallback

Keep Blacksmith runners for the upstream repository while allowing forks\nto use hosted runners by default. Forks with a Blacksmith subscription can\nopt in through HAS_BLACKSMITH=true.

* ci: skip pkg-pr previews without the GitHub App

Forks do not inherit the pkg-pr-new GitHub App installation. Skip\npreview publishing and its dependent E2E and performance jobs unless a fork\nexplicitly opts in with HAS_PKG_PR_NEW=true.

* ci: keep Windows arm release runner defined

Supply both matrix runner fields so the release workflow resolves its\nWindows ARM runner consistently with the other native package targets.

---------

Co-authored-by: ryoppippi <[email protected]>
Co-authored-by: pullfrog[bot] <226033991+pullfrog[bot]@users.noreply.github.com>
Co-authored-by: sijie-ni-0214 <[email protected]>
Co-authored-by: Ben Vargas <[email protected]>
Co-authored-by: Mint Choco <[email protected]>
Co-authored-by: Claude Opus 4.8 <[email protected]>
Co-authored-by: turtton <[email protected]>
Co-authored-by: Codesmith <[email protected]>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cursor Agent <[email protected]>
Co-authored-by: ryoppippi <[email protected]>
Co-authored-by: axisrow <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants