Skip to content

Emit per-call usage events with list-price estimates - #46

Open
abhinavgautam01 wants to merge 2 commits into
alpha-omega-security:mainfrom
abhinavgautam01:feature/per-turn-usage-events
Open

abhinavgautam01 wants to merge 2 commits into
alpha-omega-security:mainfrom
abhinavgautam01:feature/per-turn-usage-events

Conversation

@abhinavgautam01

Copy link
Copy Markdown
Contributor

Part of alpha-omega-security/scrutineer#774

Problem

Event.Usage and Event.CostUSD are only filled in on the final result event, so a caller learns what a run cost only after it ends. scrutineer#774 needs a running estimate while the run is in progress, so the worker can stop a scan once it crosses a cost cap. That check belongs in the worker rather than in each backend, so the harness has to expose usage as it happens.

Change

A new usage event kind reports the tokens one model call added, with its Model and a list-price CostUSD estimate (zero for an unknown model). Summing usage events gives the running estimate. The result event is unchanged and stays the authoritative total, so callers must not add usage events to it.

Claude

I checked the Claude CLI's stream-json against real runs before writing the parser:

  • One API message is split into several assistant lines that all repeat the same usage and that output_tokens is only the snapshot from the start of the message (3 against 273 actually used).
  • With --include-partial-messages, each main-thread call also emits a message_delta stream event carrying the message's final cumulative usage. Those deltas summed exactly to the result's output_tokens.
  • Subagent calls emit no stream events. They only appear as assistant lines with parent_tool_use_id set.

So Args now adds --include-partial-messages and the parser keeps per-stream state:

  • Usage is tracked per message id. Only the increase since the last report is emitted, so repeated lines and zero-usage <synthetic> messages emit nothing.
  • message_start, message_delta and assistant lines all feed that tracker. Subagent output tokens are therefore a lower bound until the result arrives.
  • Every other stream_event line is dropped silently, so text deltas never reach the log.

Two pricing fixes were needed for the estimate to be right:

  • Dated model ids. message.model is dated (claude-haiku-4-5-20251001), which priced at $0. normalizeModelID now strips a trailing -YYYYMMDD. This also helps any other caller of CostFromUsage that passes a dated id.
  • One-hour cache writes. Claude Code writes one-hour cache entries (cache_creation.ephemeral_1h_input_tokens), which bill at 2x the input rate, while the table's CacheWrite is the five-minute rate. Without this fix the estimate came out about 25% low on a real run. The surcharge is computed from the one-hour split without changing the public Usage struct, since callers convert into it.

Other backends

  • Copilot emits a usage event for each assistant.usage record, sub-agent calls included. The result event is unchanged and still prefers the billing checkpoint.
  • OpenCode emits a usage event before each step_finish result. It is skipped when the step has no tokens and no cost.
  • Codex reports usage only on turn.completed, which is already the result event, so it emits no usage events.

FormatEvent renders [usage] <model> in=… out=… cache_read=… cache_write=… cost=$…. The README documents the new kind and the Model field.

Verification

  • Replaying real Claude CLI output (v2.1.28x, Haiku 4.5) through the new parser:

    Stream Usage events Result
    Main-thread run with partial messages tokens 18 in / 280 out / 37822 cache read / 5884 cache write, cost $0.0169682 identical tokens, total_cost_usd $0.0169682
    Run using a subagent $0.04224 $0.04407 (subagent output is a lower bound, as documented)
  • New tests:

    • Claude: dedupe across repeated lines, the message_start / message_delta sum, subagent, synthetic and unknown-model cases, one-hour cache pricing charged once and the new arg.
    • Copilot and OpenCode usage events.
    • FormatEvent, dated id normalization and the one-hour surcharge.
  • Mutation checks: removing the per-message dedupe, the one-hour dedupe, the date stripping or the new arg each makes a test fail.

  • gofmt, go vet (also -tags integration), go test -race ./..., golangci-lint v2.14.0 (0 issues) and go mod tidy -diff all pass.

Note for callers

Callers that log every event through FormatEvent will now see one [usage] line per model call. Callers that sum CostUSD over result events are unaffected.

@abhinavgautam01

Copy link
Copy Markdown
Contributor Author

hey @andrew,

The vuln check failed on ten Go standard library vulnerabilities in go1.27.1 (net/http, net/http/internal/http2, crypto/tls, net/textproto, os), all fixed in go1.27.2. It isn't caused by this change and main fails the same check. I bumped toolchain to go1.27.2 in a separate commit and govulncheck v1.8.0 now reports no vulnerabilities.

Happy to move that commit into its own PR if you'd rather land it on main first.

@andrew

andrew commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

patrick already bumped it on main: aef582a

@abhinavgautam01
abhinavgautam01 force-pushed the feature/per-turn-usage-events branch from d486578 to b9d5fae Compare October 10, 2026 09:22
@abhinavgautam01

Copy link
Copy Markdown
Contributor Author

thanks, i dropped my bump commit and rebased onto main, so this PR is back to just the usage events change.

@andrew andrew left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The main-thread output count in handleStreamEvent depends on api_message_id being present on the message_delta line. In the current CLI schema that field is marked @internal, and its description says it is absent from older producers and on events a plugin produced or rewrote. When it is missing, reportUsage("") returns early, the delta is dropped, and output tokens stay at the message_start snapshot (3 against 273 in your own run). The running estimate then comes out far too low with no signal, which is the number the scrutineer cost cap would check, and nothing here pins the Claude CLI version.

Main-thread stream events are sequential and subagents emit none, so recording m.ID as the current message on message_start and using it when api_message_id is empty covers this. Please add a test where the message_delta line has no api_message_id.

@abhinavgautam01

Copy link
Copy Markdown
Contributor Author

Thanks, fixed. The parser now records the message that message_start opened and uses it when message_delta has no api_message_id. TestClaudeUsageEventsWithoutAPIMessageID covers two messages with deltas lacking the field and fails without the fallback.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants