Repository navigation
Emit per-call usage events with list-price estimates - #46
abhinavgautam01 wants to merge 2 commits into
Conversation
|
hey @andrew, The Happy to move that commit into its own PR if you'd rather land it on |
|
patrick already bumped it on main: aef582a |
d486578 to
b9d5fae
Compare
|
thanks, i dropped my bump commit and rebased onto main, so this PR is back to just the usage events change. |
andrew
left a comment
There was a problem hiding this comment.
The main-thread output count in handleStreamEvent depends on api_message_id being present on the message_delta line. In the current CLI schema that field is marked @internal, and its description says it is absent from older producers and on events a plugin produced or rewrote. When it is missing, reportUsage("") returns early, the delta is dropped, and output tokens stay at the message_start snapshot (3 against 273 in your own run). The running estimate then comes out far too low with no signal, which is the number the scrutineer cost cap would check, and nothing here pins the Claude CLI version.
Main-thread stream events are sequential and subagents emit none, so recording m.ID as the current message on message_start and using it when api_message_id is empty covers this. Please add a test where the message_delta line has no api_message_id.
|
Thanks, fixed. The parser now records the message that |
Part of alpha-omega-security/scrutineer#774
Problem
Event.UsageandEvent.CostUSDare only filled in on the finalresultevent, so a caller learns what a run cost only after it ends. scrutineer#774 needs a running estimate while the run is in progress, so the worker can stop a scan once it crosses a cost cap. That check belongs in the worker rather than in each backend, so the harness has to expose usage as it happens.Change
A new
usageevent kind reports the tokens one model call added, with itsModeland a list-priceCostUSDestimate (zero for an unknown model). Summing usage events gives the running estimate. Theresultevent is unchanged and stays the authoritative total, so callers must not add usage events to it.Claude
I checked the Claude CLI's stream-json against real runs before writing the parser:
assistantlines that all repeat the sameusageand thatoutput_tokensis only the snapshot from the start of the message (3 against 273 actually used).--include-partial-messages, each main-thread call also emits amessage_deltastream event carrying the message's final cumulative usage. Those deltas summed exactly to the result'soutput_tokens.assistantlines withparent_tool_use_idset.So
Argsnow adds--include-partial-messagesand the parser keeps per-stream state:<synthetic>messages emit nothing.message_start,message_deltaandassistantlines all feed that tracker. Subagent output tokens are therefore a lower bound until the result arrives.stream_eventline is dropped silently, so text deltas never reach the log.Two pricing fixes were needed for the estimate to be right:
message.modelis dated (claude-haiku-4-5-20251001), which priced at $0.normalizeModelIDnow strips a trailing-YYYYMMDD. This also helps any other caller ofCostFromUsagethat passes a dated id.cache_creation.ephemeral_1h_input_tokens), which bill at 2x the input rate, while the table'sCacheWriteis the five-minute rate. Without this fix the estimate came out about 25% low on a real run. The surcharge is computed from the one-hour split without changing the publicUsagestruct, since callers convert into it.Other backends
assistant.usagerecord, sub-agent calls included. The result event is unchanged and still prefers the billing checkpoint.step_finishresult. It is skipped when the step has no tokens and no cost.turn.completed, which is already the result event, so it emits no usage events.FormatEventrenders[usage] <model> in=… out=… cache_read=… cache_write=… cost=$…. The README documents the new kind and theModelfield.Verification
Replaying real Claude CLI output (v2.1.28x, Haiku 4.5) through the new parser:
total_cost_usd$0.0169682New tests:
message_start/message_deltasum, subagent, synthetic and unknown-model cases, one-hour cache pricing charged once and the new arg.FormatEvent, dated id normalization and the one-hour surcharge.Mutation checks: removing the per-message dedupe, the one-hour dedupe, the date stripping or the new arg each makes a test fail.
gofmt,go vet(also-tags integration),go test -race ./...,golangci-lintv2.14.0 (0 issues) andgo mod tidy -diffall pass.Note for callers
Callers that log every event through
FormatEventwill now see one[usage]line per model call. Callers that sumCostUSDoverresultevents are unaffected.