Submission checklist
Area (Required)
Related Issues / PRs
Reproduction Steps / Example Code (Python)
"""Minimal repro: the `task` tool discards the subagent's token usage.
Run with: python repro_subagent_usage.py
"""
from typing import Any
from deepagents import create_deep_agent
from langchain_core.callbacks import CallbackManagerForLLMRun
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import AIMessage, BaseMessage
from langchain_core.outputs import ChatGeneration, ChatResult
PARENT_TOKENS = 100
SUBAGENT_TOKENS = 5000
class ScriptedModel(BaseChatModel):
"""Replies from a script, reporting usage on every call like a real model."""
replies: list[AIMessage]
calls: list[int] = []
def _generate(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: CallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> ChatResult:
index = len(self.calls)
self.calls.append(index)
reply = self.replies[min(index, len(self.replies) - 1)]
return ChatResult(generations=[ChatGeneration(message=reply)])
def bind_tools(self, tools: Any, **kwargs: Any) -> "ScriptedModel":
return self # the script decides what gets called
@property
def _llm_type(self) -> str:
return "scripted"
def _usage(total: int) -> dict[str, int]:
return {"input_tokens": total, "output_tokens": 0, "total_tokens": total}
def main() -> None:
parent = ScriptedModel(
replies=[
# 1. the parent delegates to the subagent…
AIMessage(
content="",
tool_calls=[
{
"name": "task",
"args": {
"description": "do the work",
"subagent_type": "worker",
},
"id": "call-1",
"type": "tool_call",
}
],
usage_metadata=_usage(PARENT_TOKENS),
),
# 2. …then wraps up.
AIMessage(content="all done", usage_metadata=_usage(PARENT_TOKENS)),
]
)
worker = ScriptedModel(
replies=[
AIMessage(content="worked hard", usage_metadata=_usage(SUBAGENT_TOKENS))
]
)
agent = create_deep_agent(
model=parent,
tools=[],
subagents=[
{
"name": "worker",
"description": "does the work",
"prompt": "You do the work.",
"model": worker,
}
],
)
result = agent.invoke({"messages": [("user", "delegate this")]})
reported = sum(
(m.usage_metadata or {}).get("total_tokens", 0)
for m in result["messages"]
if isinstance(m, AIMessage)
)
actually_spent = PARENT_TOKENS * len(parent.calls) + SUBAGENT_TOKENS * len(
worker.calls
)
tool_messages = [m for m in result["messages"] if m.type == "tool"]
print(f"parent model calls : {len(parent.calls)}")
print(f"subagent model calls : {len(worker.calls)}")
print(f"tokens actually spent : {actually_spent}")
print(f"tokens visible in state: {reported}")
print(f"task ToolMessage usage : {[m.response_metadata for m in tool_messages]}")
print()
print(f"MISSING: {actually_spent - reported} tokens — the entire subagent spend.")
assert worker.calls, "subagent never ran"
assert reported < actually_spent, "usage was propagated after all"
if __name__ == "__main__":
main()
Error Message and Stack Trace (if applicable)
No error is raised — that is the problem. Output on 0.7.19:
parent model calls : 2
subagent model calls : 1
tokens actually spent : 5200
tokens visible in state: 200
task ToolMessage usage : [{}]
MISSING: 5000 tokens — the entire subagent spend.
96% of the run's tokens are invisible to anything reading agent state, and the `task` `ToolMessage` carries empty `response_metadata`.
Description
When the task tool finishes, the subagent's entire message history is reduced to a single ToolMessage carrying only text. Every AIMessage.usage_metadata produced inside the subagent is dropped at that boundary and is not recorded anywhere else.
deepagents/middleware/subagents.py:742-755 (0.7.19):
content = ""
for msg in reversed(result["messages"]):
if isinstance(msg, AIMessage):
text = msg.text.rstrip() if msg.text else ""
if text:
content = text
break
return Command(
update={
**state_update,
"messages": [ToolMessage(content, tool_call_id=tool_call_id)],
}
)
result["messages"] holds the full subagent history, including its token counts. It is walked purely to scrape the last non-empty text, and the returned ToolMessage gets no response_metadata.
The consequence is that any usage accounting done by summing usage_metadata over agent state is structurally incapable of seeing subagent spend — and for delegation-heavy agents the subagent is where most of the tokens go, so the number isn't slightly low, it's mostly wrong. There's no error and no warning; the total just looks plausible.
To be clear about what I am not asking for: collapsing the subagent transcript is the entire point of the task tool, and forwarding those messages would defeat context isolation. Usage metadata is orthogonal to that — it's a small fixed-size dict, not context.
Expected: the aggregated subagent usage survives the boundary, e.g. attached to the returned ToolMessage's response_metadata, so downstream consumers can account for it without each reinventing out-of-band tracking (as #5766 shows dcode having to do).
Observed in production: we run a delegation-heavy agent on deepagents and our per-run token reporting was understating spend by roughly half (~$3 reported against ~$6 actual on one run of 152 requests) until we worked around it with a run-scoped BaseCallbackHandler that catches on_llm_end across the whole callback tree. That workaround only works because it bypasses agent state entirely — it shouldn't be necessary.
Environment / System Info
> OS: Darwin
> Python Version: 3.13.12
> deepagents: 0.7.19
> langchain: 1.4.2
> langchain_core: 1.6.5
> langgraph: 1.2.12
Submission checklist
Area (Required)
Related Issues / PRs
dcodealready tracks_session_cost_usd"including nested model calls", i.e. it reconstructs this itself inlibs/codebecause the SDK does not surface it.--max-costspend cap that halts the agent #6209 — a hard--max-costspend cap; a cap can only be as accurate as the usage it can see, and today subagent spend is invisible to anything reading agent state.Reproduction Steps / Example Code (Python)
Error Message and Stack Trace (if applicable)
Description
When the
tasktool finishes, the subagent's entire message history is reduced to a singleToolMessagecarrying only text. EveryAIMessage.usage_metadataproduced inside the subagent is dropped at that boundary and is not recorded anywhere else.deepagents/middleware/subagents.py:742-755(0.7.19):result["messages"]holds the full subagent history, including its token counts. It is walked purely to scrape the last non-empty text, and the returnedToolMessagegets noresponse_metadata.The consequence is that any usage accounting done by summing
usage_metadataover agent state is structurally incapable of seeing subagent spend — and for delegation-heavy agents the subagent is where most of the tokens go, so the number isn't slightly low, it's mostly wrong. There's no error and no warning; the total just looks plausible.To be clear about what I am not asking for: collapsing the subagent transcript is the entire point of the
tasktool, and forwarding those messages would defeat context isolation. Usage metadata is orthogonal to that — it's a small fixed-size dict, not context.Expected: the aggregated subagent usage survives the boundary, e.g. attached to the returned
ToolMessage'sresponse_metadata, so downstream consumers can account for it without each reinventing out-of-band tracking (as #5766 showsdcodehaving to do).Observed in production: we run a delegation-heavy agent on deepagents and our per-run token reporting was understating spend by roughly half (~$3 reported against ~$6 actual on one run of 152 requests) until we worked around it with a run-scoped
BaseCallbackHandlerthat catcheson_llm_endacross the whole callback tree. That workaround only works because it bypasses agent state entirely — it shouldn't be necessary.Environment / System Info