Skip to content

task tool discards the subagent's token usage, making subagent spend invisible to agent-state accounting #6556

Description

@Namang1

Submission checklist

  • This is a bug, not a usage question.
  • I added a clear and descriptive title.
  • I searched existing issues and didn't find this.
  • I can reproduce this with the latest released version.
  • I included a minimal reproducible example and steps to reproduce.

Area (Required)

  • deepagents (SDK)
  • dcode
  • talon
  • acp
  • evals
  • daytona
  • modal
  • quickjs
  • runloop
  • vercel
  • langsmith-sandbox
  • Other / not sure / general

Related Issues / PRs

Reproduction Steps / Example Code (Python)

"""Minimal repro: the `task` tool discards the subagent's token usage.

Run with:  python repro_subagent_usage.py
"""

from typing import Any

from deepagents import create_deep_agent
from langchain_core.callbacks import CallbackManagerForLLMRun
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import AIMessage, BaseMessage
from langchain_core.outputs import ChatGeneration, ChatResult

PARENT_TOKENS = 100
SUBAGENT_TOKENS = 5000


class ScriptedModel(BaseChatModel):
    """Replies from a script, reporting usage on every call like a real model."""

    replies: list[AIMessage]
    calls: list[int] = []

    def _generate(
        self,
        messages: list[BaseMessage],
        stop: list[str] | None = None,
        run_manager: CallbackManagerForLLMRun | None = None,
        **kwargs: Any,
    ) -> ChatResult:
        index = len(self.calls)
        self.calls.append(index)
        reply = self.replies[min(index, len(self.replies) - 1)]
        return ChatResult(generations=[ChatGeneration(message=reply)])

    def bind_tools(self, tools: Any, **kwargs: Any) -> "ScriptedModel":
        return self  # the script decides what gets called

    @property
    def _llm_type(self) -> str:
        return "scripted"


def _usage(total: int) -> dict[str, int]:
    return {"input_tokens": total, "output_tokens": 0, "total_tokens": total}


def main() -> None:
    parent = ScriptedModel(
        replies=[
            # 1. the parent delegates to the subagent…
            AIMessage(
                content="",
                tool_calls=[
                    {
                        "name": "task",
                        "args": {
                            "description": "do the work",
                            "subagent_type": "worker",
                        },
                        "id": "call-1",
                        "type": "tool_call",
                    }
                ],
                usage_metadata=_usage(PARENT_TOKENS),
            ),
            # 2. …then wraps up.
            AIMessage(content="all done", usage_metadata=_usage(PARENT_TOKENS)),
        ]
    )
    worker = ScriptedModel(
        replies=[
            AIMessage(content="worked hard", usage_metadata=_usage(SUBAGENT_TOKENS))
        ]
    )

    agent = create_deep_agent(
        model=parent,
        tools=[],
        subagents=[
            {
                "name": "worker",
                "description": "does the work",
                "prompt": "You do the work.",
                "model": worker,
            }
        ],
    )

    result = agent.invoke({"messages": [("user", "delegate this")]})

    reported = sum(
        (m.usage_metadata or {}).get("total_tokens", 0)
        for m in result["messages"]
        if isinstance(m, AIMessage)
    )
    actually_spent = PARENT_TOKENS * len(parent.calls) + SUBAGENT_TOKENS * len(
        worker.calls
    )

    tool_messages = [m for m in result["messages"] if m.type == "tool"]
    print(f"parent model calls    : {len(parent.calls)}")
    print(f"subagent model calls  : {len(worker.calls)}")
    print(f"tokens actually spent : {actually_spent}")
    print(f"tokens visible in state: {reported}")
    print(f"task ToolMessage usage : {[m.response_metadata for m in tool_messages]}")
    print()
    print(f"MISSING: {actually_spent - reported} tokens — the entire subagent spend.")

    assert worker.calls, "subagent never ran"
    assert reported < actually_spent, "usage was propagated after all"


if __name__ == "__main__":
    main()

Error Message and Stack Trace (if applicable)

No error is raised — that is the problem. Output on 0.7.19:


parent model calls    : 2
subagent model calls  : 1
tokens actually spent : 5200
tokens visible in state: 200
task ToolMessage usage : [{}]

MISSING: 5000 tokens — the entire subagent spend.


96% of the run's tokens are invisible to anything reading agent state, and the `task` `ToolMessage` carries empty `response_metadata`.

Description

When the task tool finishes, the subagent's entire message history is reduced to a single ToolMessage carrying only text. Every AIMessage.usage_metadata produced inside the subagent is dropped at that boundary and is not recorded anywhere else.

deepagents/middleware/subagents.py:742-755 (0.7.19):

            content = ""
            for msg in reversed(result["messages"]):
                if isinstance(msg, AIMessage):
                    text = msg.text.rstrip() if msg.text else ""
                    if text:
                        content = text
                        break

        return Command(
            update={
                **state_update,
                "messages": [ToolMessage(content, tool_call_id=tool_call_id)],
            }
        )

result["messages"] holds the full subagent history, including its token counts. It is walked purely to scrape the last non-empty text, and the returned ToolMessage gets no response_metadata.

The consequence is that any usage accounting done by summing usage_metadata over agent state is structurally incapable of seeing subagent spend — and for delegation-heavy agents the subagent is where most of the tokens go, so the number isn't slightly low, it's mostly wrong. There's no error and no warning; the total just looks plausible.

To be clear about what I am not asking for: collapsing the subagent transcript is the entire point of the task tool, and forwarding those messages would defeat context isolation. Usage metadata is orthogonal to that — it's a small fixed-size dict, not context.

Expected: the aggregated subagent usage survives the boundary, e.g. attached to the returned ToolMessage's response_metadata, so downstream consumers can account for it without each reinventing out-of-band tracking (as #5766 shows dcode having to do).

Observed in production: we run a delegation-heavy agent on deepagents and our per-run token reporting was understating spend by roughly half (~$3 reported against ~$6 actual on one run of 152 requests) until we worked around it with a run-scoped BaseCallbackHandler that catches on_llm_end across the whole callback tree. That workaround only works because it bypasses agent state entirely — it shouldn't be necessary.

Environment / System Info

> OS:  Darwin
> Python Version:  3.13.12
> deepagents: 0.7.19
> langchain: 1.4.2
> langchain_core: 1.6.5
> langgraph: 1.2.12

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    org:externalIssue or pull request created by someone outside the `langchain-ai` GitHub organization.package:deepagentsChanges related to the `deepagents` SDK and agent harness.priority:backlogNot currently planned/prioritized work, often affecting a limited feature or set of users.topic:async-subagentsAsync subagent execution and orchestration.topic:performancePerformance, resource usage, and cost optimization.topic:subagentsSubagent creation, routing, and orchestration.type:featureA request, idea, or new user-facing functionality or behavior.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions