Skip to content

Summary input limit can produce a content-free compaction #6470

Description

@open-swe

Submission checklist

  • This is a bug, not a usage question.
  • I added a clear and descriptive title.
  • I searched existing issues and didn't find this.
  • I can reproduce this with the latest released version.
  • I included a minimal reproducible example and steps to reproduce.

Area (Required)

  • deepagents (SDK)
  • dcode
  • talon
  • acp
  • evals
  • daytona
  • modal
  • quickjs
  • runloop
  • vercel
  • langsmith-sandbox
  • Other / not sure / general

Related Issues / PRs

Reproduction Steps / Example Code (Python)

trim_tokens_to_summarize is a safety limit for the internal summarization request. It does not control when conversation compaction starts or how long the final summary may be. After Deep Agents chooses old messages to compact, this option limits how many tokens from those messages are sent to the summarization model, preventing that internal request from exceeding the model's context window.

The bug occurs when the limit is smaller than an individual message. LangChain may be unable to retain even a partial valid message, so non-empty history becomes an empty summarization input:

from deepagents.backends import StateBackend
from deepagents.middleware.summarization import create_summarization_middleware
from langchain.chat_models import init_chat_model
from langchain_core.messages import HumanMessage

model = init_chat_model("openai:gpt-4.1", api_key="not-used")
middleware = create_summarization_middleware(
    model,
    StateBackend(),
    trim_tokens_to_summarize=1,
)

messages = [HumanMessage(content="important production context")]

assert middleware._lc_helper._trim_messages_for_summary(messages) == []
assert middleware._create_summary(messages) == (
    "Previous conversation was too long to summarize."
)

The model API is not called. The same branch is used by _acreate_summary during normal async agent execution.

This also reproduces with trim_tokens_to_summarize=4000 and messages = [HumanMessage(content="important context " * 2000)]. The default partial-message splitter splits on newlines, so an oversized single-line message cannot be shortened to fit even with allow_partial=True. A large single-line tool result or application payload can therefore leave no valid messages for the summarizer. This is pre-summary trimming exhaustion, not a context-window error returned by the summarization model.

Error Message and Stack Trace (if applicable)

No exception is raised. Compaction is recorded as successful with the fallback "Previous conversation was too long to summarize."

Description

During automatic compaction, Deep Agents:

  1. chooses old messages to remove from the active context;
  2. saves those messages to /conversation_history/<session>.md;
  3. summarizes them for the replacement context;
  4. persists the replacement as a successful _summarization_event.

When trim_tokens_to_summarize reduces non-empty history to zero messages, step 3 returns a generic fallback instead of a summary. Deep Agents still completes step 4, replacing the evicted task context with:

The full conversation history has been saved to /conversation_history/<session>.md ...

<summary>
Previous conversation was too long to summarize.
</summary>

The event looks successful but preserves no task state or decisions. The agent may then follow the archive pointer and read back the history it just compacted, adding latency and token usage.

Expected behavior: if bounded trimming cannot retain content from non-empty history, Deep Agents should not commit a successful compaction with the generic fallback. It should retain a safe bounded portion, skip the compaction event, or fail explicitly.

The factory deliberately leaves trim_tokens_to_summarize unset by default, so its default configuration does not take this branch. The bug affects users who opt into this public safety limit and subclasses that inherit the middleware class's bounded default.

Environment / System Info

OS: Linux
Python: 3.14 (not Python-version-specific)
deepagents: 0.7.16
langchain: version resolved by deepagents 0.7.16

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

org:internalIssue or pull request created by a member of the `langchain-ai` GitHub organization.package:deepagentsChanges related to the `deepagents` SDK and agent harness.priority:backlogNot currently planned/prioritized work, often affecting a limited feature or set of users.topic:backendsFilesystem and storage backends for Deep Agents.topic:middlewareMiddleware behavior and composition.type:bugAn unexpected problem or incorrect behavior.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions