Skip to content

fix(langgraph): keep dev-runtime idle pickle writes off the host NVMe - #4

Merged
Exploitacious merged 2 commits into
mainfrom
fix/langgraph-idle-writes
Sep 27, 2026
Merged

Exploitacious merged 2 commits into
mainfrom
fix/langgraph-idle-writes

Conversation

@Exploitacious

Copy link
Copy Markdown
Owner

Stacked on #3 (base branch sync/upstream-2026-09-26). Fixes the idle disk abuse measured on VM 150.

Root cause (file:line, pinned versions)

langgraph dev always runs the in-memory runtime (it hardcodes :memory: and never reads DATABASE_URI). That runtime's daemon flush loop (langgraph_runtime_inmem 0.30.0 _persistence.py: _flush_interval = 10, _flush_loop) calls PersistentDict.sync() on every store every 10s unconditionally. sync() (langgraph_checkpoint 4.1.1 langgraph/checkpoint/memory/__init__.py) re-pickles the ENTIRE dict and atomically replaces the file, with no dirty flag anywhere in the class and no pruning. The blobs dict (.langgraph_checkpoint.3.pckl) is only pruned by an explicit delete_thread().

Measured on VM 150 (idle): .3.pckl reached 198 MB, rewritten every ~10s = 26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear.

Why the runtime's own disable switch does not work here (verified on this box)

LANGGRAPH_DISABLE_FILE_PERSISTENCE / disable_persistence cannot be set through langgraph dev in the pinned versions:

  • langgraph-cli 0.4.29 validate_config_file() strips the unknown disable_persistence key (the validated config's keys do not contain it), so config_json.get("disable_persistence", False) is always False.
  • run_server() then builds LANGGRAPH_DISABLE_FILE_PERSISTENCE=str(disable_persistence).lower() = false and applies it with patch_environment(), which refuses to overwrite from the loaded env. Observed directly: the worker subprocess env showed the flag false even when the container was launched with it true, and the .pckl files were still written.

There is also no knob for the 10s interval (_flush_interval is a hardcoded constant).

Fix

Back /app/.langgraph_api with a 1 GiB size-capped tmpfs in the compose langgraph service. The runtime still re-pickles every 10s, but into RAM, so the NVMe sees nothing. LANGGRAPH_DISABLE_FILE_PERSISTENCE=true is set as a forward-compatible no-op that takes over if upstream honors it. langgraph.json is left unchanged (an unknown key is stripped now and could be rejected by a stricter future validator). Design and rejected alternatives: docs/adr/0013.

Before/after (this box, one image build, idle ~60s each)

  • Before (no tmpfs): .langgraph_ops.pckl (8.4 MB after one seeded run) rewritten every 10s; container block-IO write climbed to 277 MB.
  • After (tmpfs mounted): same file still rewritten every 10s into RAM; tmpfs held ~8 MB; container block-IO write stayed 0 B. mount confirms tmpfs on /app/.langgraph_api.

Guard

packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.py asserts the shipped compose mounts the size-capped tmpfs and sets the forward-compat flag. It fails on the old config (no tmpfs / no flag) and passes here.

Checks

  • uv run pytest -n auto -q -m "not slow": 5179 passed, 44 skipped.

Trade-off / open item for maintainer

LangGraph run history no longer survives a container restart (the tmpfs clears). It already did not survive a recreate (/app/.langgraph_api was never volume-mounted) and the stack already ran :memory:; engagement evidence persists separately. Residual: the runtime still burns CPU/RAM re-pickling every 10s; only the disk cost is removed. The complete fix is upstream honoring the disable switch. Do not merge without maintainer review.

`langgraph dev` always runs the in-memory runtime (it hardcodes :memory: and
never reads DATABASE_URI). That runtime's daemon flush loop
(langgraph_runtime_inmem 0.30.0 _persistence.py, _flush_interval=10) calls
PersistentDict.sync() on every store every 10s unconditionally, and sync()
(langgraph_checkpoint 4.1.1 memory backend) re-pickles the ENTIRE dict and
atomically replaces the file with no dirty check and no pruning. The blobs dict
(.langgraph_checkpoint.3.pckl) is only pruned by explicit delete_thread().

Measured on VM 150 (idle): the .3.pckl blob reached 198 MB, rewritten every
~10s = 26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear.

The runtime's own disable switch does NOT work through `langgraph dev` in the
pinned versions (verified on this box, see docs/adr/0013): langgraph-cli 0.4.29
validate_config_file() strips the unknown "disable_persistence" key, and
run_server() then overwrites LANGGRAPH_DISABLE_FILE_PERSISTENCE to "false" in
the worker via patch_environment. A container env var and a .env entry are
overridden the same way. There is also no knob for the 10s interval.

Fix: back /app/.langgraph_api with a 1 GiB size-capped tmpfs in the compose
langgraph service. The runtime still re-pickles every 10s, but into RAM, so the
NVMe sees nothing. Verified idle for ~60s each on one image build: without the
tmpfs the container block-IO write climbed to 277 MB; with it, block-IO write
stayed 0 B while the pickle was still rewritten every 10s. Also set
LANGGRAPH_DISABLE_FILE_PERSISTENCE=true as a forward-compatible no-op that takes
over if upstream honors it. langgraph.json left unchanged (unknown key is
stripped now and could be rejected by a stricter future validator).

Rejected: the disable switch (proven non-functional here); a Postgres
checkpointer (unreachable from `langgraph dev`); throttling the interval
(hardcoded constant); a custom entrypoint replicating dev (fragile across bumps).

Trade-off: langgraph run history no longer survives a container restart (the
tmpfs clears). It already did not survive a recreate (/app/.langgraph_api was
never volume-mounted) and the stack already ran :memory:; engagement evidence
persists separately. Residual: the runtime still burns CPU/RAM re-pickling; only
the disk cost is removed, pending an upstream fix.

Guard: tests/unit/ops/test_langgraph_persistence_config.py asserts the shipped
compose mounts the size-capped tmpfs and sets the forward-compat flag; it fails
on the old config.
@Exploitacious
Exploitacious changed the base branch from sync/upstream-2026-09-26 to main September 27, 2026 15:36
@Exploitacious Exploitacious reopened this Sep 27, 2026
@Exploitacious
Exploitacious merged commit dbce695 into main Sep 27, 2026
28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant