Repository navigation
fix(langgraph): keep dev-runtime idle pickle writes off the host NVMe - #4
Merged
Merged
Conversation
`langgraph dev` always runs the in-memory runtime (it hardcodes :memory: and never reads DATABASE_URI). That runtime's daemon flush loop (langgraph_runtime_inmem 0.30.0 _persistence.py, _flush_interval=10) calls PersistentDict.sync() on every store every 10s unconditionally, and sync() (langgraph_checkpoint 4.1.1 memory backend) re-pickles the ENTIRE dict and atomically replaces the file with no dirty check and no pruning. The blobs dict (.langgraph_checkpoint.3.pckl) is only pruned by explicit delete_thread(). Measured on VM 150 (idle): the .3.pckl blob reached 198 MB, rewritten every ~10s = 26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear. The runtime's own disable switch does NOT work through `langgraph dev` in the pinned versions (verified on this box, see docs/adr/0013): langgraph-cli 0.4.29 validate_config_file() strips the unknown "disable_persistence" key, and run_server() then overwrites LANGGRAPH_DISABLE_FILE_PERSISTENCE to "false" in the worker via patch_environment. A container env var and a .env entry are overridden the same way. There is also no knob for the 10s interval. Fix: back /app/.langgraph_api with a 1 GiB size-capped tmpfs in the compose langgraph service. The runtime still re-pickles every 10s, but into RAM, so the NVMe sees nothing. Verified idle for ~60s each on one image build: without the tmpfs the container block-IO write climbed to 277 MB; with it, block-IO write stayed 0 B while the pickle was still rewritten every 10s. Also set LANGGRAPH_DISABLE_FILE_PERSISTENCE=true as a forward-compatible no-op that takes over if upstream honors it. langgraph.json left unchanged (unknown key is stripped now and could be rejected by a stricter future validator). Rejected: the disable switch (proven non-functional here); a Postgres checkpointer (unreachable from `langgraph dev`); throttling the interval (hardcoded constant); a custom entrypoint replicating dev (fragile across bumps). Trade-off: langgraph run history no longer survives a container restart (the tmpfs clears). It already did not survive a recreate (/app/.langgraph_api was never volume-mounted) and the stack already ran :memory:; engagement evidence persists separately. Residual: the runtime still burns CPU/RAM re-pickling; only the disk cost is removed, pending an upstream fix. Guard: tests/unit/ops/test_langgraph_persistence_config.py asserts the shipped compose mounts the size-capped tmpfs and sets the forward-compat flag; it fails on the old config.
Exploitacious
changed the base branch from
sync/upstream-2026-09-26
to
main
September 27, 2026 15:36
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #3 (base branch
sync/upstream-2026-09-26). Fixes the idle disk abuse measured on VM 150.Root cause (file:line, pinned versions)
langgraph devalways runs the in-memory runtime (it hardcodes:memory:and never readsDATABASE_URI). That runtime's daemon flush loop (langgraph_runtime_inmem0.30.0_persistence.py:_flush_interval = 10,_flush_loop) callsPersistentDict.sync()on every store every 10s unconditionally.sync()(langgraph_checkpoint4.1.1langgraph/checkpoint/memory/__init__.py) re-pickles the ENTIRE dict and atomically replaces the file, with no dirty flag anywhere in the class and no pruning. Theblobsdict (.langgraph_checkpoint.3.pckl) is only pruned by an explicitdelete_thread().Measured on VM 150 (idle):
.3.pcklreached 198 MB, rewritten every ~10s = 26.9 MiB/s, 32.3 TB in 18 days, 0.44%/day of NVMe wear.Why the runtime's own disable switch does not work here (verified on this box)
LANGGRAPH_DISABLE_FILE_PERSISTENCE/disable_persistencecannot be set throughlanggraph devin the pinned versions:langgraph-cli0.4.29validate_config_file()strips the unknowndisable_persistencekey (the validated config's keys do not contain it), soconfig_json.get("disable_persistence", False)is alwaysFalse.run_server()then buildsLANGGRAPH_DISABLE_FILE_PERSISTENCE=str(disable_persistence).lower()=falseand applies it withpatch_environment(), which refuses to overwrite from the loaded env. Observed directly: the worker subprocess env showed the flagfalseeven when the container was launched with ittrue, and the.pcklfiles were still written.There is also no knob for the 10s interval (
_flush_intervalis a hardcoded constant).Fix
Back
/app/.langgraph_apiwith a 1 GiB size-capped tmpfs in the composelanggraphservice. The runtime still re-pickles every 10s, but into RAM, so the NVMe sees nothing.LANGGRAPH_DISABLE_FILE_PERSISTENCE=trueis set as a forward-compatible no-op that takes over if upstream honors it.langgraph.jsonis left unchanged (an unknown key is stripped now and could be rejected by a stricter future validator). Design and rejected alternatives:docs/adr/0013.Before/after (this box, one image build, idle ~60s each)
.langgraph_ops.pckl(8.4 MB after one seeded run) rewritten every 10s; container block-IO write climbed to 277 MB.mountconfirmstmpfs on /app/.langgraph_api.Guard
packages/decepticon/tests/unit/ops/test_langgraph_persistence_config.pyasserts the shipped compose mounts the size-capped tmpfs and sets the forward-compat flag. It fails on the old config (no tmpfs / no flag) and passes here.Checks
uv run pytest -n auto -q -m "not slow": 5179 passed, 44 skipped.Trade-off / open item for maintainer
LangGraph run history no longer survives a container restart (the tmpfs clears). It already did not survive a recreate (
/app/.langgraph_apiwas never volume-mounted) and the stack already ran:memory:; engagement evidence persists separately. Residual: the runtime still burns CPU/RAM re-pickling every 10s; only the disk cost is removed. The complete fix is upstream honoring the disable switch. Do not merge without maintainer review.