Skip to content

[Bug]: ID map RocksDB is never compacted; ascending primary keys leave one SST and one open fd per flush #784

Description

@JasonXuDeveloper

Description

A collection's ID map (idmap.0/) is a RocksDB instance that zvec flushes but never compacts. When primary keys only increase (UUIDv7, ULID, timestamps, auto-increment counters), each flush() leaves one more SST file in the ID map. RocksDB keeps every SST file open (max_open_files=-1), so the process's file descriptor count grows linearly with the number of flushes. optimize() does not reduce it, and reopening the collection opens all of the files again.

With random keys the same workload stays at a handful of files.

What the code does on main (53c1bb6):

  • SegmentImpl flushes the ID map on every segment flush: id_map_->flush() at src/db/index/segment/segment.cc:2216-2217, which calls RocksdbContext::flush() (src/db/index/common/id_map.cc:82-94).
  • IDMap has no compaction method (src/db/index/common/id_map.h). The only caller of RocksdbContext::compact() is the inverted index (src/db/index/column/inverted_column/inverted_indexer.cc:143). CollectionImpl::optimize() (src/db/collection.cc:913) does not touch id_map_. The ID map path suffix is always 0 (src/db/collection.cc:2301), so the map is never rewritten either.
  • The ID map is opened with the shared defaults from RocksdbContext::prepare_options() (src/db/common/rocksdb_context.cc:241-306): OptimizeLevelStyleCompaction() with its default 512 MiB budget (line 247), so L0 triggers compaction at 2 files and max_bytes_for_level_base is 512 MiB. max_open_files is left at -1.

With ascending keys, a new L0 file never overlaps anything already in L1, so RocksDB moves it down without rewriting it. L1 stays far below 512 MiB, so nothing is merged after that.

Expected: the ID map's file count stays bounded, as it does with random keys, or at least optimize() compacts it.

Note that just calling RocksdbContext::compact() would not be enough. It uses a default CompactRangeOptions (src/db/common/rocksdb_context.h:52), whose bottommost_level_compaction is kIfHaveCompactionFilter. With no compaction filter, RocksDB 8.1.1 skips the bottommost level under that setting (db/db_impl/db_impl_compaction_flush.cc:1160-1172), and that is exactly where these files sit. We first checked the RocksDB behaviour outside zvec, using RocksDB 9.8.4 through the rocksdict Python binding with the same options (OptimizeLevelStyleCompaction(512 MiB), 200 flushes of 64 ascending keys each). This is a simulation, not zvec's RocksDB 8.1.1:

Remedy (simulation, RocksDB 9.8.4) SSTs before SSTs after
none 75 75
CompactRange(nullptr, nullptr), default options 75 75
CompactRange, bottommost_level_compaction = kForce 72 1
level_compaction_dynamic_level_bytes = true, no manual compaction 74 74

Confirmed in zvec's own build (main 53c1bb6, bundled RocksDB 8.1.1, macOS arm64, Release), 50 flushes of 64 ascending keys, then optimize():

Remedy (zvec build, RocksDB 8.1.1) SSTs before SSTs after
none (optimize() as on main) 50 50
ID-map CompactRange in optimize(), default options 50 50
same, bottommost_level_compaction = kForceOptimized 50 1

A fix is in the linked PR.

Steps to Reproduce

# pip install zvec==0.7.0 ; python idmap_repro.py ascending ; python idmap_repro.py random
import glob, os, sys, tempfile
import zvec
from zvec import CollectionOption, CollectionSchema, DataType, Doc, FieldSchema

mode = sys.argv[1]  # "ascending" or "random"
path = os.path.join(tempfile.mkdtemp(), "c")
schema = CollectionSchema(name="idmap", fields=[FieldSchema("v", DataType.INT64)])
opt = CollectionOption(read_only=False, enable_mmap=True)
coll = zvec.create_and_open(path=path, schema=schema, option=opt)

n = 0
for _ in range(200):                      # 200 rounds of: upsert 64 new docs, flush
    docs = []
    for _ in range(64):
        # ascending keys behave like UUIDv7 / ULID / timestamps / auto-increment
        pk = f"{n:016d}" if mode == "ascending" else os.urandom(8).hex()
        docs.append(Doc(id=pk, fields={"v": n}))
        n += 1
    coll.upsert(docs)
    coll.flush()

def ssts():
    return len(glob.glob(f"{path}/idmap.0/*.sst"))

def idmap_fds():
    fds = []
    for fd in os.listdir("/proc/self/fd"):
        try:
            fds.append(os.readlink(f"/proc/self/fd/{fd}"))
        except OSError:
            pass
    return sum("/idmap.0/" in f for f in fds)

print(f"{mode}: after 200 flushes  idmap.0 SSTs={ssts()}  fds into idmap.0={idmap_fds()}")
coll.optimize()
print(f"{mode}: after optimize()   idmap.0 SSTs={ssts()}  fds into idmap.0={idmap_fds()}")
del coll
coll = zvec.open(path=path, option=opt)
print(f"{mode}: after reopen       idmap.0 SSTs={ssts()}  fds into idmap.0={idmap_fds()}")

Logs / Stack Trace

Output of the script above with zvec 0.7.0 from PyPI. One representative run of each mode:

ascending: after 200 flushes  idmap.0 SSTs=196  fds into idmap.0=200
ascending: after optimize()   idmap.0 SSTs=196  fds into idmap.0=200
ascending: after reopen       idmap.0 SSTs=196  fds into idmap.0=200
random: after 200 flushes  idmap.0 SSTs=4  fds into idmap.0=8
random: after optimize()   idmap.0 SSTs=4  fds into idmap.0=8
random: after reopen       idmap.0 SSTs=4  fds into idmap.0=8
  • Across 4 runs of this script, ascending ended with 175, 184, 196 and 195 ID-map SSTs after 200 flushes. random ended with 4, 2 and 4 in 3 runs.
  • optimize() never changed the count.
  • In an application that flushes every 64 documents, running zvec 1ab7975 (through zvec-rust 0.7.2), 200 flushes with ascending keys left 200 ID-map SSTs. Reopening the whole collection after optimize() held 215 fds. The same workload with non-ascending keys ended with 1 ID-map SST and 72 fds.

idmap.0/OPTIONS-* from the run above (excerpt):

max_open_files=-1
write_buffer_size=134217728
max_write_buffer_number=6
level0_file_num_compaction_trigger=2
max_bytes_for_level_base=536870912
level_compaction_dynamic_level_bytes=false
periodic_compaction_seconds=0

Operating System

Ubuntu 24.04.4 LTS, x86_64, 4 vCPU, 16 GiB RAM (repro and application measurement); macOS arm64 for the source-build confirmation above.

Build & Runtime Environment

  • Repro above: Python 3.11.15, zvec 0.7.0 wheel from PyPI (cp311-cp311-manylinux_2_28_x86_64).
  • Application measurement: zvec 1ab7975, from the zvec-rust 0.7.2 prebuilt libzvec_c_api.so, with the bundled RocksDB 8.1.1.
  • Source-build confirmation: main at 53c1bb6, CMAKE_BUILD_TYPE=Release, AppleClang.
  • Code references were checked against main at 53c1bb6 (2026-09-24). rocksdb_context.{h,cc}, id_map.cc, segment.cc and collection.cc are unchanged between 1ab7975 and 53c1bb6, and rocksdb_context.cc and id_map.cc are also unchanged since v0.7.0.
  • Open PR feat(buffer): integrate database storage with shared memory budgets #770 does not change this path: it adds no compaction of the ID map.

Additional Context

  • I've checked git status — no uncommitted submodule changes (source-build confirmation at 53c1bb6)
  • I built with CMAKE_BUILD_TYPE=Debug (not tried: Release and PyPI wheel only)
  • This occurs with or without COVERAGE=ON (not tried)
  • The issue involves Python ↔ C++ integration (pybind11) (no: the repro uses the Python API, but the bug is in the C++ RocksDB options)

The repro and application numbers come from prebuilt binaries (the PyPI wheel and the zvec-rust prebuilt); the source-build confirmation used a Release build, so these items do not apply.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

  • Status
    In progress

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions