Skip to content
chengusPublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Jeserch

Jeserch explores using fast bounded-decision models as a learned attention layer for coding-agent repository context. It indexes source files into deterministic, line-addressed chunks, retrieves candidates with BM25, dense embeddings, native zvec-grep, or their hybrids, and uses Jev to decide which retrieved regions deserve the limited context window. Jev is an evaluation-model API, not a generative coding agent; this project does not claim agent repair success or superiority over cross-encoder rerankers.

repository
    ↓
zvec retrieval
    ↓
100 candidate regions
    ↓
Jev relevance decisions
    ↓
line-budgeted context
    ↓
coding model

Install

Use Python 3.12 or newer and Node.js 22 or newer. From the repository root:

uv sync --all-extras --locked
npm ci

uv installs the Python package and its locked dependencies. npm ci installs the pinned local @zvec/zvec-grep@0.2.2 runtime used by the Python bridge. Existing Python-only methods (bm25, dense, hybrid, and their Jev variants) do not require Node or npm. Do not install a global zg binary: this repository uses the pinned SDK through src/jeserch/zvec_bridge.mjs, with no server or MCP configuration changes.

The normal and optional Python dependencies are open source and listed in pyproject.toml; the npm dependency is listed in package.json and locked in package-lock.json. The optional semantic extra adds sentence-transformers, Transformers 4-compatible packages, and einops. Dense retrieval uses the pinned nomic-ai/CodeRankEmbed revision in the TOML configs. The hosted Jev service is the only hosted runtime service required for Jev reranking.

For Jev commands, put the key in .env in the current working directory:

AI_GATEWAY_API_KEY=your-key

The CLI loads this value from the process environment or the .env in the current working directory. It is required for a live Jev request. The key is never written to run configuration or output. Lexical and native retrieval without Jev work without it.

Quick start

Refresh a repository index before searching:

uv run jeserch index --repo /path/to/repository
uv run jeserch search --repo /path/to/repository \
  --task "Find where logout clears the authenticated session" \
  --method bm25 --candidates 100 --top-k 20

search refreshes the index itself before retrieval, so a separate index command is useful for inspecting or warming the cache but is not required. The default search method is zvec-jev; use --method bm25 or another explicit method when you want a Python baseline or when the Node runtime is unavailable. Native routes are zvec, zvec-fts, and zvec-vector; add -jev to any of them for reranking. --fresh bypasses the local Jev response cache.

Check the Gateway with a small live request:

uv run jeserch jev-check

CLI commands

Command Purpose
jeserch index --repo PATH Refresh the content-addressed SQLite repository index and print JSON stats.
jeserch search --repo PATH --task TEXT Refresh, retrieve, optionally rerank, pack context, and print ranked chunks.
jeserch jev-check Send the built-in two-candidate connectivity check to Jev.
jeserch benchmark prepare --suite SUITE Download pinned benchmark inputs and write frozen JSONL data.
jeserch benchmark run --config PATH Run or resume a TOML experiment; individual failures remain in results.jsonl.
jeserch report --run PATH Rebuild CSV, summary, plot, and Markdown report artifacts from a run.

The Python indexer honors Git's tracked and untracked file selection when the repository is a Git worktree. For non-Git directories it applies .gitignore patterns, including nested files, and skips binary files, symlinks, secrets such as .env, lock files, generated/output directories, and files larger than 1 MiB. It keeps tests, configuration, and documentation source text. Python chunks use Tree-sitter; malformed Python and non-Python files use deterministic windows of at most 120 lines with 20-line overlap by default. References are imports and statically inspectable call identifiers, not a complete call graph.

Retrieval methods

Method family Backend and semantics
bm25 Python BM25S lexical retrieval.
dense Python sentence-transformers retrieval with the pinned CodeRankEmbed revision.
hybrid Python lexical/dense fusion.
hybrid-graph Python hybrid retrieval plus heuristic reference expansion.
zvec Native zvec-grep hybrid route: native chunking and the local Potion Code 16M v2 embedding.
zvec-fts Native zvec-grep full-text route.
zvec-vector Native zvec-grep vector route.
*-jev The corresponding candidate set sent through the hosted Jev reranker.

The zvec bridge returns structured native context items. Jeserch verifies each returned source path, freshness status, original source text, and inclusive line range before constructing a Chunk; it does not parse human-oriented CLI output. Native zvec chunking and the Potion embedding are therefore a different end-to-end retrieval system from the older Python hybrid, rather than a pure vector-store-only ablation. zvec-jev versus zvec is the same native candidate route with the reranker added, and is the direct Jev-effect control. Native indexes and model caches live below the configured cache directory; .zvec-grep is excluded from Python indexing.

The local model is minishlab/potion-code-16M-v2 at immutable revision e9d2a44ca6a05ac6685f3b23709ea57eb7352d5b, dimension 256, cosine metric. The bridge records artifact sizes and SHA-256 values in retrieval metadata. See THIRD_PARTY.md for licensing boundaries; the local package metadata does not establish a license for this model, so this documentation makes no MIT claim for it.

Results

The headline result is from the frozen 356-task SWE-Explore test cohort in this repository. Native zvec and zvec+Jev used the same 100 native candidates; Jev only changed their ordering. Four repository tasks ran in parallel, with a shared limit of 16 in-flight Jev requests. Every one of the 712 method rows completed successfully.

SWE-Explore held-out test

Method Candidate core recall Core recall@500 Core recall@1,000 Core recall@2,000
Native zvec 0.3324 0.0813 0.1196 0.1704
Native zvec + Jev 0.3325 0.1838 0.2295 0.2778

At the 1,000-line budget, the paired gain from Jev is +10.99 percentage points, with a 95% bootstrap interval of [8.93, 13.13]. Candidate recall is unchanged (33.24% → 33.25%), which is the intended result for a reranker: Jev improves the ordering and context packing rather than discovering new candidates.

At 1,000 lines, zvec+Jev also reaches 48.3% core-file hit rate, 62.4% nDCG@100, 48.5% context efficiency, and 72.3% first-useful-hit. The full evaluator reports precision, F1, hit-region, noise, weighted coverage, nDCG, and per-repository metrics.

The run used OpenRouter's typesafe/jev-1.13 decisions endpoint. It made 8,218 Jev requests, processed 47.69M input tokens, and cost approximately $2.00 at OpenRouter's listed $0.042/M input price; Jev produces no billed output tokens. Native zvec query latency was 0.91s p50 / 2.56s p95 after indexing, with cold retriever/index builds at 25.2s p50 / 94.1s p95. zvec+Jev query latency was 2.59s p50 / 5.48s p95 under parallel execution. These are retrieval-run timings, not end-to-end coding-agent latency.

Artifacts: report.md, summary.json, results.jsonl, and the frozen configuration.

SWE-Explore paper reference results

The paper evaluates five ranked regions (K=5) across the full 848-instance multilingual benchmark. Those values are useful context, but they are not directly comparable to the 356-task run above, which returns native regions under 500/1,000/2,000-line budgets.

Explorer Line recall HitFile nDCG@500 Context efficiency Downstream resolve rate
CoSIL 0.788 0.544 0.824 0.898 59.3%
AweAgent 0.140 0.682 0.954 0.829 41.3%
LocAgent 0.191 0.540 0.950 0.799 44.7%
Claude Code 0.154 0.667 0.938 0.829 48.0%
Codex 0.194 0.649 0.901 0.762 50.3%
BM25 0.021 0.079 0.132 0.087 12.7%
TF-IDF 0.049 0.140 0.223 0.190 26.0%
Potion 0.025 0.088 0.136 0.100 23.3%

The paper's strongest line-level explorer is CoSIL; AweAgent leads file hit rate and nDCG@500, while LocAgent leads first-useful-hit. The paper reports these results in its Tables 3, 5, and 6. Jeserch's fair comparison target is first the matched BM25/TF-IDF/Potion control under the same output protocol, then the native zvec-versus-zvec+Jev paired ablation.

Agent Retrieval Bench: edit2ripple

The first native ARB adapter run covered all 58 edit2ripple positives and used ARB's official file-level evaluator, including Recall@K, MRR, and gold_coverage@8k.

Method Recall@5 Recall@10 Recall@20 MRR gold_coverage@8k
Official BM25 0.2069 0.2874 0.4986 0.1536 0.0603
Official lexical 0.4124 0.5359 0.5876 0.2428 0.0632
Native zvec 0.2184 0.3190 0.3621 0.1628 0.1782
Native zvec + Jev 0.3693 0.4670 0.5244 0.2944 0.3175

Jev adds 15.1 points Recall@5, 14.8 points Recall@10, 16.2 points Recall@20, 13.2 points MRR, and 13.9 points gold coverage over the same native zvec candidate route. Official lexical remains higher on file recall, while zvec+Jev has the strongest MRR and context coverage in this comparison.

Artifacts: summary.json, results.jsonl, and experiments/run_arb_zvec.py. The other three ARB releases have official baselines prepared; their native adapters remain future work.

The official ARB baselines across all 345 positive samples are:

Task Ranker Recall@5 Recall@10 Recall@20 MRR gold_coverage@8k
code2test (106) BM25 0.1305 0.2186 0.3066 0.1057 0.0739
code2test (106) Lexical 0.0676 0.1399 0.2469 0.0663 0.0299
comment2context (80) BM25 0.2146 0.4042 0.5292 0.1974 0.1062
comment2context (80) Lexical 0.1667 0.3479 0.4979 0.1530 0.0625
trace2code (101) BM25 0.2228 0.3218 0.4934 0.1638 0.1205
trace2code (101) Lexical 0.3432 0.4818 0.6964 0.2075 0.1287
edit2ripple (58) BM25 0.2069 0.2874 0.4986 0.1536 0.0603
edit2ripple (58) Lexical 0.4124 0.5359 0.5876 0.2428 0.0632

Development pilots

Benchmark Result
Requests / SWE-Explore core context, 3 issues Hybrid + Jev: 57.2% mean core recall@1,000; hybrid + graph + Jev: 63.9%. Descriptive only.
CoIR CoSQA, 20 queries BM25 nDCG@10 0.0829; BM25 + Jev 0.0677 over 19 successful matches. No improvement evidence.
Requests native zvec pilot, 3 issues Native zvec 0.6% mean core recall@1,000; zvec+Jev 51.5%. Descriptive only.

The three Requests methods were:

Method Mean core recall@1,000 Candidate core recall
BM25 3.0% 49.7%
CodeRankEmbed 40.9% 60.2%
Hybrid 28.9% 81.5%
Hybrid + Jev 57.2% 81.5%
Hybrid + graph 28.9% 89.3%
Hybrid + graph + Jev 63.9% 89.3%

The Requests issue-level hybrid / Jev / graph+Jev scores were 18.8% / 53.1% / 53.1%, 9.0% / 44.5% / 44.5%, and 58.8% / 74.0% / 94.1%. The three-issue native zvec pilot used Vercel Jev and measured 18.4 seconds mean native retrieval versus 37.3 seconds mean zvec+Jev reranking, across 886 Jev requests and 5.52M input tokens. These pilots are retained for engineering diagnostics and should not be presented as general benchmark wins.

Python interfaces

from pathlib import Path

from jeserch.index import RepositoryIndex, chunk_file

chunks = chunk_file("src/auth.py", source_text, max_lines=120, overlap=20)

index = RepositoryIndex(Path("."))
stats = index.refresh()
chunks = index.chunks()
print(stats["snapshot"])
index.close()

Chunk and Region are Pydantic models in jeserch.models. A Chunk has path, inclusive start_line/end_line, deterministic id, source text, symbol, kind, content_hash, and references. SearchResult carries the snapshot, method, ranked candidates/results, timings, usage, and metadata. SearchEngine supports the Python and native methods listed above.

RepositoryIndex.refresh() returns snapshot, file count, chunk count, changed-file count, and deleted-file count. The default SQLite cache is .jeserch/index.sqlite3; Jev responses, dense embeddings, and native zvec state are cached separately. Snapshot and chunk IDs are content-addressed and stable for the same path, range, and content.

Repository layout

src/jeserch/       index, retrieval, zvec adapter/bridge, Jev, benchmark, metrics, report CLI
experiments/        frozen TOML run configurations
data/               prepared benchmark JSONL, manifests, and local repository snapshots
runs/               recorded configs, results, summaries, plots, reports, and source archives
tests/              unit and integration tests
package.json        pinned native zvec-grep runtime

Benchmark source archives include the Python source, .mjs bridge, package.json, package-lock.json, and Python lock/config files needed to identify the runtime. There is no daemon, language server, or background index process. The graph-assisted method uses heuristic references and is not a complete static call graph.

Reproducibility and scope

The frozen SWE-Explore run uses revision bdb0ae45d7c337d9e1dc3ebfe2a0af6bc7c1fbd9, the Verified-backed Python test cohort, native zvec-grep 0.2.2, Potion Code 16M v2, 100 candidates, and 500/1,000/2,000-line budgets. It does not claim results for the full 848-instance multilingual paper benchmark. ARB uses its own official file-level evaluator and must not be mixed with SWE-Explore line-level metrics.

Jev is a bounded relevance decision model. It does not generate code, replace a coding agent, or establish downstream patch success. The results show repository-context retrieval and reranking behavior, not improved end-to-end repair success. A controlled coding-agent repair study remains future work. Native zvec and Python hybrid retrieval use different chunkers and embeddings, so their comparison is end-to-end. The direct causal ablations are zvec+Jev versus zvec, and hybrid+Jev versus hybrid under the same candidate set.

See docs/EXPERIMENTS.md for frozen datasets, label isolation, commands, controls, and artifact interpretation. See docs/RESULTS.md for the complete metric inventory and development history.

Validation

  • 64 automated tests pass.
  • Ruff and Node syntax checks pass.
  • Native zvec paths, freshness, source ranges, and model metadata are validated before results are accepted.
  • Jev response IDs, typed probabilities, retries, usage, and oversized-source splitting are validated.
  • Benchmark runs preserve successful rows, explicit failures, configs, source archives, summaries, and reports for auditability.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages