Jeserch explores using fast bounded-decision models as a learned attention layer for coding-agent repository context. It indexes source files into deterministic, line-addressed chunks, retrieves candidates with BM25, dense embeddings, native zvec-grep, or their hybrids, and uses Jev to decide which retrieved regions deserve the limited context window. Jev is an evaluation-model API, not a generative coding agent; this project does not claim agent repair success or superiority over cross-encoder rerankers.
repository
↓
zvec retrieval
↓
100 candidate regions
↓
Jev relevance decisions
↓
line-budgeted context
↓
coding model
Use Python 3.12 or newer and Node.js 22 or newer. From the repository root:
uv sync --all-extras --locked
npm ciuv installs the Python package and its locked dependencies. npm ci installs the pinned local @zvec/zvec-grep@0.2.2 runtime used by the Python bridge. Existing Python-only methods (bm25, dense, hybrid, and their Jev variants) do not require Node or npm. Do not install a global zg binary: this repository uses the pinned SDK through src/jeserch/zvec_bridge.mjs, with no server or MCP configuration changes.
The normal and optional Python dependencies are open source and listed in pyproject.toml; the npm dependency is listed in package.json and locked in package-lock.json. The optional semantic extra adds sentence-transformers, Transformers 4-compatible packages, and einops. Dense retrieval uses the pinned nomic-ai/CodeRankEmbed revision in the TOML configs. The hosted Jev service is the only hosted runtime service required for Jev reranking.
For Jev commands, put the key in .env in the current working directory:
AI_GATEWAY_API_KEY=your-keyThe CLI loads this value from the process environment or the .env in the current working directory. It is required for a live Jev request. The key is never written to run configuration or output. Lexical and native retrieval without Jev work without it.
Refresh a repository index before searching:
uv run jeserch index --repo /path/to/repository
uv run jeserch search --repo /path/to/repository \
--task "Find where logout clears the authenticated session" \
--method bm25 --candidates 100 --top-k 20search refreshes the index itself before retrieval, so a separate index command is useful for inspecting or warming the cache but is not required. The default search method is zvec-jev; use --method bm25 or another explicit method when you want a Python baseline or when the Node runtime is unavailable. Native routes are zvec, zvec-fts, and zvec-vector; add -jev to any of them for reranking. --fresh bypasses the local Jev response cache.
Check the Gateway with a small live request:
uv run jeserch jev-check| Command | Purpose |
|---|---|
jeserch index --repo PATH |
Refresh the content-addressed SQLite repository index and print JSON stats. |
jeserch search --repo PATH --task TEXT |
Refresh, retrieve, optionally rerank, pack context, and print ranked chunks. |
jeserch jev-check |
Send the built-in two-candidate connectivity check to Jev. |
jeserch benchmark prepare --suite SUITE |
Download pinned benchmark inputs and write frozen JSONL data. |
jeserch benchmark run --config PATH |
Run or resume a TOML experiment; individual failures remain in results.jsonl. |
jeserch report --run PATH |
Rebuild CSV, summary, plot, and Markdown report artifacts from a run. |
The Python indexer honors Git's tracked and untracked file selection when the repository is a Git worktree. For non-Git directories it applies .gitignore patterns, including nested files, and skips binary files, symlinks, secrets such as .env, lock files, generated/output directories, and files larger than 1 MiB. It keeps tests, configuration, and documentation source text. Python chunks use Tree-sitter; malformed Python and non-Python files use deterministic windows of at most 120 lines with 20-line overlap by default. References are imports and statically inspectable call identifiers, not a complete call graph.
| Method family | Backend and semantics |
|---|---|
bm25 |
Python BM25S lexical retrieval. |
dense |
Python sentence-transformers retrieval with the pinned CodeRankEmbed revision. |
hybrid |
Python lexical/dense fusion. |
hybrid-graph |
Python hybrid retrieval plus heuristic reference expansion. |
zvec |
Native zvec-grep hybrid route: native chunking and the local Potion Code 16M v2 embedding. |
zvec-fts |
Native zvec-grep full-text route. |
zvec-vector |
Native zvec-grep vector route. |
*-jev |
The corresponding candidate set sent through the hosted Jev reranker. |
The zvec bridge returns structured native context items. Jeserch verifies each returned source path, freshness status, original source text, and inclusive line range before constructing a Chunk; it does not parse human-oriented CLI output. Native zvec chunking and the Potion embedding are therefore a different end-to-end retrieval system from the older Python hybrid, rather than a pure vector-store-only ablation. zvec-jev versus zvec is the same native candidate route with the reranker added, and is the direct Jev-effect control. Native indexes and model caches live below the configured cache directory; .zvec-grep is excluded from Python indexing.
The local model is minishlab/potion-code-16M-v2 at immutable revision e9d2a44ca6a05ac6685f3b23709ea57eb7352d5b, dimension 256, cosine metric. The bridge records artifact sizes and SHA-256 values in retrieval metadata. See THIRD_PARTY.md for licensing boundaries; the local package metadata does not establish a license for this model, so this documentation makes no MIT claim for it.
The headline result is from the frozen 356-task SWE-Explore test cohort in this repository. Native zvec and zvec+Jev used the same 100 native candidates; Jev only changed their ordering. Four repository tasks ran in parallel, with a shared limit of 16 in-flight Jev requests. Every one of the 712 method rows completed successfully.
| Method | Candidate core recall | Core recall@500 | Core recall@1,000 | Core recall@2,000 |
|---|---|---|---|---|
| Native zvec | 0.3324 | 0.0813 | 0.1196 | 0.1704 |
| Native zvec + Jev | 0.3325 | 0.1838 | 0.2295 | 0.2778 |
At the 1,000-line budget, the paired gain from Jev is +10.99 percentage points, with a 95% bootstrap interval of [8.93, 13.13]. Candidate recall is unchanged (33.24% → 33.25%), which is the intended result for a reranker: Jev improves the ordering and context packing rather than discovering new candidates.
At 1,000 lines, zvec+Jev also reaches 48.3% core-file hit rate, 62.4% nDCG@100, 48.5% context efficiency, and 72.3% first-useful-hit. The full evaluator reports precision, F1, hit-region, noise, weighted coverage, nDCG, and per-repository metrics.
The run used OpenRouter's typesafe/jev-1.13 decisions endpoint. It made 8,218 Jev requests, processed 47.69M input tokens, and cost approximately $2.00 at OpenRouter's listed $0.042/M input price; Jev produces no billed output tokens. Native zvec query latency was 0.91s p50 / 2.56s p95 after indexing, with cold retriever/index builds at 25.2s p50 / 94.1s p95. zvec+Jev query latency was 2.59s p50 / 5.48s p95 under parallel execution. These are retrieval-run timings, not end-to-end coding-agent latency.
Artifacts: report.md, summary.json, results.jsonl, and the frozen configuration.
The paper evaluates five ranked regions (K=5) across the full 848-instance multilingual benchmark. Those values are useful context, but they are not directly comparable to the 356-task run above, which returns native regions under 500/1,000/2,000-line budgets.
| Explorer | Line recall | HitFile | nDCG@500 | Context efficiency | Downstream resolve rate |
|---|---|---|---|---|---|
| CoSIL | 0.788 | 0.544 | 0.824 | 0.898 | 59.3% |
| AweAgent | 0.140 | 0.682 | 0.954 | 0.829 | 41.3% |
| LocAgent | 0.191 | 0.540 | 0.950 | 0.799 | 44.7% |
| Claude Code | 0.154 | 0.667 | 0.938 | 0.829 | 48.0% |
| Codex | 0.194 | 0.649 | 0.901 | 0.762 | 50.3% |
| BM25 | 0.021 | 0.079 | 0.132 | 0.087 | 12.7% |
| TF-IDF | 0.049 | 0.140 | 0.223 | 0.190 | 26.0% |
| Potion | 0.025 | 0.088 | 0.136 | 0.100 | 23.3% |
The paper's strongest line-level explorer is CoSIL; AweAgent leads file hit rate and nDCG@500, while LocAgent leads first-useful-hit. The paper reports these results in its Tables 3, 5, and 6. Jeserch's fair comparison target is first the matched BM25/TF-IDF/Potion control under the same output protocol, then the native zvec-versus-zvec+Jev paired ablation.
The first native ARB adapter run covered all 58 edit2ripple positives and used ARB's official file-level evaluator, including Recall@K, MRR, and gold_coverage@8k.
| Method | Recall@5 | Recall@10 | Recall@20 | MRR | gold_coverage@8k |
|---|---|---|---|---|---|
| Official BM25 | 0.2069 | 0.2874 | 0.4986 | 0.1536 | 0.0603 |
| Official lexical | 0.4124 | 0.5359 | 0.5876 | 0.2428 | 0.0632 |
| Native zvec | 0.2184 | 0.3190 | 0.3621 | 0.1628 | 0.1782 |
| Native zvec + Jev | 0.3693 | 0.4670 | 0.5244 | 0.2944 | 0.3175 |
Jev adds 15.1 points Recall@5, 14.8 points Recall@10, 16.2 points Recall@20, 13.2 points MRR, and 13.9 points gold coverage over the same native zvec candidate route. Official lexical remains higher on file recall, while zvec+Jev has the strongest MRR and context coverage in this comparison.
Artifacts: summary.json, results.jsonl, and experiments/run_arb_zvec.py. The other three ARB releases have official baselines prepared; their native adapters remain future work.
The official ARB baselines across all 345 positive samples are:
| Task | Ranker | Recall@5 | Recall@10 | Recall@20 | MRR | gold_coverage@8k |
|---|---|---|---|---|---|---|
| code2test (106) | BM25 | 0.1305 | 0.2186 | 0.3066 | 0.1057 | 0.0739 |
| code2test (106) | Lexical | 0.0676 | 0.1399 | 0.2469 | 0.0663 | 0.0299 |
| comment2context (80) | BM25 | 0.2146 | 0.4042 | 0.5292 | 0.1974 | 0.1062 |
| comment2context (80) | Lexical | 0.1667 | 0.3479 | 0.4979 | 0.1530 | 0.0625 |
| trace2code (101) | BM25 | 0.2228 | 0.3218 | 0.4934 | 0.1638 | 0.1205 |
| trace2code (101) | Lexical | 0.3432 | 0.4818 | 0.6964 | 0.2075 | 0.1287 |
| edit2ripple (58) | BM25 | 0.2069 | 0.2874 | 0.4986 | 0.1536 | 0.0603 |
| edit2ripple (58) | Lexical | 0.4124 | 0.5359 | 0.5876 | 0.2428 | 0.0632 |
| Benchmark | Result |
|---|---|
| Requests / SWE-Explore core context, 3 issues | Hybrid + Jev: 57.2% mean core recall@1,000; hybrid + graph + Jev: 63.9%. Descriptive only. |
| CoIR CoSQA, 20 queries | BM25 nDCG@10 0.0829; BM25 + Jev 0.0677 over 19 successful matches. No improvement evidence. |
| Requests native zvec pilot, 3 issues | Native zvec 0.6% mean core recall@1,000; zvec+Jev 51.5%. Descriptive only. |
The three Requests methods were:
| Method | Mean core recall@1,000 | Candidate core recall |
|---|---|---|
| BM25 | 3.0% | 49.7% |
| CodeRankEmbed | 40.9% | 60.2% |
| Hybrid | 28.9% | 81.5% |
| Hybrid + Jev | 57.2% | 81.5% |
| Hybrid + graph | 28.9% | 89.3% |
| Hybrid + graph + Jev | 63.9% | 89.3% |
The Requests issue-level hybrid / Jev / graph+Jev scores were 18.8% / 53.1% / 53.1%, 9.0% / 44.5% / 44.5%, and 58.8% / 74.0% / 94.1%. The three-issue native zvec pilot used Vercel Jev and measured 18.4 seconds mean native retrieval versus 37.3 seconds mean zvec+Jev reranking, across 886 Jev requests and 5.52M input tokens. These pilots are retained for engineering diagnostics and should not be presented as general benchmark wins.
from pathlib import Path
from jeserch.index import RepositoryIndex, chunk_file
chunks = chunk_file("src/auth.py", source_text, max_lines=120, overlap=20)
index = RepositoryIndex(Path("."))
stats = index.refresh()
chunks = index.chunks()
print(stats["snapshot"])
index.close()Chunk and Region are Pydantic models in jeserch.models. A Chunk has path, inclusive start_line/end_line, deterministic id, source text, symbol, kind, content_hash, and references. SearchResult carries the snapshot, method, ranked candidates/results, timings, usage, and metadata. SearchEngine supports the Python and native methods listed above.
RepositoryIndex.refresh() returns snapshot, file count, chunk count, changed-file count, and deleted-file count. The default SQLite cache is .jeserch/index.sqlite3; Jev responses, dense embeddings, and native zvec state are cached separately. Snapshot and chunk IDs are content-addressed and stable for the same path, range, and content.
src/jeserch/ index, retrieval, zvec adapter/bridge, Jev, benchmark, metrics, report CLI
experiments/ frozen TOML run configurations
data/ prepared benchmark JSONL, manifests, and local repository snapshots
runs/ recorded configs, results, summaries, plots, reports, and source archives
tests/ unit and integration tests
package.json pinned native zvec-grep runtime
Benchmark source archives include the Python source, .mjs bridge, package.json, package-lock.json, and Python lock/config files needed to identify the runtime. There is no daemon, language server, or background index process. The graph-assisted method uses heuristic references and is not a complete static call graph.
The frozen SWE-Explore run uses revision bdb0ae45d7c337d9e1dc3ebfe2a0af6bc7c1fbd9, the Verified-backed Python test cohort, native zvec-grep 0.2.2, Potion Code 16M v2, 100 candidates, and 500/1,000/2,000-line budgets. It does not claim results for the full 848-instance multilingual paper benchmark. ARB uses its own official file-level evaluator and must not be mixed with SWE-Explore line-level metrics.
Jev is a bounded relevance decision model. It does not generate code, replace a coding agent, or establish downstream patch success. The results show repository-context retrieval and reranking behavior, not improved end-to-end repair success. A controlled coding-agent repair study remains future work. Native zvec and Python hybrid retrieval use different chunkers and embeddings, so their comparison is end-to-end. The direct causal ablations are zvec+Jev versus zvec, and hybrid+Jev versus hybrid under the same candidate set.
See docs/EXPERIMENTS.md for frozen datasets, label isolation, commands, controls, and artifact interpretation. See docs/RESULTS.md for the complete metric inventory and development history.
- 64 automated tests pass.
- Ruff and Node syntax checks pass.
- Native zvec paths, freshness, source ranges, and model metadata are validated before results are accepted.
- Jev response IDs, typed probabilities, retries, usage, and oversized-source splitting are validated.
- Benchmark runs preserve successful rows, explicit failures, configs, source archives, summaries, and reports for auditability.