Conversation
The `mcp>=1.1.0` requirement had no upper bound, so a fresh install now resolves to mcp 2.2.0 - an SDK that removed the v1 decorator handler API this release is built on. Pin to `mcp>=1.28,<2.0.0`. Also gives notice, in both the CHANGELOG and the README, that 1.0.0 is a breaking release: it requires mcp>=2.2.0 and the 2026-07-28 protocol, moves the remote transport from /sse to /mcp, and changes resource errors from -32002 to -32602. Users who want to stay on this line pin mcp-nvidia<1.0.0. Verified against mcp 1.30.0 (latest 1.x): 47 passed, 11 skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MsM3KEauXym5MWGjbRiqD9
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueNo actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (4)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe project moves to version 0.9.0 and pins the MCP SDK to the supported 1.x range. The changelog and README document the planned 1.0.0 breaking changes and the installation pin for remaining on the 0.x line. ChangesMCP compatibility release
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to This release updates version metadata, pins the compatible MCP SDK range, and documents the future migration without changing runtime behavior. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@coderabbitai review |
✅ Action performedReview finished.
|
`resources/read` failed for every real MCP client with:
'AnyUrl' object has no attribute 'startswith'
The SDK types the low-level handler as `Callable[[AnyUrl], ...]` and passes
`req.params.uri` straight through, so `uri` arrives as a pydantic AnyUrl. Our
handler called `.startswith()` on it. Confirmed identical dispatch in mcp 1.28.0
and 1.30.0, so this affected the whole supported range, and 0.5.0 before it.
Every existing test in test_resources.py calls read_resource() directly with a
Python string, bypassing the protocol layer, which is why a green suite never
caught it. Adds a regression test that drives a real ClientSession; it fails with
the exact AttributeError when the coercion is removed.
Found by smoke-testing the server over stdio while verifying the 0.9.0 release.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MsM3KEauXym5MWGjbRiqD9
Adds the design for hybrid keyword + semantic search on search_nvidia: reciprocal rank fusion (k=10) replacing the fixed 70/30 score blend, a local bge-small embedding model behind an optional [embeddings] extra, an evidence floor ahead of the min_relevance_score cutoff, and observable fallback through the existing warnings channel. Also adds the decision record on why page fetches are not cached, updates the 0.9.0 CHANGELOG and README for the new extra, and ignores .claude/worktrees/. Design only. Implementation follows on this branch; PR #8 stays draft until it lands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Adds the task-by-task implementation plan for hybrid search. Corrects the design, decision record and CHANGELOG: search_all_domains already deduplicated results, so duplicates never reached clients. The real fix is moving that existing step ahead of scoring. Also clarifies that semantic_rank returns similarity values, which the evidence floor needs, and how the floor treats queries with no keywords. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Revises the hybrid search design and plan: BM25 over title and snippet replaces both the keyword heuristic and TF-IDF as the single lexical signal, so word overlap no longer gets two votes in fusion. It uses Lucene's non-negative IDF (Okapi IDF goes negative for core query terms on query-biased candidates), stems terms, weights titles twice, keeps the domain boost, and counts expansion-only terms at half weight. RRF k moves to a provisional 30 for two signals. The plan's code was dry-run on a scratch copy: all 7 search.py and server.py edits apply exactly once, the suite reaches the expected 89 passed / 13 skipped, and ruff 0.8.4 (the pre-commit pin) reports no lint or format findings. Updates the CHANGELOG, README and decision record to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Pure functions for reciprocal rank fusion (provisional k=30 for two signals), rank-derived relevance scores and the evidence floor, with unit tests including the k=30 vs k=60 crossover. Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
BM25 over title and snippet with Porter stemming, titles weighted twice, Lucene's non-negative IDF (Okapi IDF goes negative for core query terms on query-biased candidate sets) and expansion-only terms at half weight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Lazily loads a fastembed model exactly once per process under a lock, remembers load failures until restart, encodes in a worker thread with a 10s timeout, and reports why semantic ranking is unavailable (not_installed, model_load_failed, encode_failed). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Replaces the keyword heuristic, TF-IDF and their fixed 70/30 blend with a single BM25 score fused with semantic similarity (when installed) by reciprocal rank fusion, gated by an evidence floor. relevance_score is now derived from fused position. The existing dedupe moves ahead of scoring. Semantic ranking embeds the user's original query. An installed-but-failing embedding model adds a SEMANTIC_RANKING_UNAVAILABLE warning. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Extends test_bm25_alone_ranks_the_stronger_match_first to also assert that SEMANTIC_RANKING_UNAVAILABLE is never emitted for the not_installed case, closing a gap where a regression that warned on not_installed would have passed the suite. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Adds the [embeddings] extra (fastembed), bakes the bge-small weights into the Docker image at /opt/models, and adds a CI job running the ranking tests against the real model on Python 3.10 and 3.12. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
What
Release 0.9.0:
resources/read, which failed for every MCP client;search_nvidia.It also gives notice that 1.0.0 is a breaking release.
1. Pin the MCP SDK ✅
mcp>=1.1.0has no upper bound. A freshpip install mcp-nvidiatoday resolves tomcp2.2.0, which removed the v1 decorator handler API this code is built on, so the install fails at import. Resolving the old requirement gives:The pin becomes
mcp>=1.28,<2.0.0.2. Fix
resources/read✅A stdio smoke test turned up:
The SDK passes
params.urias a pydanticAnyUrl, and the handler called.startswith()on it. The bug is present inmcp1.28.0 and 1.30.0, and in 0.5.0 before them. The test suite never caught it because every resource test called the handler directly with a Python string, bypassing the protocol layer. The fix coerces the URI to a string. A new regression test drives a realClientSession, and it fails with that exact error when the coercion is removed.3. Hybrid search — design done, implementation in progress 🚧
docs/superpowers/specs/2026-09-12-hybrid-search-design.mddocs/decisions/2026-09-12-no-page-fetch-cache.md— why there is no page-fetch cache and no vector index, with the measurements behind thatWhat it does:
cudaappears in 34 of 43 results). Terms that come only from query expansion count at half weight (measured IDF:trt1.83 vstensorrt0.47).k = 30for now: with two signals, a result that one signal ranks first still beats one both rank 15th up to k≈35. The textbook 60 was tuned on ~1000-document lists. Bothkand τ are measured before release.BAAI/bge-small-en-v1.5: onnxruntime, no torch, no API key. Opt-in through the[embeddings]extra. bge-small was chosen over MiniLM-L6 because fastembed's MiniLM pads every input to 128 tokens, while bge-small pads only to the longest input in the batch.SEMANTIC_RANKING_UNAVAILABLEwarning is added.ngc.nvidia.comandcatalog.ngc.nvidia.com) was counted twice.Behaviour changes: result order changes for every
search_nvidiaquery, with or without the extra.relevance_scoreis still an integer from 0–100 but is now derived from fused rank. Typo-tolerant fuzzy matching, the phrase bonus and URL term matching no longer affect search ranking.discover_nvidia_contentis unchanged.Notice: 1.0.0 is breaking
The CHANGELOG and README state plainly that 1.0.0 is not backward compatible with 0.x:
mcp>=2.2.0,<3and the2026-07-28protocol;/sseis removed and returns410 Gone; clients use/mcpwith"transport": "http";-32602instead of-32002;RAILWAY_PUBLIC_DOMAINorMCP_ALLOWED_HOSTS.Production logs show live clients still using
/sse, so this notice is not a formality.Testing so far
Against
mcp1.30.0, the latest 1.x:pytest tests/→ 48 passed, 11 skipped (47 before, plus the regression test)initialize→2025-11-25, both tools and 8 resources listed,resources/readreturns 5261 bytes/healthreports0.9.0Tests for hybrid search land with its implementation.
Follow-on
1.0.0, a pure protocol and SDK migration with no ranking changes, is planned on the
mcp2branch.🤖 Generated with Claude Code
https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6