Skip to content

chore: release 0.9.0 pinning mcp<2.0.0 - #8

Draft
bharatr21 wants to merge 11 commits into
mainfrom
chore/pin-mcp-v1
Draft

bharatr21 wants to merge 11 commits into
mainfrom
chore/pin-mcp-v1

Conversation

@bharatr21

@bharatr21 bharatr21 commented Sep 12, 2026 •

Copy link
Copy Markdown
Owner

Draft on purpose. This release now includes hybrid search. Its design is committed here, and the implementation is in progress on this branch. Railway deploys production from main, so merging this deploys it; it stays draft until the feature lands and is verified.

What

Release 0.9.0:

  1. pins the MCP SDK so existing installs keep working;
  2. fixes resources/read, which failed for every MCP client;
  3. adds hybrid keyword + semantic ranking to search_nvidia.

It also gives notice that 1.0.0 is a breaking release.

1. Pin the MCP SDK ✅

mcp>=1.1.0 has no upper bound. A fresh pip install mcp-nvidia today resolves to mcp 2.2.0, which removed the v1 decorator handler API this code is built on, so the install fails at import. Resolving the old requirement gives:

+ mcp==2.2.0
+ mcp-types==2.2.0
+ starlette==1.6.0

The pin becomes mcp>=1.28,<2.0.0.

2. Fix resources/read ✅

A stdio smoke test turned up:

mcp.shared.exceptions.McpError: 'AnyUrl' object has no attribute 'startswith'

The SDK passes params.uri as a pydantic AnyUrl, and the handler called .startswith() on it. The bug is present in mcp 1.28.0 and 1.30.0, and in 0.5.0 before them. The test suite never caught it because every resource test called the handler directly with a Python string, bypassing the protocol layer. The fix coerces the URI to a string. A new regression test drives a real ClientSession, and it fails with that exact error when the coercion is removed.

3. Hybrid search — design done, implementation in progress 🚧

  • Design: docs/superpowers/specs/2026-09-12-hybrid-search-design.md
  • Decision record: docs/decisions/2026-09-12-no-page-fetch-cache.md — why there is no page-fetch cache and no vector index, with the measurements behind that

What it does:

  • BM25 replaces the keyword heuristic and TF-IDF as the single lexical signal, so word overlap no longer gets two votes against meaning's one. It is stemmed, counts title terms twice and keeps the domain boost. It uses Lucene's non-negative IDF, because on these candidates the textbook Okapi IDF goes negative for core query terms (cuda appears in 34 of 43 results). Terms that come only from query expansion count at half weight (measured IDF: trt 1.83 vs tensorrt 0.47).
  • Reciprocal rank fusion over the BM25 and semantic rankings replaces the fixed 70/30 score blend. k = 30 for now: with two signals, a result that one signal ranks first still beats one both rank 15th up to k≈35. The textbook 60 was tuned on ~1000-document lists. Both k and τ are measured before release.
  • Local embeddings via fastembed with BAAI/bge-small-en-v1.5: onnxruntime, no torch, no API key. Opt-in through the [embeddings] extra. bge-small was chosen over MiniLM-L6 because fastembed's MiniLM pads every input to 128 tokens, while bge-small pads only to the longest input in the batch.
  • The model loads lazily, exactly once, under a lock, and encoding runs in a worker thread so it never blocks the event loop.
  • An evidence floor runs before the cutoff, so a query with only weak matches still returns fewer results.
  • Observable fallback: if the extra is installed but fails, results are ranked on the remaining signals and a SEMANTIC_RANKING_UNAVAILABLE warning is added.
  • Dedupe moves ahead of scoring. The existing dedupe already kept duplicates out of client output, but it ran after scoring, so a page returned by two overlapping subdomains (such as ngc.nvidia.com and catalog.ngc.nvidia.com) was counted twice.
  • Evaluation: a comparison script first, then labeled regression fixtures built from the rankings it surfaces.

Behaviour changes: result order changes for every search_nvidia query, with or without the extra. relevance_score is still an integer from 0–100 but is now derived from fused rank. Typo-tolerant fuzzy matching, the phrase bonus and URL term matching no longer affect search ranking. discover_nvidia_content is unchanged.

Notice: 1.0.0 is breaking

The CHANGELOG and README state plainly that 1.0.0 is not backward compatible with 0.x:

  • it requires mcp>=2.2.0,<3 and the 2026-07-28 protocol;
  • /sse is removed and returns 410 Gone; clients use /mcp with "transport": "http";
  • resource errors return -32602 instead of -32002;
  • self-hosted deployments must allow their hostname through RAILWAY_PUBLIC_DOMAIN or MCP_ALLOWED_HOSTS.

Production logs show live clients still using /sse, so this notice is not a formality.

Testing so far

Against mcp 1.30.0, the latest 1.x:

  • pytest tests/ → 48 passed, 11 skipped (47 before, plus the regression test)
  • Live stdio smoke test: initialize → 2025-11-25, both tools and 8 resources listed, resources/read returns 5261 bytes
  • Live HTTP smoke test: /health reports 0.9.0

Tests for hybrid search land with its implementation.

Follow-on

1.0.0, a pure protocol and SDK migration with no ranking changes, is planned on the mcp2 branch.

🤖 Generated with Claude Code

https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6

The `mcp>=1.1.0` requirement had no upper bound, so a fresh install now
resolves to mcp 2.2.0 - an SDK that removed the v1 decorator handler API
this release is built on. Pin to `mcp>=1.28,<2.0.0`.

Also gives notice, in both the CHANGELOG and the README, that 1.0.0 is a
breaking release: it requires mcp>=2.2.0 and the 2026-07-28 protocol, moves
the remote transport from /sse to /mcp, and changes resource errors from
-32002 to -32602. Users who want to stay on this line pin mcp-nvidia<1.0.0.

Verified against mcp 1.30.0 (latest 1.x): 47 passed, 11 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MsM3KEauXym5MWGjbRiqD9
@coderabbitai

coderabbitai Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: ba176ab2-db3f-4f71-a0f1-4cb1966b1252

📥 Commits

Reviewing files that changed from the base of the PR and between a727196 and e85e93f.

📒 Files selected for processing (4)
  • CHANGELOG.md
  • README.md
  • package.json
  • pyproject.toml

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The project moves to version 0.9.0 and pins the MCP SDK to the supported 1.x range. The changelog and README document the planned 1.0.0 breaking changes and the installation pin for remaining on the 0.x line.

Changes

MCP compatibility release

Layer / File(s) Summary
Version and SDK compatibility pin
package.json, pyproject.toml
The project version changes from 0.5.0 to 0.9.0. The mcp dependency changes to mcp>=1.28,<2.0.0.
Release and upgrade documentation
CHANGELOG.md, README.md
The documentation describes the planned 1.0.0 endpoint, protocol, error-code, host allow-list, and SDK changes. It also documents the mcp-nvidia<1.0.0 pin.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Merge Risk: ⚪ Minimal · up to e85e9

This release updates version metadata, pins the compatible MCP SDK range, and documents the future migration without changing runtime behavior.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the 0.9.0 release and the MCP SDK upper-bound pin, which are the main changes in the pull request.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch chore/pin-mcp-v1

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@bharatr21

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

`resources/read` failed for every real MCP client with:

    'AnyUrl' object has no attribute 'startswith'

The SDK types the low-level handler as `Callable[[AnyUrl], ...]` and passes
`req.params.uri` straight through, so `uri` arrives as a pydantic AnyUrl. Our
handler called `.startswith()` on it. Confirmed identical dispatch in mcp 1.28.0
and 1.30.0, so this affected the whole supported range, and 0.5.0 before it.

Every existing test in test_resources.py calls read_resource() directly with a
Python string, bypassing the protocol layer, which is why a green suite never
caught it. Adds a regression test that drives a real ClientSession; it fails with
the exact AttributeError when the coercion is removed.

Found by smoke-testing the server over stdio while verifying the 0.9.0 release.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MsM3KEauXym5MWGjbRiqD9
@bharatr21
bharatr21 marked this pull request as draft September 12, 2026 22:05
bharatr21 and others added 9 commits September 12, 2026 17:09
Adds the design for hybrid keyword + semantic search on search_nvidia:
reciprocal rank fusion (k=10) replacing the fixed 70/30 score blend, a
local bge-small embedding model behind an optional [embeddings] extra,
an evidence floor ahead of the min_relevance_score cutoff, and
observable fallback through the existing warnings channel.

Also adds the decision record on why page fetches are not cached,
updates the 0.9.0 CHANGELOG and README for the new extra, and ignores
.claude/worktrees/.

Design only. Implementation follows on this branch; PR #8 stays draft
until it lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Adds the task-by-task implementation plan for hybrid search.

Corrects the design, decision record and CHANGELOG: search_all_domains
already deduplicated results, so duplicates never reached clients. The
real fix is moving that existing step ahead of scoring. Also clarifies
that semantic_rank returns similarity values, which the evidence floor
needs, and how the floor treats queries with no keywords.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Revises the hybrid search design and plan: BM25 over title and snippet
replaces both the keyword heuristic and TF-IDF as the single lexical
signal, so word overlap no longer gets two votes in fusion. It uses
Lucene's non-negative IDF (Okapi IDF goes negative for core query terms
on query-biased candidates), stems terms, weights titles twice, keeps
the domain boost, and counts expansion-only terms at half weight. RRF k
moves to a provisional 30 for two signals.

The plan's code was dry-run on a scratch copy: all 7 search.py and
server.py edits apply exactly once, the suite reaches the expected 89
passed / 13 skipped, and ruff 0.8.4 (the pre-commit pin) reports no
lint or format findings.

Updates the CHANGELOG, README and decision record to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Pure functions for reciprocal rank fusion (provisional k=30 for two
signals), rank-derived relevance scores and the evidence floor, with
unit tests including the k=30 vs k=60 crossover.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
BM25 over title and snippet with Porter stemming, titles weighted twice,
Lucene's non-negative IDF (Okapi IDF goes negative for core query terms
on query-biased candidate sets) and expansion-only terms at half weight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Lazily loads a fastembed model exactly once per process under a lock,
remembers load failures until restart, encodes in a worker thread with a
10s timeout, and reports why semantic ranking is unavailable
(not_installed, model_load_failed, encode_failed).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Replaces the keyword heuristic, TF-IDF and their fixed 70/30 blend with a
single BM25 score fused with semantic similarity (when installed) by
reciprocal rank fusion, gated by an evidence floor. relevance_score is
now derived from fused position. The existing dedupe moves ahead of
scoring. Semantic ranking embeds the user's original query. An
installed-but-failing embedding model adds a
SEMANTIC_RANKING_UNAVAILABLE warning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Extends test_bm25_alone_ranks_the_stronger_match_first to also assert
that SEMANTIC_RANKING_UNAVAILABLE is never emitted for the not_installed
case, closing a gap where a regression that warned on not_installed
would have passed the suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6
Adds the [embeddings] extra (fastembed), bakes the bge-small weights into
the Docker image at /opt/models, and adds a CI job running the ranking
tests against the real model on Python 3.10 and 3.12.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013FriD63a2waebg8Qmtscj6

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant