Skip to content

[bug] web_fetch silently returns link-less Markdown when Readability.js is unavailable (Windows / npm-less environments) #5973

Description

@sjr666666

Before you start

  • I searched existing issues and this is not a duplicate.
  • I can reproduce this on the latest main (e2feca1).

Problem summary

On hosts where Node.js is installed but npm is only resolvable as npm.cmd (stock Windows installs), readabilipy's have_npm() — which calls subprocess.run(["npm", ...]) — raises FileNotFoundError ([WinError 2]), because CreateProcess does not resolve .cmd shims. The lazy node_modules installation inside the readabilipy wheel therefore never runs, have_node() returns False on every call, and every extraction silently degrades to readabilipy's pure-Python tree.

That fallback strips element attributes, which erases the href/src destinations that _resolve_html_urls (introduced for #5307) has just resolved into the markup. The model-visible Markdown returned by web_fetch ends up with link text but no destinations.

Evidence

  • Upstream root cause, already reported: Cannot find NPM in path on Windows alan-turing-institute/ReadabiliPy#115 ("Cannot find NPM in path on Windows", open since 2024-12).

  • On Windows 11 (Python 3.12.13, uv sync --locked), 22 tests in backend/tests/test_web_fetch_relative_links.py fail on main (e2feca1), e.g.:

    FAILED tests/test_web_fetch_relative_links.py::test_extract_article_resolves_link_destinations[../next-https://example.com/next]
    FAILED tests/test_web_fetch_relative_links.py::test_document_base_skips_target_only_base
    ... (22 total)
    
  • Minimal repro of the dependency behavior (no DeerFlow code involved):

    from readabilipy import simple_json_from_html_string
    html = ('<html><head><title>Guide</title></head><body><article><p>'
            + 'This article explains the documentation in detail. ' * 10
            + '</p><p><a href="../next">Next</a></p></article></body></html>')
    art = simple_json_from_html_string(html, use_readability=True)
    print(art["content"][-120:])
    # Warning: A working NPM installation was not found. ...
    # Warning: node executable not found, reverting to pure-Python mode. ...
    # ...</p><p>Next</p></article></body></div>   <- <a href="../next"> flattened to plain text

    shutil.which("npm") finds npm.CMD, subprocess.run(["npm", "version"]) raises FileNotFoundError, while subprocess.run(["npm.cmd", "version"]) succeeds.

Impact

  • Windows users and npm-less deployments (slim containers) get web_fetch output whose links are all flattened to plain text — silently, with only a print to stderr from the dependency.
  • Every extraction re-runs the availability probe, re-attempting an npm install that can never succeed in these environments (per-fetch subprocess overhead + repeated warning spam).

Proposed direction

Route unavailable-or-failing Node extraction to a link-preserving Python fallback inside deerflow/utils/readability.py (the destinations are already absolute after _resolve_html_urls; they only need to survive extraction), and cache the availability probe per process. PR incoming.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageAwaiting maintainer triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions