Skip to content

docs: an implementation plan for the URL-first pipeline - #18

Merged
ralyodio merged 2 commits into
mainfrom
docs/url-first-plan
Aug 13, 2026
Merged

ralyodio merged 2 commits into
mainfrom
docs/url-first-plan

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Adds docs/url-first-pipeline.md — the plan for taking OutreachGraph from "paste a GitHub handle" to "paste a company URL, or 100 of them".

Written against the code rather than against the idea, so every claim in it points at a file and a line.

The two findings that shape it

The blocker is not the crawler, it is the absence of a job queue. JOB_KINDS (packages/pipeline/src/jobs.ts:19) is a return-type enum with no table behind it; the worker loop (apps/server/src/index.ts:249) runs only expireSignals and processDeletion; and POST /prospects executes the whole pipeline synchronously inside the request. At 100 URLs that is not a slow request, it is an impossible one.

All the new work sits in front of the existing chain. Once a URL has produced a PersonCandidate, resolve → research → score → recommend → draft → approve already run. That is what makes this a tractable project rather than a rewrite.

Two supporting facts worth knowing before anyone argues about scope:

  • website/observe and website/refresh_research are already research_only in the capability matrix, reason "Permitted public web retrieval" (capability-matrix.ts:175). Homepage crawling is a sanctioned capability with no adapter behind it, not a new policy category.
  • enrichWithWaterfall is built and tested but unused — the API passes providers: [] (app.ts:365). Wiring providers in is configuration, not construction.

Shape

Four phases: job queue → site provider → company-first endpoints → network fan-out. Phase 2 is the high-risk one and needs a decision before code, because extraction quality is the product: deterministic parsing is cheap and brittle, an LLM pass is robust and costed per URL, and a hybrid is likely.

Docs only — no code changes.

🤖 Generated with Claude Code

ralyodio and others added 2 commits August 13, 2026 01:37
Written against the code rather than against the idea. The short version
is that the new work sits entirely in front of the existing chain — once
a URL has produced a PersonCandidate, resolve, research, score,
recommend, draft and approve all already run — and that the blocker is
not the crawler but the absence of a job queue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ralyodio
ralyodio merged commit add0f02 into main Aug 13, 2026
4 checks passed
@ralyodio
ralyodio deleted the docs/url-first-plan branch August 16, 2026 17:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant