Repository navigation
docs: an implementation plan for the URL-first pipeline - #18
Merged
Merged
Conversation
Written against the code rather than against the idea. The short version is that the new work sits entirely in front of the existing chain — once a URL has produced a PersonCandidate, resolve, research, score, recommend, draft and approve all already run — and that the blocker is not the crawler but the absence of a job queue. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
docs/url-first-pipeline.md— the plan for taking OutreachGraph from "paste a GitHub handle" to "paste a company URL, or 100 of them".Written against the code rather than against the idea, so every claim in it points at a file and a line.
The two findings that shape it
The blocker is not the crawler, it is the absence of a job queue.
JOB_KINDS(packages/pipeline/src/jobs.ts:19) is a return-type enum with no table behind it; the worker loop (apps/server/src/index.ts:249) runs onlyexpireSignalsandprocessDeletion; andPOST /prospectsexecutes the whole pipeline synchronously inside the request. At 100 URLs that is not a slow request, it is an impossible one.All the new work sits in front of the existing chain. Once a URL has produced a
PersonCandidate, resolve → research → score → recommend → draft → approve already run. That is what makes this a tractable project rather than a rewrite.Two supporting facts worth knowing before anyone argues about scope:
website/observeandwebsite/refresh_researchare alreadyresearch_onlyin the capability matrix, reason "Permitted public web retrieval" (capability-matrix.ts:175). Homepage crawling is a sanctioned capability with no adapter behind it, not a new policy category.enrichWithWaterfallis built and tested but unused — the API passesproviders: [](app.ts:365). Wiring providers in is configuration, not construction.Shape
Four phases: job queue → site provider → company-first endpoints → network fan-out. Phase 2 is the high-risk one and needs a decision before code, because extraction quality is the product: deterministic parsing is cheap and brittle, an LLM pass is robust and costed per URL, and a hybrid is likely.
Docs only — no code changes.
🤖 Generated with Claude Code