This document elucidates the architectural blueprints and theoretical underpinnings of the Perplexity History Export tool. It is designed for those who seek to understand the mechanics of our "Mightiest RAG" implementation.
- 1. High-Level Flow Diagram
- 2. RAG Cognitive Structure
- 3. Theoretical Foundations
- 4. Visualizing the Retrieval Loop
- 5. Glossary
The following diagram illustrates the lifecycle of data, from the initial extraction from Perplexity.ai to the interactive synthesis in the REPL.
%%{init: {
'theme': 'base',
'themeVariables': {
'primaryColor': '#2d1b4d',
'primaryTextColor': '#e9d5ff',
'primaryBorderColor': '#7c3aed',
'lineColor': '#a78bfa',
'secondaryColor': '#1e1b4b',
'tertiaryColor': '#4c1d95',
'mainBkg': '#0f172a',
'nodeBorder': '#8b5cf6',
'clusterBkg': '#1e1b4b',
'clusterBorder': '#7c3aed',
'titleColor': '#ddd6fe',
'edgeLabelBackground':'#1e1b4b'
}
}}%%
graph TD
User([User Query]) --> Planner[Research Planner]
subgraph "Scraping & Storage"
P[Perplexity.ai] -- Playwright --> MD[Markdown Exports]
MD --> VS[Vector Store - Vectra]
end
subgraph "Retrieval Engine"
Planner -- Semantic Queries --> VS
Planner -- HyDE Passage --> VS
Planner -- Hard Keywords --> RG[Ripgrep Search]
VS --> Fusion[RRF Fusion Ranking]
RG --> Fusion
Fusion --> Reranker[Cross-Encoder Reranker]
end
subgraph "Synthesis Layer"
Reranker --> MR[MapReduce Fact Extraction]
MR --> Narrator[Final Synthesis - Narrator]
Narrator --> Verify[Quality Verification]
end
Verify --> Final[/Mightiest AI Response/]
Final --> User
Our RAG implementation is not a simple "retrieve and stuff" pipeline. It follows a multi-stage cognitive process inspired by modern IR (Information Retrieval) and LLM orchestration patterns.
Before any retrieval, the system acts as a Research Planner. It decomposes the user's query into:
- Strategy: Selecting between
precise(targeted facts, pool of 50) orexhaustive(broad historical overview, pool of 80). - Semantic Variations: Generating multiple search phrases to cover different linguistic facets of the query.
- Hard Keywords: Identifying unique entities or technical IDs that require exact-match precision.
- HyDE Passage: Generating a short hypothetical answer passage (1-2 sentences) that would plausibly appear in stored history for this question. This passage is embedded and searched alongside the query variations, improving recall when the user's question wording diverges from how the content was originally written.
We employ a Hybrid Search strategy, combining the strengths of dense and sparse retrieval:
- Dense (Vector): Captures semantic intent and conceptual similarity using embeddings (
nomic-embed-text). Runs once per semantic query variation and once for the HyDE passage. - Sparse (Exact): Leverages
ripgrepfor high-velocity exact string matching, ensuring technical IDs or specific names are never missed.
Results from all pools are merged using Reciprocal Rank Fusion (RRF), which provides a robust ranking by combining the ordinal positions of items from different search pools without needing normalized scores.
After RRF fusion, the top candidates are passed to a cross-encoder reranker before fact extraction. Unlike the bi-encoder embeddings used for retrieval (which score query and passage independently), a cross-encoder processes the query and each passage jointly, enabling it to reason about fine-grained relevance:
- Model:
Xenova/ms-marco-MiniLM-L-6-v2. It runs locally via ONNX (@huggingface/transformers), no API key required. - Quantization: Loaded in
int8for a 3× latency reduction with <0.5% rank-correlation loss versus fp32. - Batching: Processed in chunks of 64 pairs for optimal throughput.
- Threshold & Fallback: Results with logit >= -5.0 are kept; if none pass, the top 20 by RRF score are used as a fallback, guaranteeing the pipeline always has candidates. Logs
Reranked N -> M (threshold -5.0, fallback to top-20). - Fallback: Gracefully skips reranking if the package is not installed, preserving the RRF order.
- Implementation note: The raw relevance logit is read directly from
AutoModelForSequenceClassificationoutput. The high-levelpipeline()API is intentionally bypassed as it normalizes single-class regression heads to a constantscore: 1.0.
To mitigate "lost in the middle" phenomena and context window saturation, we utilize a MapReduce approach:
- Map: Each snippet is analyzed in small, high-density batches (size 10) to extract atomic facts, code snippets, and dates. Extraction uses:
- Zod schema validation with
node_idcoercion (accepts string"3"→3). - Per-entry JSON error handling: malformed LLM output is skipped, batch continues, falls back to raw snippet.
- Diversity instruction: at most 1 fact per unique source title per batch.
- Zod schema validation with
- Filter: Post-extraction LLM gate marks facts as relevant if loosely related to Python/programming/learning. Keeps all if 0 pass; keeps top 4 if only 1 passes from >3.
- Deduplicate: Source-level dedup keeps best fact per unique source title.
- Reduce: Verified facts synthesized into final response with cited sources.
History Sources Exploredshows[Find N] title + preview. Citations use source title brackets[which big python projects...].
Our architecture is informed by several key papers and concepts in the field of AI and Information Retrieval:
- Reciprocal Rank Fusion (RRF): Based on the principle that combining multiple search orderings can significantly outperform any single ordering.
- Reference: Cormack et al., 2009. Reciprocal Rank Fusion outperforms Condorcet and Individual Rank Learning Methods. (SIGIR 2009, DOI: 10.1145/1571941.1572114)
HyDE improves retrieval by generating a hypothetical answer to a query before embedding it, closing the lexical gap between questions and historical documents. Our implementation is highly configurable via HYDE_MODE:
- Off: Standard semantic search only.
- Fusion: (Traditional HyDE) Always generate a passage and fuse results using RRF.
- Supplement: (Default) Perform initial semantic search first. Only trigger HyDE if results are weak (score <
HYDE_THRESHOLD_SCOREor count <HYDE_THRESHOLD_COUNT). This maximizes accuracy while minimizing LLM latency.
Reference: Gao et al., 2022. Precise Zero-Shot Dense Retrieval without Relevance Labels.
- Two-Stage Retrieval: The standard industry pattern of using a fast bi-encoder for candidate retrieval followed by a slower but more precise cross-encoder for reranking. The cross-encoder reads the full (query, passage) pair jointly, capturing interaction signals the bi-encoder cannot.
- General RAG Framework: We follow the core paradigm of grounding LLM outputs in external, verifiable data.
- Summarization & Synthesis: Our fact extraction layer mirrors the "MapReduce" chain pattern, effectively handling long-context retrieval by distilling information before final generation.
- Reference: Wu et al., 2021. Recursively Summarizing Books with Human Feedback. (Applying hierarchical summarization principles).
%%{init: {
'theme': 'base',
'themeVariables': {
'primaryColor': '#2d1b4d',
'primaryTextColor': '#e9d5ff',
'primaryBorderColor': '#7c3aed',
'lineColor': '#a78bfa',
'secondaryColor': '#1e1b4b',
'tertiaryColor': '#4c1d95',
'mainBkg': '#0f172a',
'nodeBorder': '#8b5cf6'
}
}}%%
sequenceDiagram
participant U as User
participant P as Planner
participant R as Retrieval (Hybrid)
participant CE as Cross-Encoder
participant MR as MapReduce (Extraction)
participant N as Narrator (Synthesis)
U->>P: Query
P->>P: Generate HyDE passage
P->>R: [Semantic Queries + HyDE Passage + Hard Keywords]
R->>R: Fusion Ranking (RRF)
R->>CE: Top N candidates (35 precise / 60 exhaustive)
CE->>CE: Joint (query, passage) scoring
CE->>MR: Reranked results
loop Fact Extraction
MR->>MR: Batch Processing
end
MR->>N: Verified Fact List
N->>U: Final Synthesized Answer
Plain-language definitions for every technical term used in this document.
| Term | Simple explanation |
|---|---|
| RAG (Retrieval-Augmented Generation) | Instead of the AI making things up from memory, it first searches your actual saved data, then uses those results to write the answer. Think: open-book exam vs. closed-book. |
| Embedding / Vector | Turning a piece of text into a list of numbers that captures its meaning. Texts with similar meaning end up with similar numbers, so you can find related content by comparing numbers. |
| Vector Store (Vectra) | A database that stores those number-lists (embeddings) and can quickly find the ones most similar to a query. Like a library that sorts books by vibe instead of title. |
| Bi-Encoder | Two separate "understanding machines": one encodes the query, one encodes each passage; and you compare the resulting numbers. Fast, but the query and passage never "see" each other during encoding, so subtle relevance signals can be missed. |
| Cross-Encoder | One "understanding machine" that reads the query and a passage together at the same time. Much more accurate at judging relevance, but slower. It is used as a second pass after bi-encoder retrieval narrows the field. |
| Reranking | Taking a first batch of search results and sorting them again with a smarter (but slower) method. The first pass casts a wide net; reranking picks the actual best ones. |
| Dense Retrieval | Search by meaning/semantics using embeddings. Good at finding conceptually related content even if the exact words differ. |
| Sparse Retrieval | Search by exact keywords or strings (e.g. ripgrep). Perfect for finding specific names, IDs, or code snippets that embeddings might miss. |
| Hybrid Search | Combining dense (meaning) and sparse (keyword) retrieval so you get the benefits of both. |
| RRF (Reciprocal Rank Fusion) | A formula for merging ranked lists from multiple search methods into one combined ranking. It rewards results that appear near the top in multiple lists, without needing to normalize scores. |
| HyDE (Hypothetical Document Embeddings) | Before searching, ask the LLM: "What would a good answer to this question look like?" Then search using that hypothetical answer instead of the raw question. Helps when your question phrasing doesn't match how the answer was originally written. Configurable via HYDE_MODE (off, fusion, supplement). |
| MapReduce | A two-step processing pattern: Map = process many small chunks independently in parallel; Reduce = combine all those results into one final output. Used here to extract facts from many snippets without overloading the LLM's context window. |
| Context Window | The maximum amount of text an LLM can read and reason about at once. Stuffing too many search results in at once causes the model to "lose" information in the middle; we use MapReduce to mitigate this. |
| Lost in the Middle | A known LLM failure mode: when given a very long input, the model tends to remember the start and end well but forgets details buried in the middle. MapReduce mitigates this by processing small chunks. |
| ONNX | A standard file format for AI models that lets them run efficiently on different hardware and in different programming languages (including Node.js) without needing Python or a GPU. |
| Quantization (int8) | Compressing a model's internal numbers from 32-bit floats down to 8-bit integers. Makes the model ~3-4× faster and smaller with minimal accuracy loss. |
| LLM (Large Language Model) | The AI model that reads text and generates responses. In this project, served locally via Ollama (e.g. mistral, llama3). |
| Ollama | A local server that runs open-source LLMs on your own machine. No API key, no cloud, no data leaving your device. |
| Playwright | A browser automation library used here to scrape conversation history from Perplexity.ai and save it as Markdown files. |
| ripgrep (rg) | An extremely fast command-line search tool that scans files for exact text matches. Used here for the sparse/keyword retrieval leg of hybrid search. |
| nomic-embed-text | The embedding model (run via Ollama) that converts text snippets into vectors for storage and similarity search in Vectra. |
| Provenance | Knowing exactly which source document a fact came from. The system tracks this so the final answer can cite where each piece of information originated. |
To achieve high-fidelity synchronization without redundant operations, the system employs a content-driven skipping mechanism:
- Stable Serialization: Raw API entries are serialized into JSON with keys sorted alphabetically to ensure a deterministic representation.
- SHA-256 Hashing: A cryptographic hash is generated from the stable JSON.
- Checkpoint Comparison: On each run (or when manually triggered via "Sync"), the system compares the fresh hash against the stored value in
.storage/checkpoint.json. - Differential Export: Only threads with mismatched hashes are re-rendered and written to disk, preserving existing files and minimizing filesystem I/O.