Skip to content

Latest commit

 

History

History
238 lines (184 loc) · 15.1 KB

File metadata and controls

238 lines (184 loc) · 15.1 KB

System Architecture & Cognitive Foundations

This document elucidates the architectural blueprints and theoretical underpinnings of the Perplexity History Export tool. It is designed for those who seek to understand the mechanics of our "Mightiest RAG" implementation.



1. High-Level Flow Diagram

The following diagram illustrates the lifecycle of data, from the initial extraction from Perplexity.ai to the interactive synthesis in the REPL.

%%{init: {
  'theme': 'base',
  'themeVariables': {
    'primaryColor': '#2d1b4d',
    'primaryTextColor': '#e9d5ff',
    'primaryBorderColor': '#7c3aed',
    'lineColor': '#a78bfa',
    'secondaryColor': '#1e1b4b',
    'tertiaryColor': '#4c1d95',
    'mainBkg': '#0f172a',
    'nodeBorder': '#8b5cf6',
    'clusterBkg': '#1e1b4b',
    'clusterBorder': '#7c3aed',
    'titleColor': '#ddd6fe',
    'edgeLabelBackground':'#1e1b4b'
  }
}}%%
graph TD
    User([User Query]) --> Planner[Research Planner]

    subgraph "Scraping & Storage"
        P[Perplexity.ai] -- Playwright --> MD[Markdown Exports]
        MD --> VS[Vector Store - Vectra]
    end

    subgraph "Retrieval Engine"
        Planner -- Semantic Queries --> VS
        Planner -- HyDE Passage --> VS
        Planner -- Hard Keywords --> RG[Ripgrep Search]
        VS --> Fusion[RRF Fusion Ranking]
        RG --> Fusion
        Fusion --> Reranker[Cross-Encoder Reranker]
    end

    subgraph "Synthesis Layer"
        Reranker --> MR[MapReduce Fact Extraction]
        MR --> Narrator[Final Synthesis - Narrator]
        Narrator --> Verify[Quality Verification]
    end

    Verify --> Final[/Mightiest AI Response/]
    Final --> User
Loading

2. RAG Cognitive Structure

Our RAG implementation is not a simple "retrieve and stuff" pipeline. It follows a multi-stage cognitive process inspired by modern IR (Information Retrieval) and LLM orchestration patterns.

Stage A: Adaptive Planning

Before any retrieval, the system acts as a Research Planner. It decomposes the user's query into:

  • Strategy: Selecting between precise (targeted facts, pool of 50) or exhaustive (broad historical overview, pool of 80).
  • Semantic Variations: Generating multiple search phrases to cover different linguistic facets of the query.
  • Hard Keywords: Identifying unique entities or technical IDs that require exact-match precision.
  • HyDE Passage: Generating a short hypothetical answer passage (1-2 sentences) that would plausibly appear in stored history for this question. This passage is embedded and searched alongside the query variations, improving recall when the user's question wording diverges from how the content was originally written.

Stage B: Hybrid Retrieval & Fusion

We employ a Hybrid Search strategy, combining the strengths of dense and sparse retrieval:

  • Dense (Vector): Captures semantic intent and conceptual similarity using embeddings (nomic-embed-text). Runs once per semantic query variation and once for the HyDE passage.
  • Sparse (Exact): Leverages ripgrep for high-velocity exact string matching, ensuring technical IDs or specific names are never missed.

Results from all pools are merged using Reciprocal Rank Fusion (RRF), which provides a robust ranking by combining the ordinal positions of items from different search pools without needing normalized scores.

Stage B¹: Cross-Encoder Reranking

After RRF fusion, the top candidates are passed to a cross-encoder reranker before fact extraction. Unlike the bi-encoder embeddings used for retrieval (which score query and passage independently), a cross-encoder processes the query and each passage jointly, enabling it to reason about fine-grained relevance:

  • Model: Xenova/ms-marco-MiniLM-L-6-v2. It runs locally via ONNX (@huggingface/transformers), no API key required.
  • Quantization: Loaded in int8 for a 3× latency reduction with <0.5% rank-correlation loss versus fp32.
  • Batching: Processed in chunks of 64 pairs for optimal throughput.
  • Threshold & Fallback: Results with logit >= -5.0 are kept; if none pass, the top 20 by RRF score are used as a fallback, guaranteeing the pipeline always has candidates. Logs Reranked N -> M (threshold -5.0, fallback to top-20).
  • Fallback: Gracefully skips reranking if the package is not installed, preserving the RRF order.
  • Implementation note: The raw relevance logit is read directly from AutoModelForSequenceClassification output. The high-level pipeline() API is intentionally bypassed as it normalizes single-class regression heads to a constant score: 1.0.

Stage C: Granular MapReduce Fact Extraction

To mitigate "lost in the middle" phenomena and context window saturation, we utilize a MapReduce approach:

  1. Map: Each snippet is analyzed in small, high-density batches (size 10) to extract atomic facts, code snippets, and dates. Extraction uses:
    • Zod schema validation with node_id coercion (accepts string "3" → 3).
    • Per-entry JSON error handling: malformed LLM output is skipped, batch continues, falls back to raw snippet.
    • Diversity instruction: at most 1 fact per unique source title per batch.
  2. Filter: Post-extraction LLM gate marks facts as relevant if loosely related to Python/programming/learning. Keeps all if 0 pass; keeps top 4 if only 1 passes from >3.
  3. Deduplicate: Source-level dedup keeps best fact per unique source title.
  4. Reduce: Verified facts synthesized into final response with cited sources. History Sources Explored shows [Find N] title + preview. Citations use source title brackets [which big python projects...].

3. Theoretical Foundations

Our architecture is informed by several key papers and concepts in the field of AI and Information Retrieval:

Hybrid Search & RRF

Hypothetical Document Embeddings (HyDE)

HyDE improves retrieval by generating a hypothetical answer to a query before embedding it, closing the lexical gap between questions and historical documents. Our implementation is highly configurable via HYDE_MODE:

  • Off: Standard semantic search only.
  • Fusion: (Traditional HyDE) Always generate a passage and fuse results using RRF.
  • Supplement: (Default) Perform initial semantic search first. Only trigger HyDE if results are weak (score < HYDE_THRESHOLD_SCORE or count < HYDE_THRESHOLD_COUNT). This maximizes accuracy while minimizing LLM latency.

Reference: Gao et al., 2022. Precise Zero-Shot Dense Retrieval without Relevance Labels.

Cross-Encoder Reranking

  • Two-Stage Retrieval: The standard industry pattern of using a fast bi-encoder for candidate retrieval followed by a slower but more precise cross-encoder for reranking. The cross-encoder reads the full (query, passage) pair jointly, capturing interaction signals the bi-encoder cannot.

Retrieval-Augmented Generation (RAG)

MapReduce for Context Compression


4. Visualizing the Retrieval Loop

%%{init: {
  'theme': 'base',
  'themeVariables': {
    'primaryColor': '#2d1b4d',
    'primaryTextColor': '#e9d5ff',
    'primaryBorderColor': '#7c3aed',
    'lineColor': '#a78bfa',
    'secondaryColor': '#1e1b4b',
    'tertiaryColor': '#4c1d95',
    'mainBkg': '#0f172a',
    'nodeBorder': '#8b5cf6'
  }
}}%%
sequenceDiagram
    participant U as User
    participant P as Planner
    participant R as Retrieval (Hybrid)
    participant CE as Cross-Encoder
    participant MR as MapReduce (Extraction)
    participant N as Narrator (Synthesis)

    U->>P: Query
    P->>P: Generate HyDE passage
    P->>R: [Semantic Queries + HyDE Passage + Hard Keywords]
    R->>R: Fusion Ranking (RRF)
    R->>CE: Top N candidates (35 precise / 60 exhaustive)
    CE->>CE: Joint (query, passage) scoring
    CE->>MR: Reranked results
    loop Fact Extraction
        MR->>MR: Batch Processing
    end
    MR->>N: Verified Fact List
    N->>U: Final Synthesized Answer
Loading

5. Glossary

Plain-language definitions for every technical term used in this document.

Term Simple explanation
RAG (Retrieval-Augmented Generation) Instead of the AI making things up from memory, it first searches your actual saved data, then uses those results to write the answer. Think: open-book exam vs. closed-book.
Embedding / Vector Turning a piece of text into a list of numbers that captures its meaning. Texts with similar meaning end up with similar numbers, so you can find related content by comparing numbers.
Vector Store (Vectra) A database that stores those number-lists (embeddings) and can quickly find the ones most similar to a query. Like a library that sorts books by vibe instead of title.
Bi-Encoder Two separate "understanding machines": one encodes the query, one encodes each passage; and you compare the resulting numbers. Fast, but the query and passage never "see" each other during encoding, so subtle relevance signals can be missed.
Cross-Encoder One "understanding machine" that reads the query and a passage together at the same time. Much more accurate at judging relevance, but slower. It is used as a second pass after bi-encoder retrieval narrows the field.
Reranking Taking a first batch of search results and sorting them again with a smarter (but slower) method. The first pass casts a wide net; reranking picks the actual best ones.
Dense Retrieval Search by meaning/semantics using embeddings. Good at finding conceptually related content even if the exact words differ.
Sparse Retrieval Search by exact keywords or strings (e.g. ripgrep). Perfect for finding specific names, IDs, or code snippets that embeddings might miss.
Hybrid Search Combining dense (meaning) and sparse (keyword) retrieval so you get the benefits of both.
RRF (Reciprocal Rank Fusion) A formula for merging ranked lists from multiple search methods into one combined ranking. It rewards results that appear near the top in multiple lists, without needing to normalize scores.
HyDE (Hypothetical Document Embeddings) Before searching, ask the LLM: "What would a good answer to this question look like?" Then search using that hypothetical answer instead of the raw question. Helps when your question phrasing doesn't match how the answer was originally written. Configurable via HYDE_MODE (off, fusion, supplement).
MapReduce A two-step processing pattern: Map = process many small chunks independently in parallel; Reduce = combine all those results into one final output. Used here to extract facts from many snippets without overloading the LLM's context window.
Context Window The maximum amount of text an LLM can read and reason about at once. Stuffing too many search results in at once causes the model to "lose" information in the middle; we use MapReduce to mitigate this.
Lost in the Middle A known LLM failure mode: when given a very long input, the model tends to remember the start and end well but forgets details buried in the middle. MapReduce mitigates this by processing small chunks.
ONNX A standard file format for AI models that lets them run efficiently on different hardware and in different programming languages (including Node.js) without needing Python or a GPU.
Quantization (int8) Compressing a model's internal numbers from 32-bit floats down to 8-bit integers. Makes the model ~3-4× faster and smaller with minimal accuracy loss.
LLM (Large Language Model) The AI model that reads text and generates responses. In this project, served locally via Ollama (e.g. mistral, llama3).
Ollama A local server that runs open-source LLMs on your own machine. No API key, no cloud, no data leaving your device.
Playwright A browser automation library used here to scrape conversation history from Perplexity.ai and save it as Markdown files.
ripgrep (rg) An extremely fast command-line search tool that scans files for exact text matches. Used here for the sparse/keyword retrieval leg of hybrid search.
nomic-embed-text The embedding model (run via Ollama) that converts text snippets into vectors for storage and similarity search in Vectra.
Provenance Knowing exactly which source document a fact came from. The system tracks this so the final answer can cite where each piece of information originated.

Content Integrity & Incremental Updates

To achieve high-fidelity synchronization without redundant operations, the system employs a content-driven skipping mechanism:

  • Stable Serialization: Raw API entries are serialized into JSON with keys sorted alphabetically to ensure a deterministic representation.
  • SHA-256 Hashing: A cryptographic hash is generated from the stable JSON.
  • Checkpoint Comparison: On each run (or when manually triggered via "Sync"), the system compares the fresh hash against the stored value in .storage/checkpoint.json.
  • Differential Export: Only threads with mismatched hashes are re-rendered and written to disk, preserving existing files and minimizing filesystem I/O.