11# PCD Pipeline
22
33Build-time pipeline that fetches PCD repositories, compiles their SQL operations, and outputs
4- structured JSON for the website to consume.
4+ structured JSON for the website to consume: the current state of every entity, plus the change
5+ history of every entity derived from the same replay.
56
67## Source
78
89```
910tooling/pcd/
10- ├── index.ts # Entry point, orchestrates fetch -> compile -> extract
11+ ├── index.ts # Entry point, orchestrates fetch -> replay -> extract
1112├── config.json # Database registry
12- ├── fetch.ts # GitHub tarball download and extraction
13- ├── build.ts # In-memory SQLite compilation
13+ ├── fetch.ts # git clone and op file to commit mapping
14+ ├── build.ts # In-memory SQLite creation and op execution
15+ ├── ops.ts # Op file parser (batch header, per-op markers)
16+ ├── diff.ts # Structural diff between two extracted entities
17+ ├── history.ts # History replay fold (pure)
18+ ├── replay.ts # SQLite adapter for the history replay
1419├── extract.ts # Entity extraction via SQL queries
1520└── types.ts # Pipeline-internal types
1621```
@@ -46,28 +51,37 @@ Each entry has:
4651## Pipeline Flow
4752
4853```
49- pnpm compile:pcd
54+ pnpm compile:pcd [-- --no-history]
5055 1. Read config.json
5156 2. For each database:
52- a. Fetch tarball from GitHub API
53- b. Read pcd.json manifest from extracted files
57+ a. Clone the repo at its branch (blobless clone, full commit history)
58+ b. Read pcd.json manifest from the checkout
5459 c. Resolve schema version from manifest dependencies
55- d. Fetch schema tarball (cached if same version as previous database)
60+ d. Clone the schema at that version tag (cached if same version as previous database)
5661 e. Create in-memory SQLite with foreign keys enabled
5762 f. Execute schema ops in numeric filename order
58- g. Execute base ops in numeric filename order
59- h. Extract all entity data via SQL queries
60- i. Write {id}.json to src/lib/data/pcd/
63+ g. Map each base op file to the commit that added it (one git log)
64+ h. Replay base ops one file at a time, recording per-entity history
65+ i. Extract all entity data via SQL queries and assert it matches the replay
66+ j. Write {id}.json and history/{id}.json to src/lib/data/pcd/
6167 3. Write index.json (nav-only data for sidebar)
6268 4. Clean up temp directories
6369```
6470
71+ With ` --no-history ` , step g and the per-file bookkeeping in h are skipped: base ops execute in one
72+ pass and only ` {id}.json ` is written. Pages and artifacts then show no History section.
73+
6574## Fetching
6675
67- Repos are fetched as tarballs via the GitHub API
68- (` https://api.github.com/repos/{owner}/{repo}/tarball/{ref} ` ). No git required at build time. Schema
69- tarballs are cached within a pipeline run since multiple databases typically pin the same schema
70- version.
76+ Repos are cloned with ` git clone --filter=blob:none --single-branch --branch {ref} ` from
77+ ` https://github.com/{owner}/{repo}.git ` . A blobless clone downloads the full commit history but only
78+ the checked-out tree's file contents, so one clone serves both the op files and the commit lookup.
79+ Git is required at build time. Schema clones are cached within a pipeline run since multiple
80+ databases typically pin the same schema version.
81+
82+ Commit metadata comes from one ` git log --name-only --diff-filter=A -- ops ` per repo, mapping each
83+ op file to the commit that added it (hash, author date, subject). An op file with no matching commit
84+ falls back to its ` @exportedAt ` header and no commit link.
7185
7286## Schema Resolution
7387
@@ -87,6 +101,41 @@ numeric filename prefix (`0.schema.sql` before `1.languages.sql` before `10.some
87101
88102No custom SQLite functions are needed. Exported PCD ops use plain SQL with name-based WHERE clauses.
89103
104+ ## History
105+
106+ A database repo's ` ops/ ` folder is an append-only log. The first file is a bulk import with no
107+ markers. Every later file is one Profilarr export batch (in practice one commit) with a header
108+ (` -- @name: ` , ` -- @exportedAt: ` , ` -- @opIds: ` ) and each op wrapped in markers naming the entity it
109+ touches:
110+
111+ ``` sql
112+ -- --- BEGIN op 3587 ( update regular_expression "Special Edition" )
113+ update " regular_expressions" set " pattern" = ' ...' where " name" = ' Special Edition' ;
114+ -- --- END op 3587
115+ ```
116+
117+ ` ops.ts ` parses a file into its header, ops (verb, entity type, name, SQL) and, per op, the other
118+ same-type names the SQL mentions (the old name in a rename's WHERE clause). Test entities
119+ (` test_entity ` , ` test_release ` ) are ignored: the site does not extract them.
120+
121+ ` history.ts ` replays the files in order. After each file it re-reads only the entities the markers
122+ touched, using the per-type extractors in ` extract.ts ` with a name filter, and diffs each against
123+ its previous state (` diff.ts ` , which matches array items by name so a changed condition reads as one
124+ change). The kind of each entry comes from the state transition, not the marker verb: absent then
125+ present is ` created ` , present then absent is ` deleted ` , both present is ` updated ` . An entity that
126+ appeared while exactly one name its ops mention disappeared is a ` renamed ` entry, and the old name's
127+ history moves under the new name; chains through temporary names inside one file resolve to the
128+ original. When a regex or custom format disappears, the custom formats or profiles that referenced
129+ it are re-read too, so cascades are attributed to the file that caused them. A file with no markers
130+ (the bulk import) or an unlabeled op is diffed in full instead, with no related links.
131+
132+ After the last file, a full extraction must deep-equal the replayed state. A mismatch fails the
133+ build naming the differing entities. This is what guarantees a page's History section can never
134+ disagree with the entity it sits under. Entities that no longer exist are dropped from the output,
135+ and related links only point at entities that still exist.
136+
137+ Per database the replay costs about one second on top of the normal compile.
138+
90139## Extraction
91140
92141After compilation, the pipeline queries the SQLite database for each entity type with appropriate
@@ -105,13 +154,21 @@ database. These are consumed by `+page.server.ts` load functions for entity deta
105154` +layout.server.ts ` to populate the sidebar. Kept separate to avoid shipping full entity data to
106155every page.
107156
157+ ** Per-database history** (` history/{id}.json ` ): ` EntityHistory ` from ` src/lib/types/pcd.ts ` , keyed
158+ by ` {entityType}:{name} ` (entity types as they appear in op markers, e.g. ` custom_format ` ,
159+ ` radarr_naming ` ), each a list of ` HistoryEntry ` in replay order. Lives in its own folder so the
160+ routes' ` pcd/*.json ` database globs never see it. Loaded by
161+ ` src/lib/shared/utils/pcd/history-data.ts ` , which tolerates the folder being absent.
162+
108163## Shared Types
109164
110165` src/lib/types/pcd.ts ` defines the compiled data shape, used by both the pipeline and the SvelteKit
111166app. ` CompiledDatabase ` retains both the database manifest version and its pinned schema dependency
112167version. Key interfaces:
113168
114- - ` CompiledDatabase ` - top-level container with metadata and all entity collections
169+ - ` CompiledDatabase ` - top-level container with metadata (including the source ` repo ` and ` branch ` ,
170+ used for commit links) and all entity collections
171+ - ` EntityHistory ` , ` HistoryEntry ` , ` EntityChange ` - the per-entity change log
115172- ` CustomFormat ` - name, description, tags, conditions (with discriminated union for condition data)
116173- ` QualityProfile ` - name, scoring, quality list with groups, languages
117174- ` RegularExpression ` - name, pattern, description, tags
0 commit comments