Repository navigation
feat(data-access): batched semantic index writer and suggestion lookup (LLMO-7445) - #1966
Conversation
…p (LLMO-7445) Moves the semantic-index helpers to the redesigned tables (match_type / match_field_type) and removes the semantic type registries from utils.
|
This PR will trigger a minor release when merged. |
MysticatBot
left a comment
There was a problem hiding this comment.
Hey @tb-adbe,
⚠ Degraded review - no spec document was found for this change (searched the PR links, the touched repos' docs, the architecture/guidelines docs, and linked Jira). This review covers code-level quality but could not validate the change against an agreed design, so confidence is reduced. Add a spec link (PR template section 4) and re-request review for a full-confidence pass.
Verdict: Request changes - the batched writer is solid and well tested, but the read-side embedQueries skips the input bounds the writer applies.
Complexity: HIGH - large diff (about 1,500 changed lines) across two packages with a breaking API reshape.
Changes: replaces the per-table semantic helpers with one batched indexSemanticTexts writer, an embedQueries read helper and a new lookupSuggestionsByVectors, and removes the semantic type registries from utils (12 files).
Note: Recommend a human read before merge - this change amends docs/adr/0002-embedding-client-and-semantic-index-utils.md (architectural decision record). The bot review is a complement to, not a replacement for, a human read here.
Note: CI checks were still pending at review time.
Must fix before merge
- [Important]
embedQueriesdoes not applyMAX_SOURCE_TEXT_LENGTHto query texts, so long queries re-embed on every call and never cache -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:718(details inline) - [Important]
embedQuerieshas no cap on text count and sends all misses in one unbatched embed call -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:733(details inline)
Non-blocking (10): minor issues and suggestions
- suggestion: the stored-row diff ignores
model/dims, so changingSEMANTIC_MATCHING_CONFIG.embeddingleaves old vectors marked unchanged and search silently stops matching; select and compare them -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:336 - suggestion: unchanged rows never get a corrected
entityType, soentityTypesfilters can miss them; compare storedentity_typeor document that it is immutable -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:614 - suggestion: mixed input like
['ok', <3000 chars>]silently drops the bad text while full-replace still prunes the group; throw or return a rejected count -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:365 - suggestion: the length check runs before
toLowerCase(), which can lengthen some strings; check the normalized key against the limit -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:241 - suggestion: the cache-read catch swallows any error into
result.cacheErrorwith no log; narrow it or accept alogand warn likeembedQueries-packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:632 - suggestion:
timeoutMsis not validated;0, negative orNaNfails every miss asEmbeddingUnavailableError-packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:709 - suggestion:
entityIdis only checked for non-empty; validating it as a UUID fails fast and keeps.in()filters well-formed -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:358 - suggestion: removed exports (
syncOpportunitySemantic,SEMANTIC_SEARCH_RPC, the*_SEMANTIC_*registries) ship as a minor; add aBREAKING CHANGE:footer or an explicit note in the squash commit body -packages/spacecat-shared-utils/src/index.js:92 - suggestion: the semantic exports have no
.d.tsdeclarations (pre-existing, but the surface grew) -packages/spacecat-shared-data-access/src/service/index.js:75 - nit:
indexSemanticTexts/embedQueriespinmodel/dimsfrom config while the lookups still require callers to pass them; pick one convention -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:858
Open design questions (3): not blocking
- A
semantic_query_embeddingread failure fails everyembedQueriescall, while the writer treats the same failure as best-effort; is that intended for the synchronous API path? -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:720 - The global query cache stores normalized query text from every tenant; is SELECT restricted to the writer role, and are the
hits/missescounts kept internal? -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:754 - ADR-0002 still describes adding types to the removed registries above the amendment; should that text carry an inline superseded note? -
docs/adr/0002-embedding-client-and-semantic-index-utils.md:15
Skill: pr-review | Model: us.anthropic.claude-opus-5-5[1m] | Duration: 1m 1s | Cost: $4.50 | Commit: 7f6abf8726f5b79151d7b40082ea9b31baf8c289
If this code review was useful, please react with 👍. Otherwise, react with 👎.
| } | ||
| const { model, dims } = SEMANTIC_MATCHING_CONFIG.embedding; | ||
| const cacheKey = { model, dims }; | ||
| const normalized = texts.map((text) => hashQueryText(text).normalized); |
There was a problem hiding this comment.
issue (blocking): embedQueries only runs hashQueryText, which rejects empty strings but not over-length ones. The writer bounds texts at MAX_SOURCE_TEXT_LENGTH through cleanTopicText, and that JSDoc says both sides "reject the same input".
Failure: a query longer than 2048 characters misses the cache and gets embedded. The cache upsert then breaks the text CHECK that the :71 comment says this constant matches. That error is swallowed as a warn at :745, so every repeat of the query pays for a new embedding and logs again. If the provider rejects the length instead, the caller sees EmbeddingUnavailableError (an outage) for what is really a bad request.
Fix: run each text through cleanTopicText and throw ValidationError on null. Reject rather than truncate, so the stored hash matches what was embedded. Add a test with a 2049-character query.
Blocking basis: round1
There was a problem hiding this comment.
Fixed in d65e289: embedQueries now runs every text through cleanTopicText and throws ValidationError (texts[i] must be a non-empty string of at most 2048 characters) instead of truncating, so the cached hash always matches what was embedded. cleanTopicText now checks the normalized key (what is stored), so the writer and reader share the exact bound. Tests cover a 2049-character query.
| if (misses.length > 0) { | ||
| let embedded; | ||
| try { | ||
| embedded = await embedWithinBudget(embeddingClient, misses, timeoutMs); |
There was a problem hiding this comment.
issue (blocking): texts has no upper bound. All misses go to createEmbeddings in one call, with no dedupe and none of the writer's EMBEDDING_BATCH_SIZE batching.
Failure: a request with more distinct uncached texts than the per-request input cap noted at :68 always fails as EmbeddingUnavailableError. Below that cap, one request can still drive one large paid embed call, and duplicate misses are sent more than once. The timeout at :686 only abandons the race; it does not cancel the underlying request.
Fix: dedupe misses by normalized text and fan the vectors back out to their positions. Then either reject inputs above a MAX_QUERY_TEXTS (at most EMBEDDING_BATCH_SIZE) with a ValidationError before the cache read, or batch within the time budget.
Blocking basis: round1
There was a problem hiding this comment.
Fixed in d65e289: misses are deduped by normalized text and fanned back out to their positions, and more than MAX_QUERY_TEXTS (= EMBEDDING_BATCH_SIZE, 256) distinct texts is rejected with a ValidationError before the cache read, so the misses always fit one embed call. hits / misses now count distinct texts. api-service caps topics at 100, so it's unaffected. Tests cover the dedupe, the cap and exactly MAX_QUERY_TEXTS.
…ns (LLMO-7445) indexSemanticTopics / lookupOpportunitiesByTopic / lookupSuggestionsByTopic set match_type internally; callers pass only matchFieldType.
…eration in lookups (LLMO-7445)
|
Thanks. Both blocking findings are fixed in d65e289 (inline replies). Note the review ran on the first commit; Non-blocking
Design questions
|
MysticatBot
left a comment
There was a problem hiding this comment.
Hey @tb-adbe,
⚠ Degraded review - no spec document was found for this change (searched the PR links, the touched repos' docs, the architecture/guidelines docs, and linked Jira). This review covers code-level quality but could not validate the change against an agreed design, so confidence is reduced. Add a spec link (PR template section 4) and re-request review for a full-confidence pass.
Verdict: Approve - both earlier blocking findings are fixed and tested. What is left is ADR wording, a few test gaps and optional hardening.
Complexity: HIGH - large diff (about 2,800 diff lines) across two packages with a breaking API reshape.
Changes: moves the semantic-index helpers to the redesigned tables with a batched indexSemanticTopics writer, embedQueries, and topic lookups for opportunities and suggestions, and removes the semantic type registries from utils (12 files).
Note: Recommend a human read before merge - this change amends docs/adr/0002-embedding-client-and-semantic-index-utils.md (architectural decision record). The bot review is a complement to, not a replacement for, a human read here.
Note: CI checks were still pending at review time. Merge once the Test job passes the coverage thresholds.
Non-blocking (12): minor issues and suggestions
- nit: this Consequences bullet still names
lookupOpportunitiesByVectorsand caller-passedmodel/dims, with no superseded marker. Rename it tolookupOpportunitiesByTopic/lookupSuggestionsByTopicor point it at the amendment -docs/adr/0002-embedding-client-and-semantic-index-utils.md:71 - nit: these
_Superseded by the amendment below:_markers sit in the middle of sentences and are followed by a double space. Move them to the start of the paragraph, as at lines 43 and 72 -docs/adr/0002-embedding-client-and-semantic-index-utils.md:48 - nit: the Alternatives entry that rejects embedding inside data-access is reversed by the amendment but has no inline marker -
docs/adr/0002-embedding-client-and-semantic-index-utils.md:80 - nit: the error says "at most 2048 characters", but the check runs on the normalized key. Add "once normalized" so the message matches the JSDoc -
packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:746 - suggestion: with no
log, a failed query-cache read inembedQueriesleaves no trace, and the catch also swallows code bugs. ReturncacheErroras the writer does, or narrow the catch toDataAccessError-packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:772 - suggestion: only distinct texts are capped. Many copies of one text cost a single embed call but return one vector each, and
searchByVectorsthen sends one sequential RPC per 20 vectors. Captexts.lengthorvectors.lengthatMAX_QUERY_TEXTS-packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:798 - suggestion:
embedWithinBudgetabandons the race on timeout but does not abort the paidcreateEmbeddingscall, and timed-out misses are never cached, so retries pay again. Pass anAbortSignalif the client supports one -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:702 - suggestion:
MAX_QUERY_TEXTS = EMBEDDING_BATCH_SIZEcouples the read-side API limit to a write-side batching constant. A separate constant with a "must be <= EMBEDDING_BATCH_SIZE" comment keeps the two decisions independent -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:74 - suggestion: the writer now copies vectors from the query cache into persistent index rows, so a bad cache row now persists past cache eviction. Record in the ADR that this relies on the internal-only PostgREST trust boundary -
packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:616 - suggestion: add a test for a stored old-model row whose text is not resubmitted, asserting it is deleted by id -
packages/spacecat-shared-data-access/test/unit/util/semantic-index.utils.test.js:430 - suggestion: add an
embedQueriestest that mixes cache hits with duplicate inputs (for example['Cached', 'cached ', 'fresh']) and assertshits,misses, the fanned-back vectors and one touched hash -packages/spacecat-shared-data-access/test/unit/util/semantic-index.utils.test.js:564 - suggestion: in the exactly-
MAX_QUERY_TEXTStest, also assert that the misses went out in a single embed call -packages/spacecat-shared-data-access/test/unit/util/semantic-index.utils.test.js:578
Open design questions (2): not blocking
- After a model change, rows for entities that are never refreshed (closed opportunities, suggestions under resolved opportunities) stay forever and are invisible to search. Is a one-off sweep (delete where
modelis not current) part of the model-change runbook? -docs/adr/0002-embedding-client-and-semantic-index-utils.md:71 - Should
upsertQueryEmbeddingsstay a public export, now that both readers and the writer trust the global cache it writes to? -packages/spacecat-shared-data-access/src/service/index.js:54
Previously flagged, now resolved
embedQueriesnow runs every text throughcleanTopicTextand rejects empty or over-length normalized keys withValidationError. A test covers a 2049-character query.embedQueriesdedupes misses, rejects more thanMAX_QUERY_TEXTSdistinct texts before the cache read, and maps vectors back to input order. Tests cover the cap and the at-limit-plus-repeat case.- Stored rows from another model or dims are re-embedded and overwritten in place.
timeoutMs,siteIdandentityIdare validated, and a query-cache read failure on the read path is now best-effort.
Skill: pr-review | Model: us.anthropic.claude-opus-5-5[1m] | Duration: 0m 48s | Cost: $4.72 | Commit: d65e28968f42d2c6f79150a1e2a38e8d96e0e747
If this code review was useful, please react with 👍. Otherwise, react with 👎.
…che helpers internal (LLMO-7445)
|
Thanks. Addressed in e960850, all but one:
Design questions:
|
MysticatBot
left a comment
There was a problem hiding this comment.
Hey @tb-adbe,
⚠ Degraded review - no spec document was found for this change (searched the PR links, the touched repos' docs, the architecture/guidelines docs, and linked Jira). This review covers code-level quality but could not validate the change against an agreed design, so confidence is reduced. Add a spec link (PR template section 4) and re-request review for a full-confidence pass.
Verdict: Approve - the new commit addresses the earlier suggestions with targeted tests. What is left is wording, an optional guard test, and light hardening.
Complexity: HIGH - large diff (about 2,800 diff lines) across two packages with a breaking API reshape.
Changes: moves the semantic-index helpers to the redesigned tables with a batched indexSemanticTopics writer, embedQueries, and topic lookups for opportunities and suggestions, and removes the semantic type registries from utils (12 files).
Note: Recommend a human read before merge - this change amends docs/adr/0002-embedding-client-and-semantic-index-utils.md (architectural decision record). The bot review is a complement to, not a replacement for, a human read here.
Note: CI checks were still pending at review time.
Non-blocking (4): minor issues and suggestions
- suggestion: "nothing outside data-access can write vectors that the writer and readers trust" claims more than the code enforces. The package has no
exportsmap, so a deep import ofsrc/util/semantic-index.utils.jsstill reachesupsertQueryEmbeddings. Drop theexportkeyword from the three cache helpers (tests can use a test-only re-export), or soften the line to match the ADR: not part of the public API, trusted on the assumption that onlyembedQuerieswrites the cache -packages/spacecat-shared-data-access/README.md:194 - nit: "logged and returned as
cacheError), like the writer" reads as if the writer also logs.indexSemanticTopicsreturnscacheErrorwithout logging it. Attach "like the writer" to the all-misses behavior only -docs/adr/0002-embedding-client-and-semantic-index-utils.md:67 - suggestion:
MAX_QUERY_TEXTS <= EMBEDDING_BATCH_SIZEis only a comment, and the single-embed-call test stays green if either constant changes. Addexpect(MAX_QUERY_TEXTS).to.be.at.most(EMBEDDING_BATCH_SIZE)-packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:74 - suggestion:
cacheErroris the rawDataAccessError, which carries table names and the PostgREST cause. Add a JSDoc note that callers log it and never return it to clients -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:807
Open design questions (1): not blocking
- Will the old-generation row sweep in the model-change runbook be a script, or only runbook prose? -
docs/adr/0002-embedding-client-and-semantic-index-utils.md:72
Previously flagged, now resolved
MAX_QUERY_TEXTSis its own constant, and the cap now counts every entry, not just distinct ones, checked before any work. A test covers repeated inputs.embedQueriesreturnscacheErroron a cache read failure, and the error message says "once normalized".- The query-cache helpers are no longer exported from
service/index.js, and the ADR records the cache-copy trust boundary and the model-change sweep. - ADR superseded markers now start their paragraphs, and the Consequences and Alternatives entries point at the new lookups and the amendment.
- New tests cover deleting an old-generation row, fanning cached repeats back out, and sending one embed call at the limit.
Skill: pr-review | Model: us.anthropic.claude-opus-5-5[1m] | Duration: 0m 46s | Cost: $2.74 | Commit: e960850ac120b4c38d3b09c20541e7b489be0bcc
If this code review was useful, please react with 👍. Otherwise, react with 👎.
rarescheseli
left a comment
There was a problem hiding this comment.
Hey @tb-adbe,
⚠ Degraded review - no spec document was found for this change (searched the PR links, the touched repos' docs, the architecture/guidelines docs, and linked Jira). This review covers code-level quality but could not validate the change against an agreed design, so confidence is reduced. Add a spec link (PR template section 4) and re-request review for a full-confidence pass.
Verdict: Approve - the writer, embedQueries and both topic lookups are correct and well tested; what is left is type declarations and doc accuracy, none blocking.
Complexity: HIGH - large diff across two packages with a breaking reshape of the semantic API.
Changes: moves the semantic-index helpers to the redesigned tables with a batched indexSemanticTopics writer, embedQueries, and topic lookups for opportunities and suggestions, and removes the semantic type registries from utils (12 files).
Note: Recommend a human read before merge - this change amends docs/adr/0002-embedding-client-and-semantic-index-utils.md (architectural decision record), and the docs claim a cache trust boundary the code does not enforce (heuristic). The bot review is a complement to, not a replacement for, a human read here.
Non-blocking (5): minor issues and suggestions
- suggestion: the root
typesentry (src/index.d.ts) still declares none of the semantic exports, while the README tells consumers to importindexSemanticTopics,embedQueries,lookupOpportunitiesByTopic,lookupSuggestionsByTopic,SEMANTIC_TARGETS,MAX_QUERY_TEXTSandEmbeddingUnavailableErrorfrom the package root. The gap predates this PR, but the PR rewrites the whole surface. Please file the follow-up you offered (semantic + url-index declarations) before api-service adopts this API -packages/spacecat-shared-data-access/src/index.d.ts:13 - suggestion: "nothing outside data-access can write vectors that the writer and readers trust" is stronger than the code.
getQueryEmbeddings/upsertQueryEmbeddings/touchQueryEmbeddingsare stillexported atsemantic-index.utils.js:430,:475,:519, andpackage.jsonhas noexportsmap, so a deep import reaches them. Drop theexportkeywords (tests can import via a test-only re-export) or reword to "not part of the public API" -packages/spacecat-shared-data-access/README.md:194 - nit: "(logged and returned as
cacheError), like the writer" reads as if the writer logs too;indexSemanticTopicsonly returnscacheError. Attach "like the writer" to the all-misses behavior only -docs/adr/0002-embedding-client-and-semantic-index-utils.md:67 - suggestion:
MAX_QUERY_TEXTS <= EMBEDDING_BATCH_SIZEis only a comment, and the one-embed-call test stays green if either constant changes. Addexpect(MAX_QUERY_TEXTS).to.be.at.most(EMBEDDING_BATCH_SIZE)-packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:74 - suggestion:
cacheErroris the rawDataAccessError(table names, PostgREST cause). Note in the JSDoc that callers log it and never return it to clients -packages/spacecat-shared-data-access/src/util/semantic-index.utils.js:729
## [@adobe/spacecat-shared-data-access-v4.39.0](https://github.com/adobe/spacecat-shared/compare/@adobe/spacecat-shared-data-access-v4.38.0...@adobe/spacecat-shared-data-access-v4.39.0) (2026-10-02) ### Features * **data-access:** batched semantic index writer and suggestion lookup (LLMO-7445) ([#1966](#1966)) ([3479136](3479136)), closes [mysticat-data-service#1153](https://github.com/adobe/mysticat-data-service/issues/1153) [mysticat-data-service#1153](https://github.com/adobe/mysticat-data-service/issues/1153) [#1153](#1153) [#1153](#1153) [#1153](#1153)
|
🎉 This PR is included in version @adobe/spacecat-shared-data-access-v4.39.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
## [@adobe/spacecat-shared-utils-v1.128.0](https://github.com/adobe/spacecat-shared/compare/@adobe/spacecat-shared-utils-v1.127.0...@adobe/spacecat-shared-utils-v1.128.0) (2026-10-02) ### Features * **data-access:** batched semantic index writer and suggestion lookup (LLMO-7445) ([#1966](#1966)) ([3479136](3479136)), closes [mysticat-data-service#1153](https://github.com/adobe/mysticat-data-service/issues/1153) [mysticat-data-service#1153](https://github.com/adobe/mysticat-data-service/issues/1153) [#1153](#1153) [#1153](#1153) [#1153](#1153)
|
🎉 This PR is included in version @adobe/spacecat-shared-utils-v1.128.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
…LMO-7445) (#1967) <!-- mysticat-pr-skill --> ## 1. Abstract Renames the data-access export `MAX_SOURCE_TEXT_LENGTH` to `MAX_TEXT_LENGTH`. ## 2. Reasoning The semantic-index redesign (#1966) renamed the `source_text` / `source_hash` columns to `text` / `text_hash`, but this constant kept the old name. It is the last `source`-era identifier in the semantic-index code. api-service is about to import it (spacecat-api-service#79 review: derive the topic length limit from data-access instead of keeping its own 2048), so it's renamed before anything depends on it. ## 3. High-level overview of the changes - `MAX_SOURCE_TEXT_LENGTH` → `MAX_TEXT_LENGTH` (still 2048, matching the `text` DB CHECKs), in the source, the package export, the tests and ADR-0002. - No behaviour change, and no alias: nothing outside data-access imports the old name (checked api-service, audit-worker and project-elmo-ui). - Removed export, for the squash commit body: `MAX_SOURCE_TEXT_LENGTH` (replaced by `MAX_TEXT_LENGTH`). Released as a patch for the same reason as #1966: the only consumers are the LLMO-7445 code paths, updated in lockstep. ## 4. Required information - Jira / issue: [LLMO-7445](https://jira.corp.adobe.com/browse/LLMO-7445) - Spec: [ADR-0002](https://github.com/adobe/spacecat-shared/blob/main/docs/adr/0002-embedding-client-and-semantic-index-utils.md) (amended 2026-10-01) - Other: [spacecat-shared#1966](#1966), [spacecat-api-service#79](https://github.com/Adobe-AEM-Sites/spacecat-api-service/pull/79) ## 5. Spec deviations - None ## 6. Affected / used mysticat-workspace projects - spacecat-api-service: consumer. #79 will import `MAX_TEXT_LENGTH` for its topic length limit once this is released. ## 8. Test plan (a) Rename only; no runtime behaviour changes, so nothing beyond the unit tests applies. (b) None per environment: the value and every code path are unchanged. ## 9. Deployment & merge order - [spacecat-api-service#79](https://github.com/Adobe-AEM-Sites/spacecat-api-service/pull/79): related. After this releases, #79 bumps data-access and replaces its own `MAX_TOPIC_LENGTH = 2048` with the imported `MAX_TEXT_LENGTH`. Order: this PR (release) → bump data-access in spacecat-api-service#79. 🤖 Generated with [review-kit](https://github.com/adobe/experience-success-skills) Co-authored-by: Tanase Butcaru <tbutcaru@adobe.com>
## [@adobe/spacecat-shared-data-access-v4.39.1](https://github.com/adobe/spacecat-shared/compare/@adobe/spacecat-shared-data-access-v4.39.0...@adobe/spacecat-shared-data-access-v4.39.1) (2026-10-02) ### Bug Fixes * **data-access:** rename MAX_SOURCE_TEXT_LENGTH to MAX_TEXT_LENGTH (LLMO-7445) ([#1967](#1967)) ([d20b120](d20b120)), closes [#1966](#1966) [spacecat-api-service#79](https://github.com/adobe/spacecat-api-service/issues/79) [#79](#79) [#79](#79)
1. Abstract
Moves the semantic-index helpers to the redesigned tables: one batched writer and two readers for opportunities and suggestions. The type registries are removed from utils.
2. Reasoning
The first version of the semantic index had separate shapes per table, a
source_typecolumn, and type registries in utils, so every new producer or text kind needed a shared release. Each producer would also have had to embed, batch, and dedupe on its own. The data-service redesign (#1153) gives all three tables the same text columns and splitssource_typeintomatch_typeandmatch_field_type. This PR moves the shared helpers to that shape and puts the embed and write logic in one place.3. High-level overview of the changes
utils
OPPORTUNITY_SEMANTIC_SOURCE_TYPES,OPPORTUNITY_SEMANTIC_ENTITY_TYPES, and the suggestion equivalents, along with their.d.tstypes. This reverts feat(utils): add semantic lookup source/entity type registries (LLMO-7445) #1959.data-access (
semantic-index.utils.js)match_typeinternally, so callers can't index or search under the wrong lookup. Today that's topics; claims follow the same pattern when that dimension lands.SEMANTIC_MATCH_TYPESis internal.indexSemanticTopics(pg, embeddingClient, { target, siteId, entities: [{ entityId, entityType, fields: [{ matchFieldType, texts }] }] })replacessyncOpportunitySemantic.targetisopportunityorsuggestion.matchFieldType) group: it keeps unchanged texts, embeds and inserts new ones, and deletes stale ones. Fields that aren't listed (and other match types) are left alone, andtexts: []clears a field.semantic_query_embeddingbut never writes to it.embedQueries(pg, embeddingClient, texts, { timeoutMs, log })is the query side. It normalizes the texts, serves cached vectors, embeds the distinct misses in one call within a time budget, and returns the cache writes as a best-effortcacheWritespromise.MAX_QUERY_TEXTS(256) texts with aValidationError, and treats a cache read failure as all misses (logged and returned ascacheError).lookupOpportunitiesByTopicand the newlookupSuggestionsByTopicreplacelookupOpportunitiesByVectors. They take optionalmatchFieldTypes/entityTypesfilters (free-form, deduped, max 100 values).lookupSuggestionsByTopicchecksstatusesagainstSuggestion.STATUSESand addsopportunityStatuses.embedQueriesand both lookups readSEMANTIC_MATCHING_CONFIG.embeddingthemselves.siteIdandentityIdmust be UUIDs, and each writer field reports arejectedcount (texts dropped as empty or too long).entityTypeandmatchFieldTypeare free-form, length-bounded strings, so there are no registries left.SEMANTIC_TARGETS,SEMANTIC_SEARCH_RPCS(target → RPC),MAX_QUERY_TEXTSandEmbeddingUnavailableError.getQueryEmbeddings,upsertQueryEmbeddingsandtouchQueryEmbeddingsare now internal, so onlyembedQuerieswrites the cache that the writer and readers trust. The cache row stores the normalizedtext.Release type. This is a breaking change to the semantic API in both packages, released as
feat(minor). No external code consumes these helpers yet: the writer in audit-worker is behindOFFSITE_SEMANTIC_INDEX_ENABLED(off), consumers pin exact versions, and their PRs adopting this API come next. The old helpers also stop working once mysticat-data-service#1153 runs.Removed exports (to list in the squash commit body):
OPPORTUNITY_SEMANTIC_SOURCE_TYPES,OPPORTUNITY_SEMANTIC_ENTITY_TYPES,SUGGESTION_SEMANTIC_SOURCE_TYPES,SUGGESTION_SEMANTIC_ENTITY_TYPESand their.d.tstypes;syncOpportunitySemantic,lookupOpportunitiesByVectors,SEMANTIC_SEARCH_RPC, the re-exportedOPPORTUNITY_SEMANTIC_SOURCE_TYPES/OPPORTUNITY_SEMANTIC_ENTITY_TYPES, and the cache helpersgetQueryEmbeddings/upsertQueryEmbeddings/touchQueryEmbeddings(no consumer imports them); the lookups no longer acceptmodel/dims.4. Required information
docs/adr/0002-embedding-client-and-semantic-index-utils.md(Decision 4 amended; outdated paragraphs marked superseded)6. Affected / used mysticat-workspace projects
p_match_type,p_match_field_types) from feat: implement strict 500 version limit for configurations #1153.indexSemanticTopicswhen it bumps data-access (spacecat-audit-worker#27, draft; the writer is flag-off today).embedQueries,lookupOpportunitiesByTopicandmatchFieldTypes, and dropssourceTypes, when it bumps data-access (spacecat-api-service#79, draft).8. Test plan
(a) Nothing beyond unit tests yet. The helpers are first exercised end to end when the audit-worker and api-service follow-ups run against a local data-service with #1153 applied.
(b) dev: after #1153 is deployed and the consumer PRs ship, enable
OFFSITE_SEMANTIC_INDEX_ENABLEDand run an offsite audit. Check that rows land inopportunity_semantic_embeddingwithmatch_type='topic', then call the by-topics endpoint and confirm it returns matches.9. Deployment & merge order
Order: data-service #1153 → this PR (release) → audit-worker + api-service bumps.
🤖 Generated with review-kit