You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
With the Redis cache enabled, OpenMetadata can report an entity as missing while it still exists, and can cancel the workflows of an entity that was never deleted. There are two separate causes.
Background
The "not found" marker. When a lookup finds no entity, OpenMetadata stores a short-lived marker in Redis (30 seconds by default, notFoundTtlSeconds), so repeated lookups of the same missing entity answer 404 without querying the database. A hard delete also writes it, so the deleted entity is not served from a stale cache (fix(cache): seal hard-delete stale-read window via NotFoundCache marker #28197).
Uncached reads. Code calls find or findByName with fromCache=false when it needs the current state from the database, not a cached copy. Governance workflows read everything this way (through FreshReadScope), because their answers decide whether an entity needs approval.
Problem 1: a delete inside a larger transaction finalizes too early
A hard delete ends with two steps that are only correct once the delete is permanent: cancel the entity's workflow instances, and write the "not found" marker. EntityRepository.cleanup() and the bulk subtree delete run them right after writing their own changes. When the delete is part of a larger operation that runs in one database transaction, nothing is committed yet at that point. This happens today, for example when an App's cleanup deletes its ingestion pipelines, or a Table's cleanup deletes leftover test cases.
If the operation rolls back, the entity is still there but answers 404 for 30 seconds, and its workflows stay cancelled, because the workflow engine (Flowable) commits the cancellation on its own.
If it is retried after a deadlock, the steps run again on every attempt.
Problem 2: uncached reads trust the marker
find and findByName check the marker before reading the database, even when the caller asked for the database. For these reads the marker can only be stale or redundant: the database is the source of truth, and the marker is just a remembered earlier answer. A stale marker, such as one left by problem 1, makes them report an existing entity as missing, and the caller acts on it (a governance gate, for example). The admin cache-recovery endpoint in SystemResource already relies on the right behaviour: it expects fromCache=false to find the entity "even when NotFoundCache mistakenly says it doesn't".
The check also costs a Redis request on every such read and almost never saves a query. These reads never write a marker themselves, because the database layer throws on a missing row first, so a marker is only there if other code wrote one in the last 30 seconds.
How to reproduce
With Redis enabled (-Pcache-tests), in code running in the server:
// Problem 1: the schema is still in the database afterwards, but GET returns 404schemas.executeInTransaction(() -> {
schemas.delete("admin", schemaId, true, true);
thrownewIllegalStateException("the larger operation failed");
});
// Problem 2: throws EntityNotFoundException for a table that existsCacheBundle.getNotFoundCache().markNotFoundById(Entity.TABLE, liveTableId);
tables.find(liveTableId, Include.NON_DELETED, false);
Expected
Workflow cancellation and the marker happen only after the outermost transaction commits, and never if it rolls back.
Uncached reads are answered by the database. Cached reads keep using the marker, which is what it was built for: repeated lookups of names that do not exist.
Summary
With the Redis cache enabled, OpenMetadata can report an entity as missing while it still exists, and can cancel the workflows of an entity that was never deleted. There are two separate causes.
Background
notFoundTtlSeconds), so repeated lookups of the same missing entity answer 404 without querying the database. A hard delete also writes it, so the deleted entity is not served from a stale cache (fix(cache): seal hard-delete stale-read window via NotFoundCache marker #28197).findorfindByNamewithfromCache=falsewhen it needs the current state from the database, not a cached copy. Governance workflows read everything this way (throughFreshReadScope), because their answers decide whether an entity needs approval.Problem 1: a delete inside a larger transaction finalizes too early
A hard delete ends with two steps that are only correct once the delete is permanent: cancel the entity's workflow instances, and write the "not found" marker.
EntityRepository.cleanup()and the bulk subtree delete run them right after writing their own changes. When the delete is part of a larger operation that runs in one database transaction, nothing is committed yet at that point. This happens today, for example when an App's cleanup deletes its ingestion pipelines, or a Table's cleanup deletes leftover test cases.Problem 2: uncached reads trust the marker
findandfindByNamecheck the marker before reading the database, even when the caller asked for the database. For these reads the marker can only be stale or redundant: the database is the source of truth, and the marker is just a remembered earlier answer. A stale marker, such as one left by problem 1, makes them report an existing entity as missing, and the caller acts on it (a governance gate, for example). The admin cache-recovery endpoint inSystemResourcealready relies on the right behaviour: it expectsfromCache=falseto find the entity "even when NotFoundCache mistakenly says it doesn't".The check also costs a Redis request on every such read and almost never saves a query. These reads never write a marker themselves, because the database layer throws on a missing row first, so a marker is only there if other code wrote one in the last 30 seconds.
How to reproduce
With Redis enabled (
-Pcache-tests), in code running in the server:Expected