Component Selection
Describe the Bug
The DWRF selective reader does not support nested schema evolution from BIGINT to VARCHAR or from VARCHAR to BOOLEAN. A query that reads an older DWRF/ORC physical schema through a newer requested schema can therefore fail native validation or crash while decoding the nested field.
Sparse struct projections expose an additional issue: SelectiveStructColumnReader uses ScanSpec::channel() as a requested-schema child index. The channel represents the output position, so it can differ from the schema position when only a subset of fields is projected. This can select the wrong child reader and lead to an invalid stream access.
Reproduction Steps
-
Write DWRF data with this physical schema:
ROW<chat_id:BIGINT,meta_details:ROW<chatter_id:BIGINT,is_manual_set_nickname:VARCHAR>>
-
Read it with this requested schema:
ROW<chat_id:BIGINT,meta_details:ROW<chatter_id:VARCHAR,is_manual_set_nickname:BOOLEAN>>
-
Apply an IS NOT NULL filter to the nested struct and project only meta_details, so its output channel is 0 while its requested-schema index is 1.
-
Read both direct and dictionary-encoded string stripes.
Before the fix, schema compatibility rejects the conversions or the selective reader can select the wrong child and fail during native decoding. A regression test is included in the linked PR.
Bolt Version / Commit ID
main at c7bd647cec78c9b6a646b4565d17c811433068c7
System Configuration
- OS: Linux
- Compiler: repository CI toolchain
- Build Type: Release, Spark-compatible
- CPU Arch: x86_64
- Framework: Apache Spark 3.2.1 with Gluten
Logs / Stack Trace
Native validation failed: schema mismatch while reading nested DWRF fields
After allowing the requested conversions, a sparse projection could fail in native decoding because the output channel was used as the requested-schema child index.
Expected Behavior
The DWRF reader should:
- convert
BIGINT values to their VARCHAR representation;
- parse Spark-compatible boolean strings when reading
VARCHAR as BOOLEAN, returning null for invalid values;
- map sparse projected struct children by their requested-schema field position, independently of their output channel; and
- support direct and dictionary string encodings without changing the physical reader type.
Additional context
The proposed fix was validated with focused DWRF unit tests and an end-to-end Spark run that exercised the full scan, aggregation, and join path. The Spark application completed all 27 jobs successfully with no schema mismatch, native validation error, decoder crash, or sparse-projection assertion.
Component Selection
Describe the Bug
The DWRF selective reader does not support nested schema evolution from
BIGINTtoVARCHARor fromVARCHARtoBOOLEAN. A query that reads an older DWRF/ORC physical schema through a newer requested schema can therefore fail native validation or crash while decoding the nested field.Sparse struct projections expose an additional issue:
SelectiveStructColumnReaderusesScanSpec::channel()as a requested-schema child index. The channel represents the output position, so it can differ from the schema position when only a subset of fields is projected. This can select the wrong child reader and lead to an invalid stream access.Reproduction Steps
Write DWRF data with this physical schema:
Read it with this requested schema:
Apply an
IS NOT NULLfilter to the nested struct and project onlymeta_details, so its output channel is0while its requested-schema index is1.Read both direct and dictionary-encoded string stripes.
Before the fix, schema compatibility rejects the conversions or the selective reader can select the wrong child and fail during native decoding. A regression test is included in the linked PR.
Bolt Version / Commit ID
mainatc7bd647cec78c9b6a646b4565d17c811433068c7System Configuration
Logs / Stack Trace
Expected Behavior
The DWRF reader should:
BIGINTvalues to theirVARCHARrepresentation;VARCHARasBOOLEAN, returning null for invalid values;Additional context
The proposed fix was validated with focused DWRF unit tests and an end-to-end Spark run that exercised the full scan, aggregation, and join path. The Spark application completed all 27 jobs successfully with no schema mismatch, native validation error, decoder crash, or sparse-projection assertion.