Skip to content

[Bug] DWRF selective reader fails nested BIGINT-to-VARCHAR and VARCHAR-to-BOOLEAN schema evolution #941

Description

@fzhedu

Component Selection

  • Core Engine (Expression eval, Memory, Vector)
  • Connectors / File Formats (Hive, Parquet, etc.)
  • API / Bindings (Python, etc.)
  • Build
  • Other

Describe the Bug

The DWRF selective reader does not support nested schema evolution from BIGINT to VARCHAR or from VARCHAR to BOOLEAN. A query that reads an older DWRF/ORC physical schema through a newer requested schema can therefore fail native validation or crash while decoding the nested field.

Sparse struct projections expose an additional issue: SelectiveStructColumnReader uses ScanSpec::channel() as a requested-schema child index. The channel represents the output position, so it can differ from the schema position when only a subset of fields is projected. This can select the wrong child reader and lead to an invalid stream access.

Reproduction Steps

  1. Write DWRF data with this physical schema:

    ROW<chat_id:BIGINT,meta_details:ROW<chatter_id:BIGINT,is_manual_set_nickname:VARCHAR>>
    
  2. Read it with this requested schema:

    ROW<chat_id:BIGINT,meta_details:ROW<chatter_id:VARCHAR,is_manual_set_nickname:BOOLEAN>>
    
  3. Apply an IS NOT NULL filter to the nested struct and project only meta_details, so its output channel is 0 while its requested-schema index is 1.

  4. Read both direct and dictionary-encoded string stripes.

Before the fix, schema compatibility rejects the conversions or the selective reader can select the wrong child and fail during native decoding. A regression test is included in the linked PR.

Bolt Version / Commit ID

main at c7bd647cec78c9b6a646b4565d17c811433068c7

System Configuration

  • OS: Linux
  • Compiler: repository CI toolchain
  • Build Type: Release, Spark-compatible
  • CPU Arch: x86_64
  • Framework: Apache Spark 3.2.1 with Gluten

Logs / Stack Trace

Native validation failed: schema mismatch while reading nested DWRF fields

After allowing the requested conversions, a sparse projection could fail in native decoding because the output channel was used as the requested-schema child index.

Expected Behavior

The DWRF reader should:

  • convert BIGINT values to their VARCHAR representation;
  • parse Spark-compatible boolean strings when reading VARCHAR as BOOLEAN, returning null for invalid values;
  • map sparse projected struct children by their requested-schema field position, independently of their output channel; and
  • support direct and dictionary string encodings without changing the physical reader type.

Additional context

The proposed fix was validated with focused DWRF unit tests and an end-to-end Spark run that exercised the full scan, aggregation, and join path. The Spark application completed all 27 jobs successfully with no schema mismatch, native validation error, decoder crash, or sparse-projection assertion.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions