Private-memory lowering (LowerGEPForPrivMem) miscompiles single-index GEPs whose source element type is an i8 array ([N x i8]) instead of bare i8. The generic aggregate path multiplies the index by the array size, producing an N-times over-stride. Wrong results are computed silently; nothing is diagnosed.
This is newly important because LLVM 23's InstCombine canonicalizes every single-index GEP over a non-i8 type into exactly that shape, so any kernel built by an LLVM 23+ SPIR-V toolchain that indexes a private array can hit it.
Root cause
In IGC/Compiler/CISACodeGen/LowerGEPForPrivMem.cpp:
TransposeHelper::handleGEPInst (~line 1164): the byte-offset fast path is gated on getSourceElementType()->isIntegerTy(8). A GEP with source element type [4 x i8] is an ArrayType, not bare i8, so it misses the fast path and falls into the generic aggregate path, which computes pScalarizedIdx = (0 + %i) * arr_sz with arr_sz = 4. HandleAllocaSources then consumes that index in units of the alloca base type (i32, 4 bytes), giving a 16-byte stride where 4 was intended.
The bail-out guards key on bare i8 as well, so [N x i8] slips past all of them:
SOALayoutChecker::MismatchDetected: if (gepSrcTy->isIntegerTy(8)) { canUseSOALayout = false; }.
- The dynamic non-promoted-type-GEP loop in the same function strips arrays via
extractArrayOrVecEleType, but only walks direct users of the alloca; the affected GEP hangs off an OpBitcast-derived pointer.
- The earlier hardening commits
2c42b1b7580d ("Skip SOA promotion for variable i8 GEPs", 2025-09-24) and b7581328b599 ("Detect type mismatch in SOA promotion when alloca with multi-byte elements is accessed through i8 GEPs", 2025-12-22) share the same blind spot.
Commit 402dd07778f6 (2026-08-10, "Fix byte-offset GEP indexing into vector SoA base types") does not fix this: it is inside the isIntegerTy(8) branch and never fires for [N x i8]. Checked against b5d21888aef4.
Affected versions
| Device |
IGC |
Compute runtime |
Result |
| Arc A380 (DG2) |
intel-igc-core-2 2.36.3 |
intel-opencl-icd 26.22.38646.4-0 |
wrong |
| UHD 770 (ADL-S) |
intel-igc-core-2 2.38.2 (latest release, 2026-07-14) |
intel-opencl-icd 26.27.39122.11-0 |
wrong |
Reproducer
Kernel _Z12sneaky_snakePKjS0_Piii (HeCBench snake-cuda sneaky_snake), which indexes private uint[8] arrays. With identical input sequences and error threshold 0 the correct output is 1; the miscompiled build writes 0.
Two SPIR-V modules differ only in the shape of five access chains.
Failing (array-of-uchar source element type):
%_arr_uchar_uint_4 = OpTypeArray %uchar %uint_4
%p = OpBitcast %_ptr_Function__arr_uchar_uint_4 %RefsPerThread
%q = OpInBoundsPtrAccessChain %_ptr_Function__arr_uchar_uint_4 %p %idx
Working (bare uchar, index pre-scaled to bytes):
%p = OpBitcast %_ptr_Function_uchar %RefsPerThread
%byteidx = OpShiftLeftLogical %uint %idx %uint_2
%q = OpInBoundsPtrAccessChain %_ptr_Function_uchar %p %byteidx
The two compute the same address (index 3 -> base+12 in both).
Evidence this is an optimizer bug, not invalid input
spirv-val --target-env spv1.3 (SPIRV-Tools v2025.3) exits 0 on both modules.
- The failing module computes the correct result on Intel CPU OpenCL (2024.18.12.0.05_160000).
- The failing module computes the correct result on the same GPU with
-cl-opt-disable.
- The shape comes from LLVM itself:
InstCombinerImpl::visitGetElementPtrInst in llvm/lib/Transforms/InstCombine/InstructionCombining.cpp, the block commented "Canonicalize gep %T to gep [sizeof(%T) x i8]". Both the in-tree SPIR-V backend and SPIRV-LLVM-Translator render the canonicalized IR into the access-chain form above.
Build/run of the two modules, with ocl_run.c a plain clCreateProgramWithIL + single-enqueue harness:
cc -o ocl_run ocl_run.c -lOpenCL
./ocl_run bad_arrayof_uchar.spv # writes 0 (wrong)
./ocl_run good_uchar_byteoffset.spv # writes 1 (correct)
./ocl_run bad_arrayof_uchar.spv -cl-opt-disable # writes 1 (correct)
Full .spv and .ll reproducers plus the GPU and CPU harnesses are available on request; GitHub issues do not accept file attachments through the CLI.
For context, one downstream project currently works around this by rewriting the affected GEPs back to the bare-i8 byte-offset form before handing SPIR-V to IGC.
Private-memory lowering (
LowerGEPForPrivMem) miscompiles single-index GEPs whose source element type is an i8 array ([N x i8]) instead of barei8. The generic aggregate path multiplies the index by the array size, producing an N-times over-stride. Wrong results are computed silently; nothing is diagnosed.This is newly important because LLVM 23's InstCombine canonicalizes every single-index GEP over a non-i8 type into exactly that shape, so any kernel built by an LLVM 23+ SPIR-V toolchain that indexes a private array can hit it.
Root cause
In
IGC/Compiler/CISACodeGen/LowerGEPForPrivMem.cpp:TransposeHelper::handleGEPInst(~line 1164): the byte-offset fast path is gated ongetSourceElementType()->isIntegerTy(8). A GEP with source element type[4 x i8]is anArrayType, not barei8, so it misses the fast path and falls into the generic aggregate path, which computespScalarizedIdx = (0 + %i) * arr_szwitharr_sz = 4.HandleAllocaSourcesthen consumes that index in units of the alloca base type (i32, 4 bytes), giving a 16-byte stride where 4 was intended.The bail-out guards key on bare
i8as well, so[N x i8]slips past all of them:SOALayoutChecker::MismatchDetected:if (gepSrcTy->isIntegerTy(8)) { canUseSOALayout = false; }.extractArrayOrVecEleType, but only walks direct users of the alloca; the affected GEP hangs off anOpBitcast-derived pointer.2c42b1b7580d("Skip SOA promotion for variable i8 GEPs", 2025-09-24) andb7581328b599("Detect type mismatch in SOA promotion when alloca with multi-byte elements is accessed through i8 GEPs", 2025-12-22) share the same blind spot.Commit
402dd07778f6(2026-08-10, "Fix byte-offset GEP indexing into vector SoA base types") does not fix this: it is inside theisIntegerTy(8)branch and never fires for[N x i8]. Checked againstb5d21888aef4.Affected versions
Reproducer
Kernel
_Z12sneaky_snakePKjS0_Piii(HeCBench snake-cudasneaky_snake), which indexes privateuint[8]arrays. With identical input sequences and error threshold 0 the correct output is 1; the miscompiled build writes 0.Two SPIR-V modules differ only in the shape of five access chains.
Failing (array-of-uchar source element type):
Working (bare uchar, index pre-scaled to bytes):
The two compute the same address (index 3 -> base+12 in both).
Evidence this is an optimizer bug, not invalid input
spirv-val --target-env spv1.3(SPIRV-Tools v2025.3) exits 0 on both modules.-cl-opt-disable.InstCombinerImpl::visitGetElementPtrInstinllvm/lib/Transforms/InstCombine/InstructionCombining.cpp, the block commented "Canonicalize gep %T to gep [sizeof(%T) x i8]". Both the in-tree SPIR-V backend and SPIRV-LLVM-Translator render the canonicalized IR into the access-chain form above.Build/run of the two modules, with
ocl_run.ca plainclCreateProgramWithIL+ single-enqueue harness:Full
.spvand.llreproducers plus the GPU and CPU harnesses are available on request; GitHub issues do not accept file attachments through the CLI.For context, one downstream project currently works around this by rewriting the affected GEPs back to the bare-
i8byte-offset form before handing SPIR-V to IGC.