-
Notifications
You must be signed in to change notification settings - Fork 296
OpenSep 23, 2026
No due date
•Last updated 9% complete
List view
0 of 70 selected 0 issues of 70 selected
SDPA: CP-accurate numerics mode (chunk_size param, bitwise match to TE CP merge)
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-backendcuDNN backend API, graph execution, descriptors, engines, or backend integration.cuDNN backend API, graph execution, descriptors, engines, or backend integration.Status: Open.#752 In NVIDIA/cudnn-frontend;- Status: Open (in progress).
- Status: Open (in progress).
- Status: Open (in progress).
sdpa fwd: ragged Q over paged KV rides the d128 decode tile (nvbug 6607857)
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).[DSv4.1] Fuse BF16 dGLU and checkpoint activation recomputation
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1187#1187 In NVIDIA/cudnn-frontend;Rule 8 (gated_attention_block): quant words in a workspace slot, per-device graph handle with locked re-stream
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1186#1186 In NVIDIA/cudnn-frontend;Rule 8 (HSTU): caller-owned workspace for block-sparse metadata and the bwd fp32 accumulator, R5 declines, compile-time fakes
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1185#1185 In NVIDIA/cudnn-frontend;Rule 8 (flex_attention): compile() from fake tensors, no device memory at compile
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1184#1184 In NVIDIA/cudnn-frontend;Rule 8 (NSA/CSA): plan-time host ints, SWA caller-owned workspace and THD ragged offsets; wrappers unchanged
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1183#1183 In NVIDIA/cudnn-frontend;Rule 8 (DSA): plan-style classes own no device memory -- required outputs, caller-owned workspaces, plan-time host ints; wrappers unchanged
op: DSADSA relatedDSA relatedorig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1182#1182 In NVIDIA/cudnn-frontend;Rule 8 core: caller-owned workspaces and no host blocking in the graph API, FROST MoE GEMM and SDPA, grouped GEMM and Hopper KDA; R9 detector and recipes R10/R11
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1181#1181 In NVIDIA/cudnn-frontend;BSA: add native SM120 blk128 Sage FP8 forward
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 1163#1163 In NVIDIA/cudnn-frontend;Frost sdpa fwd d64
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).fix: acquire SM100 grouped-wgrad tensor maps before TMA use
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).- Status: Open (in progress).
frost(sdpa_bwd_sm80): serve THD ports at their own head stride
orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Draft (not ready).NVIDIA/cudnn-frontendnumber 1108#1108 In NVIDIA/cudnn-frontend;frost(sdpa_bwd_sm80): read the bound ragged offsets on device (issue #737)
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Draft (not ready).NVIDIA/cudnn-frontendnumber 1058#1058 In NVIDIA/cudnn-frontend;Propagate shared memory limit during engine selection
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).Add experimental JAX KDA using Frost launch hosts
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).Adding Sub-channel Scaling Grouped GEMMs
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).Fix D192 BF16 short-sequence performance regression
cat-perf-bugPerformance regressions or cases where behavior is correct but too slow.Performance regressions or cases where behavior is correct but too slow.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).ops: add standalone NVFP4 block-scale conversion
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 814#814 In NVIDIA/cudnn-frontend;gemm(cutedsl): add standalone Lightning W4A16 grouped GEMM
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open (in progress).NVIDIA/cudnn-frontendnumber 811#811 In NVIDIA/cudnn-frontend;[Bug]: cuDNN FP8 SDPA fails on Windows because cuDNN Frontend looks for Linux CUDA runtime libraries
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-infraInfrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.Infrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.orig-externalReported or requested by an external user, customer, or community contributor.Reported or requested by an external user, customer, or community contributor.Status: Open.#809 In NVIDIA/cudnn-frontend;