Repository navigation
[Jetson][Sandbox] Sandbox GPU passthrough proof fails for the non-root sandbox user on JetPack 6.2 IGX Orin — onboarding aborts #7610
Description
Activity
- addedarea: onboardingOnboarding FSM, provider setup, sandbox launch, or first-run flowOnboarding FSM, provider setup, sandbox launch, or first-run flowarea: sandboxOpenShell sandbox lifecycle, runtime, config, or recoveryOpenShell sandbox lifecycle, runtime, config, or recoveryplatform: jetsonAffects Jetson AGX Thor or OrinAffects Jetson AGX Thor or Orin
on Jul 28, 2026 ✨ Thanks for the detailed report. The reproduction steps, environment, and isolation findings are clear. Maintainers will investigate the non-root sandbox user GPU device access on JetPack 6.2 IGX Orin.
- added a commit that references this issue
on Jul 28, 2026 Retested on v0.0.97 — still reproduces (render-group fix insufficient)
Retested on 2026-07-30 on an NVIDIA IGX Orin Developer Kit (aarch64, JetPack 6.2 / L4T R36.5.1) with NemoClaw v0.0.97 (OpenShell 0.0.85). The bug still reproduces — the render-device-group fix is not sufficient on IGX Orin JP6.2.
The fix is applied — onboarding now grants the sandbox user the render group:
Granting sandbox user the detected Jetson GPU device groups via --group-add 44, 109, 995 (so CUDA can initialize as a non-root user)(44 =
video, 109 =render, 995 =debug)But the end-to-end sandbox GPU proof still fails for the non-root sandbox user with the same error as the original report:
Sandbox CUDA proof failed: cuInit(0) via libcuda.so.1 NvRmMemInitNvmap failed with Permission denied 356: Memory Manager Not supported ****NvRmMemMgrInit failed**** error type: 196626 cuInit(0)=999 GPU proof failed inside an executable sandboxOnboarding still aborts;
NEMOCLAW_SANDBOX_GPU=0is still required to complete onboarding.Root cause of the residual failure
The Tegra GPU nodes under
/dev/nvgpu/igpu0/*are ownedroot:video (44)androot:debug (995), which the group-add already covers. ButcuInitstill fails atNvRmMemInitNvmap— the Tegra nvmap memory-manager path. Adding the render / video / debug groups does not grant the non-root sandbox user whatever accessNvRmMemInitNvmaprequires on IGX Orin JP6.2.The fix likely needs to also cover the nvmap access path for the non-root sandbox user, not just the render device group.
71 remaining items
Completion status (2026-09-15)
The product direction is now accepted in #8910: OpenShell native CDI is the single authority for Jetson device injection, mounts, supplemental groups, and hardware-derived policy; NemoClaw will consume and verify the released contract rather than ship the legacy compatibility recreation. Supported hardware remains AGX Thor and IGX Orin.
The remaining path is upstream-gated:
- feat(core): add CDI policy resolver OpenShell#2775 is still open at
6b7e889. Branch checks and Linux ARM/AMD GPU E2E passed, but its required core E2E is red on three apparently unrelated failures (OIDC SQLite lock, rust-docker timeout, and WSL GPU create timeout). I requested a maintainer rerun at feat(core): add CDI policy resolver OpenShell#2775 (this account cannot invoke the admin-only rerun). - fix(gpu): derive CUDA-required Jetson sysfs policy OpenShell#3348 now tracks the minimum CUDA-required Jetson sysfs policy without broad
/sysaccess. - fix(policy): support legitimate CDI policies above 256 paths OpenShell#3349 now tracks the legitimate 301-path Thor CDI policy exceeding the current 256-path limit. Both await upstream ownership/design acceptance.
Once those changes land in a released OpenShell build, #8910 must be reduced to native-CDI adoption and then run the accepted hardware matrix on the exact supported artifacts: official ARM64 image, onboarding exit 0, non-root
nvidia-smi,/proc/<pid>/task/<tid>/commwrite,cuInit(0)=0, restart/resume/rebuild, non-GPU negative, and invalid/missing CDI fail-closed behavior on both AGX Thor and IGX Orin. The earlier Thor GPU-stage success is directional evidence only; it did not complete onboarding and does not close this issue.- feat(core): add CDI policy resolver OpenShell#2775 is still open at
- added and removed
on Sep 16, 2026 QA question rather than a reopen request — I may be missing context that justifies the close, and I would rather ask than flip the state on you.
I attempted to re-verify this on NemoClaw v0.0.131 (IGX Orin, JetPack 6.2 / L4T R36.5.1, aarch64) and could not complete the end-to-end run: onboarding aborts before the sandbox GPU proof with an unrelated dashboard-port/forward condition on that host, so the GPU gate was never exercised. I am making no claim about whether the symptom still occurs.
Three things I noticed while preparing that run, which is why I am asking:
-
This issue has no closing PR. The 2026-10-06 close carries no linked pull request. The earlier close (2026-07-28, PR fix(onboard): add Jetson render device group #7762) was reopened two days later, so an ancestry check against a release tag reports fix(onboard): add Jetson render device group #7762 as "contained" and looks green — which is misleading for anyone using release containment as a gate. That is what prompted me to look at the history rather than trust the check.
-
Two of the three upstream gates you listed are still open. OpenShell refactor(cli): migrate status and tunnel commands to oclif #2775 closed on 2026-10-05, the day before this issue was closed, so I assume that is what unblocked it. But OpenShell feat(cli): add resource profiles to nemoclaw onboard and nemoclaw resources command #3348 (minimum CUDA-required Jetson sysfs policy) and docs(reference): document remote-deploy env vars #3349 (legitimate 301-path CDI policy above the 256 limit) are both still open. Your 2026-09-15 comment listed all three as gating.
-
Which OpenShell build carries the fix? NemoClaw v0.0.131 ships
openshell 0.0.116. If the CDI policy resolver from OpenShell refactor(cli): migrate status and tunnel commands to oclif #2775 is what resolves this, knowing the first OpenShell build that contains it would let QA target the right NemoClaw release instead of guessing.
The question: was the acceptance matrix from your 2026-09-15 comment actually run — official ARM64 image, onboarding exit 0, non-root
nvidia-smi,/proc/<pid>/task/<tid>/commwrite,cuInit(0)=0, restart/resume/rebuild, non-GPU negative, and fail-closed on invalid/missing CDI — on IGX Orin? Or did tracking move upstream, with this issue closed as "no longer tracked here"? Either answer is fine; QA just needs to know which, because the two imply different next steps on our side.One data point from the host, since it is relevant to the boundary you isolated. On this IGX Orin under v0.0.131, direct Docker with the same image and a non-root user (
uid 998,--group-add 44 --group-add 995) returns:cuInit(0) = 0That is the "direct Docker succeeds" half of the comparison you described. I was not able to produce the OpenShell-exec half, because onboarding does not get that far on this host.
Happy to run the full acceptance matrix on IGX Orin once I can get onboarding past the forward-ownership blocker — just say the word and I will report either way.
-
Answering my own question above, empirically — I got the end-to-end run to complete on IGX Orin, and the symptom still reproduces on v0.0.131.
✓ GPU proof passed: nvidia-smi when available ✓ GPU proof passed: /proc/<pid>/task/<tid>/comm write ⚠ GPU proof inconclusive: cuInit(0) via libcuda.so.1 NvRmMemInitNvmap failed with Permission denied 356: Memory Manager Not supported ****NvRmMemMgrInit failed**** error type: 196626 cuInit(0)=999 ⚠ Sandbox CUDA proof failed: cuInit(0) via libcuda.so.1 GPU proof failed inside an executable sandbox.Same two passing proofs, same signature as the original report.
Two details that matter for where this sits:
- It was not the compatibility path. The run used
NEMOCLAW_DOCKER_GPU_PATCH=0, and the log confirmsDirect sandbox GPU enabled; allowing OpenShell GPU policy enrichment.So this is the native-injection direction from feat(onboard): preserve Jetson GPU device groups #8910, against the OpenShell build v0.0.131 ships. - The boundary you isolated still holds. On the same host, same image, same unprivileged identity, directly under Docker:
cuInit(0) = 0. Only the sandbox execution path fails.
I filed this as #12806 rather than asking you to reopen here, since this issue has been closed twice and a comment on a closed thread is easy to miss. If you would rather track it on this issue, say so and I will close #12806 as a duplicate — no attachment to either arrangement.
Earlier in this thread I had only the paper trail and explicitly made no claim about the symptom. That gap is now closed.
- It was not the compatibility path. The run used
Description
On JetPack 6.2 (L4T R36.5.1) IGX Orin, GPU detection succeeds but the end-to-end sandbox GPU passthrough proof fails for the unprivileged in-sandbox user, so
nemoclaw onboardaborts at the GPU proof gate. The user can only complete onboarding by forcing CPU behavior withNEMOCLAW_SANDBOX_GPU=0.Isolation shows the passthrough plumbing is correct and the gap is the non-root user: running
cuInit(0)as root inside the same container returns 0 (success), while the unprivileged sandbox user (even with the--group-addvideo+995 the onboarder applies) cannot initialize CUDA.Platform scope: Reproduced on JetPack 6.2 (R36.5.1) IGX Orin only; other Jetson variants / JetPack versions not tested.
cuInitas root succeeds, so this may be IGX-Orin-specific device-access strictness.Regression: Unknown — earlier versions not tested.
Environment
Steps to Reproduce
curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash[6/8] Creating sandbox→ the GPU proof step.Expected Result
The sandbox GPU proof passes and onboarding completes with sandbox GPU enabled.
Actual Result
Detection / preflight pass:
But the CUDA proof fails, aborting onboarding:
Isolation (device nodes ARE mounted via CSV:
/dev/nvmap,/dev/nvgpu/igpu0/*,/dev/nvidia0,/dev/nvidiactl,/dev/nvsciipc,/dev/nvhost-*; libcuda present):Groups alone (video=44 + 995) are insufficient for the non-root user to init CUDA on this IGX Orin.
Workaround:
NEMOCLAW_SANDBOX_GPU=0(CPU) lets onboarding complete.Logs
Related