Skip to content

[Jetson][Sandbox] Sandbox GPU passthrough proof fails for the non-root sandbox user on JetPack 6.2 IGX Orin — onboarding aborts #7610

Description

@wangericnv

Description

On JetPack 6.2 (L4T R36.5.1) IGX Orin, GPU detection succeeds but the end-to-end sandbox GPU passthrough proof fails for the unprivileged in-sandbox user, so nemoclaw onboard aborts at the GPU proof gate. The user can only complete onboarding by forcing CPU behavior with NEMOCLAW_SANDBOX_GPU=0.

Isolation shows the passthrough plumbing is correct and the gap is the non-root user: running cuInit(0) as root inside the same container returns 0 (success), while the unprivileged sandbox user (even with the --group-add video+995 the onboarder applies) cannot initialize CUDA.

Platform scope: Reproduced on JetPack 6.2 (R36.5.1) IGX Orin only; other Jetson variants / JetPack versions not tested. cuInit as root succeeds, so this may be IGX-Orin-specific device-access strictness.

Regression: Unknown — earlier versions not tested.

Environment

Device:        NVIDIA IGX Orin Developer Kit
OS:            Ubuntu 22.04 (JetPack 6.2, L4T R36.5.1)
Architecture:  aarch64
Node.js:       v22.23.1
npm:           10.9.8
Docker:        29.6.2 (installed via official convenience script; nvidia-container-toolkit 1.20.0~rc.1, CSV mode; nvidia runtime configured via `nvidia-ctk runtime configure --runtime=docker`)
OpenShell CLI: 0.0.85
NemoClaw:      v0.0.93
OpenClaw:      2026.7.1

Steps to Reproduce

  1. Fresh install on IGX Orin JetPack 6.2:
    curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash
  2. Configure the NVIDIA container runtime for Docker:
    sudo nvidia-ctk runtime configure --runtime=docker
    sudo systemctl restart docker
  3. Onboard against NVIDIA Endpoints (any cloud model):
    nemoclaw onboard --name jetson-vdr --agent openclaw
  4. Watch [6/8] Creating sandbox → the GPU proof step.

Expected Result

The sandbox GPU proof passes and onboarding completes with sandbox GPU enabled.

Actual Result

Detection / preflight pass:

OK NVIDIA GPU detected (Orin (nvgpu), 54790 MB); Sandbox GPU: enabled (auto)
OK Granting sandbox user access to Jetson Tegra GPU device nodes via --group-add 44, 995
OK Docker container mode: --runtime nvidia (NVIDIA_VISIBLE_DEVICES=all)
OK GPU proof passed: nvidia-smi when available
OK GPU proof passed: /proc/{pid}/task/{tid}/comm write

But the CUDA proof fails, aborting onboarding:

Sandbox CUDA proof failed: cuInit(0) via libcuda.so.1
  NvRmMemInitNvmap failed with Permission denied 356: Memory Manager Not supported
  ****NvRmMemMgrInit failed**** error type: 196626 cuInit(0)=999
GPU proof failed inside an executable sandbox
-> throws at docker-gpu-sandbox-create.js verifyGpuOrExit; onboarding does not finish.

Isolation (device nodes ARE mounted via CSV: /dev/nvmap, /dev/nvgpu/igpu0/*, /dev/nvidia0, /dev/nvidiactl, /dev/nvsciipc, /dev/nvhost-*; libcuda present):

- cuInit(0) as ROOT in the container                 -> 0   (CUDA_SUCCESS)
- cuInit(0) as sandbox user (uid 998, +grp 44,995)   -> 801 (CUDA_ERROR_NOT_SUPPORTED)
- onboard proof (sandbox user)                       -> 999 (NvRmMemInitNvmap Permission denied)

Groups alone (video=44 + 995) are insufficient for the non-root user to init CUDA on this IGX Orin.

Workaround: NEMOCLAW_SANDBOX_GPU=0 (CPU) lets onboarding complete.

Logs

Sandbox CUDA proof failed: cuInit(0) via libcuda.so.1
  NvRmMemInitNvmap failed with Permission denied 356: Memory Manager Not supported
  ****NvRmMemMgrInit failed**** error type: 196626 cuInit(0)=999
proof_error=Sandbox GPU proof returned failed status: cuInit(0) via libcuda.so.1

Related

Activity

  1. added
    area: onboardingOnboarding FSM, provider setup, sandbox launch, or first-run flow
    area: sandboxOpenShell sandbox lifecycle, runtime, config, or recovery
    on Jul 28, 2026
  2. wscurran commented on Jul 28, 2026

    @wscurran
    Contributor

    ✨ Thanks for the detailed report. The reproduction steps, environment, and isolation findings are clear. Maintainers will investigate the non-root sandbox user GPU device access on JetPack 6.2 IGX Orin.

  3. added theissue type on Jul 28, 2026
  4. added a commit that references this issue on Jul 28, 2026
    0fe7ed7
  5. wangericnv commented on Jul 30, 2026

    @wangericnv
    CollaboratorAuthor

    Retested on v0.0.97 — still reproduces (render-group fix insufficient)

    Retested on 2026-07-30 on an NVIDIA IGX Orin Developer Kit (aarch64, JetPack 6.2 / L4T R36.5.1) with NemoClaw v0.0.97 (OpenShell 0.0.85). The bug still reproduces — the render-device-group fix is not sufficient on IGX Orin JP6.2.

    The fix is applied — onboarding now grants the sandbox user the render group:

    Granting sandbox user the detected Jetson GPU device groups via --group-add 44, 109, 995
    (so CUDA can initialize as a non-root user)
    

    (44 = video, 109 = render, 995 = debug)

    But the end-to-end sandbox GPU proof still fails for the non-root sandbox user with the same error as the original report:

    Sandbox CUDA proof failed: cuInit(0) via libcuda.so.1
      NvRmMemInitNvmap failed with Permission denied 356: Memory Manager Not supported
      ****NvRmMemMgrInit failed**** error type: 196626 cuInit(0)=999
    GPU proof failed inside an executable sandbox
    

    Onboarding still aborts; NEMOCLAW_SANDBOX_GPU=0 is still required to complete onboarding.

    Root cause of the residual failure

    The Tegra GPU nodes under /dev/nvgpu/igpu0/* are owned root:video (44) and root:debug (995), which the group-add already covers. But cuInit still fails at NvRmMemInitNvmap — the Tegra nvmap memory-manager path. Adding the render / video / debug groups does not grant the non-root sandbox user whatever access NvRmMemInitNvmap requires on IGX Orin JP6.2.

    The fix likely needs to also cover the nvmap access path for the non-root sandbox user, not just the render device group.

  6. 71 remaining items

  7. cjagwani commented on Sep 15, 2026

    @cjagwani
    Collaborator

    Completion status (2026-09-15)

    The product direction is now accepted in #8910: OpenShell native CDI is the single authority for Jetson device injection, mounts, supplemental groups, and hardware-derived policy; NemoClaw will consume and verify the released contract rather than ship the legacy compatibility recreation. Supported hardware remains AGX Thor and IGX Orin.

    The remaining path is upstream-gated:

    Once those changes land in a released OpenShell build, #8910 must be reduced to native-CDI adoption and then run the accepted hardware matrix on the exact supported artifacts: official ARM64 image, onboarding exit 0, non-root nvidia-smi, /proc/<pid>/task/<tid>/comm write, cuInit(0)=0, restart/resume/rebuild, non-GPU negative, and invalid/missing CDI fail-closed behavior on both AGX Thor and IGX Orin. The earlier Thor GPU-stage success is directional evidence only; it did not complete onboarding and does not close this issue.

  8. wangericnv commented on Oct 8, 2026

    @wangericnv
    CollaboratorAuthor

    QA question rather than a reopen request — I may be missing context that justifies the close, and I would rather ask than flip the state on you.

    I attempted to re-verify this on NemoClaw v0.0.131 (IGX Orin, JetPack 6.2 / L4T R36.5.1, aarch64) and could not complete the end-to-end run: onboarding aborts before the sandbox GPU proof with an unrelated dashboard-port/forward condition on that host, so the GPU gate was never exercised. I am making no claim about whether the symptom still occurs.

    Three things I noticed while preparing that run, which is why I am asking:

    1. This issue has no closing PR. The 2026-10-06 close carries no linked pull request. The earlier close (2026-07-28, PR fix(onboard): add Jetson render device group #7762) was reopened two days later, so an ancestry check against a release tag reports fix(onboard): add Jetson render device group #7762 as "contained" and looks green — which is misleading for anyone using release containment as a gate. That is what prompted me to look at the history rather than trust the check.

    2. Two of the three upstream gates you listed are still open. OpenShell refactor(cli): migrate status and tunnel commands to oclif #2775 closed on 2026-10-05, the day before this issue was closed, so I assume that is what unblocked it. But OpenShell feat(cli): add resource profiles to nemoclaw onboard and nemoclaw resources command #3348 (minimum CUDA-required Jetson sysfs policy) and docs(reference): document remote-deploy env vars #3349 (legitimate 301-path CDI policy above the 256 limit) are both still open. Your 2026-09-15 comment listed all three as gating.

    3. Which OpenShell build carries the fix? NemoClaw v0.0.131 ships openshell 0.0.116. If the CDI policy resolver from OpenShell refactor(cli): migrate status and tunnel commands to oclif #2775 is what resolves this, knowing the first OpenShell build that contains it would let QA target the right NemoClaw release instead of guessing.

    The question: was the acceptance matrix from your 2026-09-15 comment actually run — official ARM64 image, onboarding exit 0, non-root nvidia-smi, /proc/<pid>/task/<tid>/comm write, cuInit(0)=0, restart/resume/rebuild, non-GPU negative, and fail-closed on invalid/missing CDI — on IGX Orin? Or did tracking move upstream, with this issue closed as "no longer tracked here"? Either answer is fine; QA just needs to know which, because the two imply different next steps on our side.

    One data point from the host, since it is relevant to the boundary you isolated. On this IGX Orin under v0.0.131, direct Docker with the same image and a non-root user (uid 998, --group-add 44 --group-add 995) returns:

    cuInit(0) = 0
    

    That is the "direct Docker succeeds" half of the comparison you described. I was not able to produce the OpenShell-exec half, because onboarding does not get that far on this host.

    Happy to run the full acceptance matrix on IGX Orin once I can get onboarding past the forward-ownership blocker — just say the word and I will report either way.

  9. wangericnv commented on Oct 8, 2026

    @wangericnv
    CollaboratorAuthor

    Answering my own question above, empirically — I got the end-to-end run to complete on IGX Orin, and the symptom still reproduces on v0.0.131.

    ✓ GPU proof passed: nvidia-smi when available
    ✓ GPU proof passed: /proc/<pid>/task/<tid>/comm write
    ⚠ GPU proof inconclusive: cuInit(0) via libcuda.so.1
      NvRmMemInitNvmap failed with Permission denied 356: Memory Manager Not supported
      ****NvRmMemMgrInit failed**** error type: 196626 cuInit(0)=999
    ⚠ Sandbox CUDA proof failed: cuInit(0) via libcuda.so.1
    GPU proof failed inside an executable sandbox.
    

    Same two passing proofs, same signature as the original report.

    Two details that matter for where this sits:

    • It was not the compatibility path. The run used NEMOCLAW_DOCKER_GPU_PATCH=0, and the log confirms Direct sandbox GPU enabled; allowing OpenShell GPU policy enrichment. So this is the native-injection direction from feat(onboard): preserve Jetson GPU device groups #8910, against the OpenShell build v0.0.131 ships.
    • The boundary you isolated still holds. On the same host, same image, same unprivileged identity, directly under Docker: cuInit(0) = 0. Only the sandbox execution path fails.

    I filed this as #12806 rather than asking you to reopen here, since this issue has been closed twice and a comment on a closed thread is easy to miss. If you would rather track it on this issue, say so and I will close #12806 as a duplicate — no attachment to either arrangement.

    Earlier in this thread I had only the paper trail and explicitly made no claim about the symptom. That gap is now closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

VDRLinked to VDR findingarea: onboardingOnboarding FSM, provider setup, sandbox launch, or first-run flowarea: sandboxOpenShell sandbox lifecycle, runtime, config, or recoveryplatform: jetsonAffects Jetson AGX Thor or Orin

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions