Skip to content

feat: support shared-GPU simulation evaluation - #6

Merged
ZhouAo-ZA merged 1 commit into
BUAA-CI-LAB:mainfrom
yufoo1:feat/shared-gpu-evaluation
Oct 9, 2026
Merged

ZhouAo-ZA merged 1 commit into
BUAA-CI-LAB:mainfrom
yufoo1:feat/shared-gpu-evaluation

Conversation

@yufoo1

@yufoo1 yufoo1 commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Pin MuJoCo rendering independently of HTTP inference replicas, allowing GPU0 to render while already-running services on GPU0 and GPU1 both serve inference. Statically shard episodes across process-isolated workers and keep each worker's sessions on one replica.

Reuse the existing SimulationService lifecycle and original inference batcher. Add an evaluation CLI, paired placement configurations, result accounting, worker timeout and failure handling. Leave EmbodiInfer unchanged.

This integrates the deployment used by the earlier external-driver 1.762x warmed-loop result (1.485x including reset). This native runner has not been GPU-benchmarked.

Validation before removing the added test file from this commit: 1044 tests passed, 45 optional-dependency tests skipped; Ruff checks passed. Runtime code is unchanged by the documentation/test removal.

Summary

Type of change

  • Bug fix
  • New capability (robot, simulator, model, backend, or runtime feature)
  • Documentation
  • Refactor with no behaviour change
  • Build, packaging, or CI

Verification

Paste the exact commands you ran and their result. State what you did not
verify (GPU, checkpoint, real hardware) rather than implying it works.

$ uv run pytest -q

Process boundaries

  • client, deployment, application, devices, model_services
    remain the canonical domains, and services.* still imports.
  • No inference-engine runtime code was imported into src/embodirun.
  • Robot/simulator adapters and policy bindings stay separate.
  • New robot- or model-specific dependencies are optional and declared in
    pyproject.toml.

Checklist

  • I updated docs/en/support-matrix.md
    for any new combination, with versions, configuration, hardware,
    checkpoint, exact command, and observed result.
  • I kept the English and Chinese READMEs in sync, if positioning, install
    steps, or the support matrix changed.
  • No checkpoints, datasets, recordings, credentials, addresses, or personal
    data are included.
  • I agree that this contribution is licensed under Apache-2.0, per
    CONTRIBUTING.md.

Pin MuJoCo rendering independently of HTTP inference replicas, allowing
GPU0 to render while already-running services on GPU0 and GPU1 both serve
inference. Statically shard episodes across process-isolated workers and
keep each worker's sessions on one replica.

Reuse the existing SimulationService lifecycle and original inference
batcher. Add an evaluation CLI, paired placement configurations, result
accounting, worker timeout and failure handling. Leave EmbodiInfer unchanged.

This integrates the deployment used by the earlier external-driver 1.762x
warmed-loop result (1.485x including reset). This native runner has not
been GPU-benchmarked.

Validation before removing the added test file from this commit: 1044
tests passed, 45 optional-dependency tests skipped; Ruff checks passed.
Runtime code is unchanged by the documentation/test removal.
@hootandy321

hootandy321 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

我看了下spec,并且在gpu服务器上跑了下应该没啥问题,让codex看了下代码这可能还有两个小点可以修改一下

GPU 小规模测试情况:

  • 环境:Linux、NVIDIA L20,使用单卡。
  • 模型与后端:π0.5 LIBERO 微调权重,固定版本 8e174154;原生 Pi05ServingAdapter 与 HTTP 服务,8 维状态输入、7 维动作输出,10 步去噪。
  • Tokenizer:使用官方哈希一致的 SentencePiece 文件生成本地配置,242 项指令、状态边界及截断检查通过。
  • 测试方式:对比 main 顺序执行、PR 单 worker 和双 worker。每组 2 个 episode、每个推进 3 步,共 18 次真实策略 HTTP 请求,使用受控采样随机性。
  • 结果:三组短轨迹的状态、双相机图像哈希和实际动作逐字段一致,正常结束时 session 均成功释放。

@hootandy321 hootandy321 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. 第 205–207 行:异常退出时的 session 清理
    超时或其他 worker 报错时,父进程会直接 terminate 子进程,可能跳过 finally 里的 DELETE。当前服务端没有 session TTL,遗留 session 会占住配额,累积后新请求就会返回 429,这个在 macOS 和 Linux 上都复现了。
    建议先让 worker 协作退出,再由父进程追踪并兜底清理 session。只补 finally 还不够,因为硬终止时它仍可能不执行。
  2. 第 146 行:ID 加后缀后可能超长
    simulator_id 追加“-0”等 worker 后缀后,原本合法的 127/128 字符 ID 会超过接口的 128 字符上限,创建 session 返回 400。可以用 metadata 区分 worker,或者生成长度受限、不会冲突的派生 ID。
    麻烦给这两项补一下回归测试:超时、其他 worker 失败后的清理和配额恢复,以及 ID 长度边界。这两项修好、测试和 CI 通过后,就可以合并这版功能。

@ZhouAo-ZA
ZhouAo-ZA merged commit 4c5c49d into BUAA-CI-LAB:main Oct 9, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants