Repository navigation
feat: support shared-GPU simulation evaluation - #6
Merged
Merged
Conversation
Pin MuJoCo rendering independently of HTTP inference replicas, allowing GPU0 to render while already-running services on GPU0 and GPU1 both serve inference. Statically shard episodes across process-isolated workers and keep each worker's sessions on one replica. Reuse the existing SimulationService lifecycle and original inference batcher. Add an evaluation CLI, paired placement configurations, result accounting, worker timeout and failure handling. Leave EmbodiInfer unchanged. This integrates the deployment used by the earlier external-driver 1.762x warmed-loop result (1.485x including reset). This native runner has not been GPU-benchmarked. Validation before removing the added test file from this commit: 1044 tests passed, 45 optional-dependency tests skipped; Ruff checks passed. Runtime code is unchanged by the documentation/test removal.
ZhouAo-ZA
approved these changes
Oct 7, 2026
Collaborator
|
我看了下spec,并且在gpu服务器上跑了下应该没啥问题,让codex看了下代码这可能还有两个小点可以修改一下 GPU 小规模测试情况:
|
hootandy321
reviewed
Oct 8, 2026
hootandy321
left a comment
Collaborator
There was a problem hiding this comment.
- 第 205–207 行:异常退出时的 session 清理
超时或其他 worker 报错时,父进程会直接 terminate 子进程,可能跳过 finally 里的 DELETE。当前服务端没有 session TTL,遗留 session 会占住配额,累积后新请求就会返回 429,这个在 macOS 和 Linux 上都复现了。
建议先让 worker 协作退出,再由父进程追踪并兜底清理 session。只补 finally 还不够,因为硬终止时它仍可能不执行。 - 第 146 行:ID 加后缀后可能超长
simulator_id 追加“-0”等 worker 后缀后,原本合法的 127/128 字符 ID 会超过接口的 128 字符上限,创建 session 返回 400。可以用 metadata 区分 worker,或者生成长度受限、不会冲突的派生 ID。
麻烦给这两项补一下回归测试:超时、其他 worker 失败后的清理和配额恢复,以及 ID 长度边界。这两项修好、测试和 CI 通过后,就可以合并这版功能。
hootandy321
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pin MuJoCo rendering independently of HTTP inference replicas, allowing GPU0 to render while already-running services on GPU0 and GPU1 both serve inference. Statically shard episodes across process-isolated workers and keep each worker's sessions on one replica.
Reuse the existing SimulationService lifecycle and original inference batcher. Add an evaluation CLI, paired placement configurations, result accounting, worker timeout and failure handling. Leave EmbodiInfer unchanged.
This integrates the deployment used by the earlier external-driver 1.762x warmed-loop result (1.485x including reset). This native runner has not been GPU-benchmarked.
Validation before removing the added test file from this commit: 1044 tests passed, 45 optional-dependency tests skipped; Ruff checks passed. Runtime code is unchanged by the documentation/test removal.
Summary
Type of change
Verification
Paste the exact commands you ran and their result. State what you did not
verify (GPU, checkpoint, real hardware) rather than implying it works.
Process boundaries
client,deployment,application,devices,model_servicesremain the canonical domains, and
services.*still imports.src/embodirun.pyproject.toml.Checklist
docs/en/support-matrix.mdfor any new combination, with versions, configuration, hardware,
checkpoint, exact command, and observed result.
steps, or the support matrix changed.
data are included.
CONTRIBUTING.md.