Skip to content

Add OpenAI-compatible server and AIPerf benchmarks - #9

Closed
HanFa wants to merge 0 commit into
mainfrom
codex/aiperf-benchmarks
Closed

HanFa wants to merge 0 commit into
mainfrom
codex/aiperf-benchmarks

Conversation

@HanFa

@HanFa HanFa commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Add an OpenAI-compatible HTTP entry point so clients can send concurrent requests to a shared tiny-vLLM Engine and benchmark it through a public endpoint.

  • Add tiny-vllm-serve / python -m tiny_vllm.server with /v1/completions, /v1/models, and /health. A single model worker runs engine steps while HTTP stays responsive, with bounded admission, context checks, JSON errors, and graceful draining.
  • Support the current engine's non-streaming, greedy, fixed-length completions with accurate usage counts. Unsupported options are rejected. Engine, scheduler, KV-cache, and generation data contracts are unchanged.
  • Add native AIPerf configs for seven baseline scenarios, KV pressure, and mock smoke testing, plus an optional manifest/result-validation wrapper. AIPerf uses only the completions endpoint; server metadata can be supplied separately.

Validation

  • Local suite with model dependencies: 82 passed, 2 opt-in integration tests skipped. Fresh [dev] environment: 72 passed, 4 optional-model skips; Ruff and mypy passed. Include the HTTP dependency in [dev] so the standard checks also cover serving.
  • OpenAI Python SDK model discovery, real-model completion, and error handling passed against sshleifer/tiny-gpt2 on CPU.
  • All three AIPerf configs validated. The seven baseline scenarios each completed eight measured requests with zero request errors or output-length mismatches. The mock smoke wrapper also passed.

These runs verify functionality. GPU performance and a live vLLM comparison remain unverified because the GPU host was unreachable. Streaming/chat APIs and internal profiler hooks are outside this slice; prompt preflight currently adds a tokenization pass to API latency.

@HanFa

HanFa commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

/profile

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

GPU profiling queued for fe833e410b34cc53fd28540d4d6ebd5ff72d2115. Run. The job waits for the RTX 5090; it does not stop other workloads.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

GPU profiling did not complete (skipped). Tested commit: fe833e410b34cc53fd28540d4d6ebd5ff72d2115. Check GPU availability and run logs.

Logs and raw artifacts

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

GPU profiling queued for fe833e410b34cc53fd28540d4d6ebd5ff72d2115. Run. The job waits for the RTX 5090; it does not stop other workloads.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

GPU profiling did not complete (build: cancelled; profile: cancelled). Tested commit: fe833e410b34cc53fd28540d4d6ebd5ff72d2115. Check the failed job's logs.

Run logs and available artifacts

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

GPU profiling queued for fe833e410b34cc53fd28540d4d6ebd5ff72d2115. Run. The job waits for the RTX 5090; it does not stop other workloads.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

tiny-vLLM / vLLM GPU profile

Source: fe833e410b34cc53fd28540d4d6ebd5ff72d2115. Run: gh-36935497541.
GPU: NVIDIA GeForce RTX 5090 on sutro-gpu1; CUDA, float16.
Model: openai-community/gpt2 at 607a30d783dfa663caf39e06633721c8d4cfcd7e.
3 repeats × 100 measured requests per scenario; warmup excluded. Values are medians across runs; percentile columns are medians of per-run percentiles, not pooled percentiles.

Scenario tiny p50 ms vLLM p50 ms tiny p95 ms vLLM p95 ms tiny output tok/s vLLM output tok/s vLLM / tiny tok/s
batch_c2 229.83 74.94 234.34 78.10 276.96 842.67 3.04×
batch_c4 268.93 78.91 284.36 83.48 469.86 1598.17 3.40×
decode_32_128 983.13 276.27 1018.16 296.76 517.39 1826.24 3.53×
latency_c1 194.27 70.87 218.26 74.43 161.41 443.99 2.75×
mixed_prefill 271.91 79.14 289.42 85.16 461.60 1585.83 3.44×
prefill_512_16 256.20 76.27 262.78 78.52 245.34 812.73 3.31×
queue_c8 538.60 145.26 548.38 151.81 471.03 1728.09 3.67×

Both servers ran sequentially on the same GPU. All measured requests succeeded with matching input/output lengths. Generation was greedy, fixed-length and non-streaming. TTFT and inter-token latency are unavailable.

vLLM 0.10.2 used FLASH_ATTN, eager mode, and disabled prefix caching. Both used 4 active requests, a 64-token step budget and 256 logical blocks of 16 tokens. This is a controlled eager baseline, not a comparison against vLLM's fastest default configuration.

This GPT-2 workload measures one model and serving configuration, including scheduling and kernel-launch overhead. It does not predict larger-model throughput. The node also hosts other CPU workloads; the GPU resource was exclusively reserved by each benchmark server.

Images and environment

  • tiny: ghcr.io/sutro-planet/tiny-vllm/profile-server@sha256:660333066235b79ab63ed14bd18210fa9c5537b6cb0126624f7237fed7fe402f; endpoint http://tiny-gh-36935497541.tiny-vllm-profile.svc:8000/v1/completions.
  • vllm: docker.io/vllm/vllm-openai@sha256:df2607b26bdda2875de4832f4d08da0055b4b6e3570347f3a849bcc652771dd6; endpoint http://vllm-gh-36935497541.tiny-vllm-profile.svc:8000/v1/completions.

See manifest.json, each server's environment.json and server.log, and raw AIPerf exports for exact versions, GPU state and request metrics.

Run logs and available artifacts

@HanFa HanFa closed this Oct 1, 2026
@HanFa
HanFa force-pushed the codex/aiperf-benchmarks branch from fe833e4 to fa8cdf4 Compare October 1, 2026 23:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant