Conversation
|
/profile |
|
GPU profiling queued for |
|
GPU profiling did not complete (skipped). Tested commit: |
|
GPU profiling queued for |
|
GPU profiling did not complete (build: cancelled; profile: cancelled). Tested commit: |
|
GPU profiling queued for |
tiny-vLLM / vLLM GPU profileSource:
Both servers ran sequentially on the same GPU. All measured requests succeeded with matching input/output lengths. Generation was greedy, fixed-length and non-streaming. TTFT and inter-token latency are unavailable. vLLM 0.10.2 used FLASH_ATTN, eager mode, and disabled prefix caching. Both used 4 active requests, a 64-token step budget and 256 logical blocks of 16 tokens. This is a controlled eager baseline, not a comparison against vLLM's fastest default configuration. This GPT-2 workload measures one model and serving configuration, including scheduling and kernel-launch overhead. It does not predict larger-model throughput. The node also hosts other CPU workloads; the GPU resource was exclusively reserved by each benchmark server. Images and environment
See |
fe833e4 to
fa8cdf4
Compare
Summary
Add an OpenAI-compatible HTTP entry point so clients can send concurrent requests to a shared tiny-vLLM Engine and benchmark it through a public endpoint.
tiny-vllm-serve/python -m tiny_vllm.serverwith/v1/completions,/v1/models, and/health. A single model worker runs engine steps while HTTP stays responsive, with bounded admission, context checks, JSON errors, and graceful draining.Validation
[dev]environment: 72 passed, 4 optional-model skips; Ruff and mypy passed. Include the HTTP dependency in[dev]so the standard checks also cover serving.sshleifer/tiny-gpt2on CPU.These runs verify functionality. GPU performance and a live vLLM comparison remain unverified because the GPU host was unreachable. Streaming/chat APIs and internal profiler hooks are outside this slice; prompt preflight currently adds a tokenization pass to API latency.