Runs, not claims.
Reviewed studies first, controls and partial runs clearly labeled, and operational evidence kept separate. Every card says what kind of conclusion it can support.
Published studies
Completed runs with bounded conclusions and reproducible evidence.
Qwen3-30B-A3B FP8: H200 prefill → MI300X decode
Dynamo and vLLM on both nodes across 21 exact-length points, prefix caching disabled, with independent KV transfer byte verification.
Nemotron 3.5: Dynamo versus llm-d P/D
Matched prefill/decode matrices across two Spark nodes at concurrency 1–32, with identical cache policy and zero request errors.
Qwen3.5-27B framework and precision sweep
vLLM and SGLang across BF16, FP8, and GPTQ-Int4, including documented unsupported configurations.
Single-replica controls
Per-replica capacity and wrapper overhead—not routing or horizontal scaling.
Qwen3.8-27B aggregated serving controls
One model replica per arm at fixed concurrency. This isolates frontend/orchestration behavior; it does not test multi-worker routing.
Nemotron 3.5 aggregated serving controls
One model replica per arm across concurrency 1–32. Use for per-replica throughput and wrapper overhead only.
Partial and operational work
Transparent notebook evidence, separated from reviewed findings.
Qwen3.8-27B P/D recovery set
Dynamo completed 18/18 points; 5/18 llm-d points were recovered before the MI300X host became unavailable.
Serving optimization and ISL/OSL sweeps
Qwen2.5-7B, Nemotron-120B, and Qwen3.5-27B studies spanning prefix caching, chunked prefill, KV precision, and concurrency scaling.
Observability, SLO, and failure simulation
Fifteen operational tasks covering DCGM, OpenTelemetry, Grafana, alerting, traces, and inference failure drills.
Quality and release checks
CI-fed model evaluation and baseline comparison.
Inspect charts, experiment registry, and machine-readable result files.