CVE-2026-54234
vllm-project vllm
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter's input ids; that out-of-vocabulary value is later consumed by the model's embedding and attention path and crashes the engine worker with a GPU device-side assertion. The same triggerin...
- CVSS
- 7.5
- EPSS
- 0.36% 28.9% percentile
- CISA KEV
- Not listed
- Published
- 2026.07.07