1 matches found
PT-2026-64142
Summary A frontend-legal multi-request speculative workload can make vLLM produce an out-of-vocabulary recovered token equal to vocab size, convert that value to -1 when choosing the next live token for a request, and then feed that -1 back into the next drafter input ids. On Qwen3 GPTQ this...
7.5CVSS5.6AI score0.00616EPSS
SaveExploits1References8
20