-
-
Notifications
You must be signed in to change notification settings - Fork 21.3k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Bugfix] Fix inverted text validation in FireRedASR2/FunASR processors
bug
Something isn't working
#54001
opened Aug 27, 2026 by
hungnnvidia
Contributor
•
Draft
[Bugfix] Ignore non-dict JSON in model redirect file
bug
Something isn't working
#54000
opened Aug 27, 2026 by
hungnnvidia
Contributor
•
Draft
[Bugfix] Raise clear error on interleaved multimodal placeholder overcount
bug
Something isn't working
frontend
#53999
opened Aug 27, 2026 by
hungnnvidia
Contributor
•
Draft
[Quantization] Skip absent fused-layer shards in the GPTQ per-layer check (Gemma 4 k_eq_v)
quantization
#53996
opened Aug 27, 2026 by
CySpiegel
Loading…
[ROCm] add AITER MLA prefill CI coverage
ci/build
rocm
Related to AMD ROCm
#53995
opened Aug 27, 2026 by
CuongTranXuan
Loading…
fix: define math constants for FlashMLA MSVC builds
ci/build
#53994
opened Aug 27, 2026 by
CuongTranXuan
Loading…
[Bugfix] Fix out-of-bounds attrIdxs in C++ batch memcpy path (cache_kernels.cu)
bug
Something isn't working
#53991
opened Aug 27, 2026 by
Vikas8719
Loading…
4 tasks
[Bugfix][XPU] Fall back to torch query when getMemoryInfo reports zero free memory
bug
Something isn't working
intel-gpu
Related to Intel GPU
#53990
opened Aug 27, 2026 by
CySpiegel
Loading…
[Model][XPU] Enable fused QK-norm+RoPE+gate Triton kernel on XPU
intel-gpu
Related to Intel GPU
qwen
Related to Qwen models
#53989
opened Aug 27, 2026 by
CySpiegel
Loading…
[Spec Decode][ROCm] Add FLy: entropy-gated deferred verification for draft-model speculative decoding
documentation
Improvements or additions to documentation
rocm
Related to AMD ROCm
speculative-decoding
verified
Run pre-commit for new contributors without triggering other tests
#53987
opened Aug 27, 2026 by
eecspan
Loading…
[XPU][CI] increase timeout of shutdown
intel-gpu
Related to Intel GPU
#53986
opened Aug 27, 2026 by
zhenwei-intel
Contributor
•
Draft
4 tasks
Fix: Prevent SageMaker logger standard from suppressing vLLM startup logs
frontend
#53985
opened Aug 27, 2026 by
DresdenGman
Loading…
[Bugfix] Route directly constructed embeddings through CPU offload
bug
Something isn't working
#53981
opened Aug 27, 2026 by
waizuichougou
Loading…
[XPU][TEST]Add entrypoints test in Intel GPU CI
ci/build
intel-gpu
Related to Intel GPU
#53980
opened Aug 27, 2026 by
zxd1997066
Contributor
Loading…
4 tasks done
[Attention][Spec Decode] NVFP4 KV: open the FA2 non-causal prefill wrapper for DFlash-family drafters (sm12x)
dflash
nvidia
quantization
#53979
opened Aug 27, 2026 by
TechPrototyper
Loading…
[Bugfix][Spec Decode] DFlash2: survive spec warmup with unfilled draft buffers (embed outside compile; clamp selector gathers)
bug
Something isn't working
dflash
mrv2
Model Runner V2 specific
qwen
Related to Qwen models
speculative-decoding
#53978
opened Aug 27, 2026 by
TechPrototyper
Loading…
[Bugfix][Spec Decode] Mask out-of-range input ids in VocabParallelEmbedding at tp_size==1
bug
Something isn't working
#53977
opened Aug 27, 2026 by
TechPrototyper
Loading…
[Bugfix][Test] Fix off-by-one error in sampled token rank causing flaky logprobs test
bug
Something isn't working
#53976
opened Aug 27, 2026 by
mayuyuace
Contributor
Loading…
Revert "[Bugfix][ROCm][Disagg] Fix MoRIIO shared KV memory region registration" (#53698)
bug
Something isn't working
kv-connector
rocm
Related to AMD ROCm
#53974
opened Aug 27, 2026 by
vllm-agent
Contributor
•
Draft
[Model][Perf] Use broadcast mHC pre for DeepSeek V4 DSpark
deepseek
Related to DeepSeek models
dflash
DSv4
#53972
opened Aug 27, 2026 by
liuyao0322
Loading…
4 tasks done
[Bugfix] Fix out-of-bounds attrIdxs in C++ batch memcpy path
bug
Something isn't working
#53971
opened Aug 27, 2026 by
dzc2000
Loading…
[BUGFIX] Fix dflash batched token budget
bug
Something isn't working
dflash
mrv2
Model Runner V2 specific
speculative-decoding
#53970
opened Aug 27, 2026 by
Orvid
Loading…
[Bugfix] Support NoPE models on FLASHINFER_MLA_SPARSE_SM120 and validate effective topk buffer width
bug
Something isn't working
nvidia
#53969
opened Aug 27, 2026 by
hamiltongaianimd
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.