RSS Anyway
Sign in
RSS Anyway
Hot
Latest
Following
Status
About
Sign in
RSS Anyway
Hot
Latest
Following
Status
About
vllm.ai
Sign in to follow
vllm.ai
RSS
Atom
JSON
items
|
feeds
01.
A Preview of Production-Scale Kimi K3 Support on vLLM
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 22
02.
Beyond a Single Model: Building Mixture-of-Models Systems with vLLM Semantic Router
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 21
03.
Keeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release Process
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 16
04.
TML Inkling on vLLM: Day-0 Support with Optimized Performance
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 15
05.
vLLM x TileRT: Specialized Decode for Latency-Critical Serving
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 14
06.
EAGLE-3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD Quark
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 13
07.
vime + ROCm: End-to-End RL Post-Training on AMD Instinct™ GPUs
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 10
08.
vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 6
09.
Experience and Lessons Learned from Serving Multi-Stage Qwen3-Omni in vLLM-Omni
vllm.ai
·
/blog/rss.xml
▲ 0
· Jul 1
10.
Micro-Agent: Beat Frontier Models with Collaboration inside Model API
vllm.ai
·
/blog/rss.xml
▲ 81
· Jun 29
·
22 comments on HN
11.
Engineering TTS Inference in vLLM-Omni
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 23
12.
Beyond One Model: Fusion in vLLM Semantic Router
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 16
13.
MiniMax M3 in vLLM: Day-0 Serving for 1M-Token Multimodal Reasoning
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 12
14.
DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 10
15.
Announcing vime: A Simple, Stable, and Efficient RL Framework for LLMs
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 9
16.
vLLM Semantic Router v0.3 Themis: From Signals to Stateful Production Routing
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 5
17.
Announcing Day-0 Support for NVIDIA Nemotron 3 Ultra on vLLM
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 4
18.
Fast & Efficient LLM Inference with vLLM: A New Course with DeepLearning.AI
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 3
19.
Accelerating vLLM-Omni Inference with AutoRound Quantization
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 2
20.
Session-Aware Agentic Routing: Continuity-Aware Model Selection for Long-Horizon LLM Agents
vllm.ai
·
/blog/rss.xml
▲ 0
· Jun 2
21.
vLLM on the DGX Spark: Architecture, Configuration, and Local Evaluation
vllm.ai
·
/blog/rss.xml
▲ 0
· May 31
22.
Speculators v0.5.0: DFlash Support and Online Training
vllm.ai
·
/blog/rss.xml
▲ 0
· May 28
23.
From Text to Multimodal Routing: Hardening Vision Signals in vLLM Semantic Router
vllm.ai
·
/blog/rss.xml
▲ 0
· May 28
24.
Native RL APIs in vLLM
vllm.ai
·
/blog/rss.xml
▲ 0
· May 28
25.
Accelerating Laguna XS.2 Inference with vLLM, Speculators, and LLM Compressor
vllm.ai
·
/blog/rss.xml
▲ 0
· May 28
26.
EAGLE 3.1: Advancing Speculative Decoding Through Collaboration Between the EAGLE Team, vLLM, and TorchSpec
vllm.ai
·
/rss
▲ 69
· May 26
·
23 comments on HN
27.
vLLM x Novita AI: PegaFlow for Production-Grade External KV Cache
vllm.ai
·
/blog/rss.xml
▲ 0
· May 18
28.
Announcing VeRL-Omni: Easy, Fast, and Stable RL Training for Diffusion and Omni-Modality Models
vllm.ai
·
/blog/rss.xml
▲ 0
· May 14
29.
Elastic Expert Parallelism in vLLM
vllm.ai
·
/blog/rss.xml
▲ 0
· May 14
30.
vLLM Tops the Artificial Analysis Leaderboard
vllm.ai
·
/blog/rss.xml
▲ 0
· May 11
page 1
next →