Tags
performance62ecosystem33hardware25model-support21multimodal19speculative-decoding17large-scale-serving15disaggregation11quantization10moe8reinforcement-learning6community6kv_cache5developer5vllm-omni4speculators3post-training3models3attention3inference3distributed2agentic2parallelism2rlhf2dflash2prefix caching2agentic-routing2watermarking1sampling1apple-silicon1qwen3.81kernels1glm1kv-cache1fastvideo1fasth31amd1cosmos31qwen3.51speculative_decoding1peagle1dspark1mixture-of-models1semantic-router1ci1evaluation1release1hpc-ops1minimax1day-0-support1long-context1model1learning1dgx-spark1nemotron1deployment1computex1llm-compressor1async-rl1production-serving1elastic-ep1expert-parallelism1fault-tolerance1turboquant1benchmarking1kernel-fusion1fp81mamba1engineering1triton1frontend1