moonshotai/Kimi-K2.5
Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes
Multimodal agentic MoE model with DeepSeek-V3 backbone and MLA attention
Guide
Overview
Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.
Prerequisites
- vLLM version: >= 0.15.0 (speculative decoding with Eagle3 requires >= 0.18.0)
- Hardware (BF16): 8x H200 GPUs (verified), or equivalent aggregate VRAM (~640 GB)
- Hardware (NVFP4): 4x Blackwell GPUs (e.g. GB200)
- AMD support: 8x MI300X / MI325X / MI355X with ROCm 7.2.1 and Python 3.12
Install vLLM
Pip (NVIDIA):
uv venv
source .venv/bin/activate
uv pip install vllm --torch-backend auto
Pip (AMD ROCm):
uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm
Docker (NVIDIA):
docker pull vllm/vllm-openai:latest
AMD MI300X/MI325X
On 8x MI300X or MI325X (gfx942), use the standard W4A16 MoE path with AITER
and INT4 QuickReduce.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
vllm serve moonshotai/Kimi-K2.5 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--tensor-parallel-size 8 \
--tool-call-parser kimi_k2 \
--enable-auto-tool-choice \
--reasoning-parser kimi_k2 \
--mm-encoder-tp-mode data
AMD MI350X/MI355X
On 8x MI350X or MI355X (gfx950), add --moe-backend flydsl to use the
optimized FlyDSL W4A16 MoE kernel. Keep LoRA disabled for this path.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
vllm serve moonshotai/Kimi-K2.5 \
--tensor-parallel-size 8 \
--trust-remote-code \
--mm-encoder-tp-mode data \
--moe-backend flydsl \
--compilation-config '{"pass_config": {"fuse_allreduce_rms": false}}'
Notes:
- The FlyDSL INT4 MoE path does not support expert parallelism; do not add
--enable-expert-parallel. - Keep
--compilation-config '{"pass_config": {"fuse_allreduce_rms": false}}'; it is required for this FlyDSL path on MI350X / MI355X. - vLLM has tuned MI350X/MI355X FlyDSL configs for this Kimi shape at TP=8 and TP=4.
- Keep vLLM's default block size unless you are tuning long-context
throughput;
--block-size 64is safe to try.
AMD MI355X (MXFP4 checkpoint)
The MXFP4 variant serves the AMD Quark checkpoint
amd/Kimi-K2.5-MXFP4 — a
different quantization from the packed-INT4 moonshotai/Kimi-K2.5 FlyDSL path
above. It runs the AITER MXFP4 MoE path and adds an fp8 KV cache.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
vllm serve amd/Kimi-K2.5-MXFP4 \
--tensor-parallel-size 8 \
--trust-remote-code \
--mm-encoder-tp-mode data \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.90
Notes:
- fp8 KV cache is the main MI355X win: it doubles resident KV capacity and halves KV read bandwidth. Measured +27% throughput at 8k/1k input with no accuracy change (GSM8K 0.9697, unchanged from bf16 KV).
- AITER RMSNorm stays on. A local TP4 GSM8K check on the current nightly measured 0.9697.
- Old firmware only: if
rocm-smi --showfwreports a MEC firmware version below 177, addexport HSA_NO_SCRATCH_RECLAIM=1— older firmware can't reclaim RCCL scratch and vLLM crashes without it. Current MI355X firmware doesn't need it. - For maximum throughput on fixed-length benchmark 8k/1k or 1k/1k workloads, see the benchmark reproduction section below. These are throughput-sweep tunings; leave vLLM's defaults for general serving.
MXFP4 benchmark reproduction (InferenceX MI355X sweep)
The command above is the default serving configuration. The SemiAnalysis InferenceX MI355X benchmark lane for this checkpoint is a separate high-concurrency benchmark configuration — not the recipe default. To reproduce that sweep, start from the default command above and apply these deltas:
- Use
--tensor-parallel-size 4(the sweep runs TP4; the default recipe keeps TP8 for KV-cache and multimodal-encoder headroom). - Add
--kv-cache-dtype fp8. - Use the benchmark block size and scheduler knobs:
--block-size 16,--max-num-batched-tokens 16384,--max-num-seqs 512,--async-scheduling, and--no-enable-prefix-caching. - Run on a current
vllm/vllm-openai-rocm:nightlythat contains the required AITER Kimi MXFP4 backend. - Export the AITER and INT4 quantized all-reduce env vars before launch:
# Kernel selection and benchmark runtime knobs.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
export VLLM_ROCM_USE_SKINNY_GEMM=0
export VLLM_ROCM_USE_AITER_RMSNORM=0
export AITER_MXFP4_INTERMEDIATE=1
export AITER_BYPASS_TUNE_CONFIG=0
export AITER_MOE_SORT_BACKEND=auto
export OMP_NUM_THREADS=1
The full reproduction serve command is then:
vllm serve amd/Kimi-K2.5-MXFP4 \
--port "${PORT:-8000}" \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.85 \
--max-model-len "$MAX_MODEL_LEN" \
--kv-cache-dtype fp8 \
--block-size 16 \
--max-num-batched-tokens 16384 \
--max-num-seqs 512 \
--async-scheduling \
--trust-remote-code \
--no-enable-prefix-caching \
--mm-encoder-tp-mode data
The benchmark script enforces MAX_MODEL_LEN >= 9472 so the 1k/1k
configuration matches the locally validated server. The env vars and scheduler
knobs are benchmark-specific throughput settings, not default serving choices.
Tuned AITER MXFP4 MoE (MI355X)
For higher MoE throughput on MI355X, run the AITER MXFP4 (RadeonFlow)
intermediate GEMM path with fused shared experts. This requires a
vllm/vllm-openai-rocm:nightly built after vLLM
#48683 (AITER v0.1.16.post5,
which includes the ROCm/aiter#3832
gfx950 MXFP4 MoE backend). Export the AITER MoE kernel-selection vars and add the
batching/scheduling flags below. AITER_MXFP4_INTERMEDIATE=1 is what selects the
RadeonFlow MXFP4 intermediate path — without it (and VLLM_ROCM_USE_AITER=1) the
MoE oracle falls back to a different kernel. Runs at TP4 or TP8; for TP < 8
disable AITER RMSNorm.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
export VLLM_ROCM_USE_SKINNY_GEMM=0
export AITER_MXFP4_INTERMEDIATE=1
export AITER_BYPASS_TUNE_CONFIG=0
export AITER_MOE_SORT_BACKEND=auto
export OMP_NUM_THREADS=1
# TP < 8 only:
export VLLM_ROCM_USE_AITER_RMSNORM=0
vllm serve amd/Kimi-K2.5-MXFP4 \
--tensor-parallel-size 4 \
--trust-remote-code \
--gpu-memory-utilization 0.90 \
--kv-cache-dtype fp8 \
--block-size 16 \
--async-scheduling \
--no-enable-prefix-caching \
--mm-encoder-tp-mode data
These AITER MXFP4 MoE env vars are kernel-selection choices (same MXFP4 math, faster kernel), not precision reductions.
For fixed-length throughput sweeping, additionally cap --max-model-len to the
working sequence length (e.g. 9472 for an 8k/1k workload) and set
--max-num-batched-tokens 16384 --max-num-seqs 512. Leave these off for general
serving.
Client Usage
Once the vLLM server is running, consume it via the OpenAI-compatible API:
import time
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:8000/v1",
timeout=3600
)
messages = [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://ofasys-multimodal-wlcb-3-toshanghai.oss-accelerate.aliyuncs.com/wpf272043/keepme/image/receipt.png"
}
},
{
"type": "text",
"text": "Read all the text in the image."
}
]
}
]
start = time.time()
response = client.chat.completions.create(
model="moonshotai/Kimi-K2.5",
messages=messages,
max_tokens=2048
)
print(f"Response costs: {time.time() - start:.2f}s")
print(f"Generated text: {response.choices[0].message.content}")
Troubleshooting
- OOM errors: Lower
--gpu-memory-utilizationor adjust TP/EP to match your GPU count. - Vision encoder performance: Use
--mm-encoder-tp-mode datato run the vision encoder in data-parallel mode. The encoder is small, so TP adds communication overhead with little gain. - Unique multimodal inputs: Pass
--mm-processor-cache-gb 0to avoid caching overhead. For repeated inputs,--mm-processor-cache-type shmuses host shared memory for better performance at high TP settings. - MoE kernel tuning: Use the
benchmark_moescript from vLLM to tune Triton kernels for your specific hardware. - Async scheduling: Enabled by default for better throughput. Disable if you encounter issues, and file a bug report to vLLM.