Skip to content
xLLM
Getting Started
Hardware
User Guide
CookBook
Developer Guide
CLI Reference
Search
Ctrl
K
Cancel
GitHub
EN
EN
中
Home
Getting Started
Quick Start
Launch xllm
Multi-Machine Deployment
Online Service
Offline Inference
Supported Models
Hardware
Hardware Platforms
NVIDIA GPU
Ascend NPU
Cambricon MLU
Hygon DCU
MetaX MACA
Mthreads MUSA
User Guide
Advanced Features
Async schedule
Multi-stream parallel
ChunkedPrefill Scheduler
Zero Evict Scheduler
Disaggregated PD
Prefix Cache Optimization
Global Multi-Level KV Cache
Multimodal Support
EP Parallelism
MoE Load Balancing (EPLB)
MTP Speculative Inference
Graph Mode
FlashComm
xLLM Service
CookBook
Autoregressive Models
Qwen
Qwen3.5
Qwen3
Qwen3-Next
Qwen3-VL
Qwen2.5-VL
DeepSeek
DeepSeek-V4
DeepSeek-V3.2
DeepSeek-V3.1
DeepSeek-V3
DeepSeek-R1
GLM
GLM-5.3-Flash
GLM-5.2
GLM-5.1
GLM-5
GLM-4.7
GLM-4.7-Flash
GLM-4.6
GLM-4.6V
GLM-4.5
GLM-4.5V
Kimi
Kimi2
Kimi-K2.5 / Kimi-K2.6
MinMax
MiniMax-M2.7
Diffusion Models
Flux
Flux
Flux2
Wan
Wan2.1
Wan2.2
Qwen-Image
Qwen-Image
Developer Guide
Development
Code Architecture
xLLM Ascend TileLang Kernel Development Guide
Online Profiling
AI Coding Workflow
Auto Config Development Guide
Design
Graph Mode Design Document
Generative Recommendation Design Document
C++ Serving Framework + Python Model Execution Architecture Decision
CLI Reference
GitHub
Select language
EN
中
GLM-5.2
Copy page
GLM-5.2 inference recipes are consolidated into
GLM-5 / GLM-5.1 / GLM-5.2
.