Featured
Compare webshops (1)
€ 15.62
Shops with other sizes
Pages: 138, Paperback, Independently published
Independently Published
GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization...
VLLM Quickstart Guide of HOS: High Performance LLM Inference for Production
LLM Inference Engineering: Quantization, KV-Cache Optimization, and High-Throughput Serving: A Production Engineer's...
vLLM Deployment Blueprint: Deploy, Optimize, and Scale High Performance LLM Inference Systems
Back to top