Understand where an inference engine sits between model weights and the application
Compare the architectures of vLLM, TensorRT-LLM, SGLang, llama.cpp and others
Learn how KV cache management, continuous batching, speculative decoding, tensor parallelism and quantization work
Choose and configure the right engine, from a single local GPU to high-throughput production clusters
Inference engines are the systems that turn a trained large language model into a usable service. They sit between the model weights and the application, handling token generation, memory allocation, batching, scheduling, caching, quantization, and hardware execution. Their job is not to make the model smarter, but to make it faster, cheaper, more scalable, and more reliable when serving real users.
In this session, we will look at how modern inference engines such as vLLM, TensorRT-LLM, SGLang, llama.cpp, and others work under the hood. We will compare their architectures, explore concepts such as KV cache management, continuous batching, speculative decoding, tensor parallelism, and quantization, and learn how to choose and configure an engine for different deployment scenarios, from a single local GPU to high-throughput production clusters.
I'm a Senior MLOps and LLMOps Engineer with 6+ years in AI, holding a master's in Informatics from Nile University. I've worked at Unifonic, Valeo, Aiactive Technologies, and Advanced Programs Co. I'm an MLOps instructor at ITI Cairo and founder of the MLOps MENA Community.