# Introduction to LLM Inference Engines

> Inference engines are the systems that turn a trained large language model into a usable service. They sit between the model weights and the application, handling token generation, memory…

- **الرابط:** https://zomra.io/ar/free-sessions/introduction-to-llm-inference-engines
- **النوع:** جلسة مجانية
- **آخر تحديث:** 3 أكتوبر 2026
- **تبدأ:** 8 أكتوبر 2026 في 5:00 م
- **المدة:** 90 دقيقة
- **يقدّمها:** Aya Nasser Salama
- **الجهة المستضيفة:** MLOps MENA Community
- **عدد المسجّلين:** 139
- **السعر:** مجاني

## لماذا تحضر

Inference engines are the systems that turn a trained large language model into a usable service. They sit between the model weights and the application, handling token generation, memory allocation, batching, scheduling, caching, quantization, and hardware execution. Their job is not to make the model smarter, but to make it faster, cheaper, more scalable, and more reliable when serving real users.

In this session, we will look at how modern inference engines such as vLLM, TensorRT-LLM, SGLang, llama.cpp, and others work under the hood. We will compare their architectures, explore concepts such as KV cache management, continuous batching, speculative decoding, tensor parallelism, and quantization, and learn how to choose and configure an engine for different deployment scenarios, from a single local GPU to high-throughput production clusters.

## ما الذي ستتعلمه

- Understand where an inference engine sits between model weights and the application
- Compare the architectures of vLLM, TensorRT-LLM, SGLang, llama.cpp and others
- Learn how KV cache management, continuous batching, speculative decoding, tensor parallelism and quantization work
- Choose and configure the right engine, from a single local GPU to high-throughput production clusters

## المتحدثون

### Mohamed Rashad

Co-Founder and CTO

- https://www.linkedin.com/in/rashaddism

## التسجيل

سجّل عبر https://zomra.io/ar/free-sessions/introduction-to-llm-inference-engines
