Partitioning and scheduling a fixed pool of GPUs across training, inference, and experimentation
Managing CUDA/driver versioning without breaking running jobs
Monitoring utilization to catch idle or contended hardware
Designing around power, cooling, and bandwidth — the constraints the cloud hides from you
When on-prem isn’t optional: data residency, air-gapped environments, regulation
🔹 In our region, on-prem isn't the exception — it's the default. Banks, telecoms, government entities, and healthcare organizations across MENA operate under data residency and regulatory requirements that make on-prem mandatory. This is a skill the local job market actually needs.
🔹 Most MLOps content assumes the cloud. Every tutorial you'll find takes elastic resources and managed services for granted. The moment you move on-prem, most of those best practices break — and the material that addresses this is rare and scattered.
🔹 This is a playbook from real production, not theory. Every point in the session comes from actual deployments in regulated industries, where every inference had to run on hardware the organization controlled directly. You'll hear the decisions that were made and the trade-offs that came with them — not generic slides.
🔹 You'll leave with practical answers to problems you actually face: How do you partition a fixed set of GPUs across training, inference, and experimentation? How do you upgrade CUDA without breaking running jobs? How do you detect hardware sitting idle or contended?
🔹 A chance to ask a CTO directly. Mohamed Rashad has been working with on-prem LLMs since they first appeared, and his work deploying Llama-2 models on-prem was recognized by Meta's Engineering Blog. The Q&A alone is worth attending for.
👥 This session is for you if you're:
An MLOps or DevOps engineer moving into AI workloads • An ML engineer who wants to understand what happens beneath the infrastructure • A tech lead deciding between cloud and on-prem • Working at an organization whose data can't leave the premises
I'm a Senior MLOps and LLMOps Engineer with 6+ years in AI, holding a master's in Informatics from Nile University. I've worked at Unifonic, Valeo, Aiactive Technologies, and Advanced Programs Co. I'm an MLOps instructor at ITI Cairo and founder of the MLOps MENA Community.