Data & AIData & AI
Conference45min
INTERMEDIATE

Optimizing LLM Inference for the Rest of Us

This session presents practical methods to optimize Kubernetes clusters for cost‑efficient, scalable LLM inference. It covers container and model tuning, accelerator management, data and network optimization, and observability—equipping attendees with strategies to balance performance, availability, and cost when deploying AI applications without hyperscale infrastructure.

talk.summaryAiDisclaimer

Abdel Sghiouar
Abdel SghiouarGoogle Cloud

talkDetail.whenAndWhere

Wednesday, April 29, 16:30-17:15
CINEMA
talks.description
Not every organization operates with the hyperscale resources of Anthropic, Google, or OpenAI. For the majority of businesses integrating Large Language Models (LLMs) into their critical paths, the high costs and scarcity of GPU/TPU accelerators present a significant challenge. Striking the balance between performance, availability, scalability, and cost-efficiency is a must.

While Kubernetes is a ubiquitous runtime for modern workloads, deploying LLM inference effectively demands a specialized approach. This session dives deep into practical strategies for optimizing your Kubernetes clusters and LLM Inference workloads to run efficiently and cost effectively. We will explore:

- Container and Model Optimization
- Accelerator Management
- Data & Storage
- Network & Load Balancing
- Observability

Attendees will leave with practical techniques for maximizing cost/performance for LLM inference for their AI-powered applications on Kubernetes.
kubernetes
scalability
optimization
inference
talks.speakers
Abdel Sghiouar

Abdel Sghiouar

Google Cloud

Sweden

Abdel Sghiouar is a senior Developer @Google Cloud. A co-host of the Kubernetes Podcast by Google and a CNCF Ambassador. His focused areas are High Scale distributed systems on Kubernetes, Service Mesh, and Serverless. With a background in datacenter scale architecture and operations and consulting. He spends most of his time working on optimizing GenAI Apps for large scale operations using Cloud Native Technologies and producing content targeting developers and ops professionals.