Agentic Engineering & ToolingAgentic Engineering & Tooling
Conference50min
BEGINNER

Owning the inference layer: When and how to run your own models

This talk compares hosted model APIs and self-hosted inference for GenAI, covering tradeoffs in cost, latency, flexibility, and data control. It highlights open tools like vLLM and llm-d, and offers a practical framework for deciding when self-hosting is worth the operational complexity.

talk.summaryAiDisclaimer

Maarten Vandeperre
Maarten VandeperreRed Hat
Camille Nigon
Camille NigonRed Hat

talkDetail.whenAndWhere

Thursday, October 8, 16:30-17:20
TBA 2
talks.roomOccupancytalks.noOccupancyInfo
talks.description
Hosted model APIs are the fastest way to start building with Gen AI models, but they are not always the best long-term fit. As workloads grow, teams start running into harder questions around latency, cost, deployment flexibility, data boundaries, and performance tuning.

This talk looks at what changes when you move from calling a hosted model API to running inference yourself. We’ll break down the practical tradeoffs between hosted and self-hosted approaches, then examine how modern open inference technologies such as vLLM and llm-d are making self-hosted AI more realistic for production systems.

Rather than treating this as a debate between two camps, the session focuses on decision-making: when self-hosting is worth the added operational complexity, when hosted APIs still win, and what teams take on when they choose to own the inference layer.

We'll also unpack what sovereignty is actually worth beyond the compliance checkbox: the concrete value of controlling your data boundaries, your model lifecycle, and your independence from a single provider's roadmap and pricing, and where that value is real rather than theoretical.

You’ll leave with a practical framework for evaluating cost, control, performance, and architectural fit in your own environment.


sovereignty
latency
inference
self-hosting
talks.speakers
Maarten Vandeperre

Maarten Vandeperre

Red Hat

Belgium

Maarten Vandeperre is an experienced software professional who recently joined Red Hat as an Appdev & AI Specialized Solutions Architect. With a strong background in software development and architecture, he brings a wealth of expertise to his role. Maarten's primary focus is on application development and AI, with a particular emphasis on leveraging Red Hat's OpenShift platform from a developer's perspective.

One of Maarten's true passions lies in advocating for "clean architecture" as a guiding principle in software development. He firmly believes in the importance of designing software systems that are modular, maintainable, and scalable. As part of his dedication to this approach, Maarten strives to map these principles to infrastructure solutions, ensuring that the underlying technology supports and enhances the overall architecture. His deep understanding of integration technologies, such as API Gateways, Keycloak, Kafka, service mesh, and Camel, enables him to create seamless connections between systems while adhering to clean architectural principles, which empower organizations to thrive in the ever-evolving digital landscape.
Camille Nigon

Camille Nigon

Red Hat

Switzerland

Camille is a Solutions Architect at Red Hat with a background in AI/ML and a keen interest in open-source projects. Passionate about knowledge sharing, Camille engages in technical workshops, public speaking, and hands-on projects to help others navigate the evolving AI landscape.