Keynote20min
Half of a billion Idle GPUs: The Case for On-Device AI
This keynote argues that Apple Silicon devices form a massive, underused AI compute fleet. Drawing on the MLX ecosystem, it shows how on-device inference can handle frontier models through advanced optimization, and why privacy, latency, cost, and sovereignty will shift more AI workloads from cloud to device.
talk.summaryAiDisclaimer
Prince CanumaNeywa Labs
talkDetail.whenAndWhere
Wednesday, October 7, 10:10-10:30
TBA 7
talks.roomOccupancytalks.noOccupancyInfo
The most powerful compute fleet ever assembled wasn't built by a cloud provider. It's the half a billion-plus Apple Silicon devices already in laps, pockets, and offices — each with a GPU, a neural engine, and unified memory that would have qualified as a workstation five years ago. Collectively they represent more silicon than any datacenter on Earth. And for AI, they sit almost entirely idle.
This keynote makes the case that the next era of AI won't be defined by bigger clusters, but by inference moving to where the data already lives. I'll tell that story through building the MLX open-source ecosystem (mlx-vlm, mlx-audio, nativ — 5M+ downloads, used by Google DeepMind, Hugging Face, and LM Studio): what it actually takes to make consumer hardware run frontier models, why the server playbook fails on-device, and the hard-won engineering — quantization that holds accuracy, KV-cache compression, speculative decoding across chip generations — that closes the gap.
You'll leave with a new mental model for where AI workloads belong: which belong in the cloud, which belong on-device, and why the answer is about to flip for more of them than you think, for reasons of privacy, latency, cost, and sovereignty.
For every engineer and leader deciding where their AI features will live.
This keynote makes the case that the next era of AI won't be defined by bigger clusters, but by inference moving to where the data already lives. I'll tell that story through building the MLX open-source ecosystem (mlx-vlm, mlx-audio, nativ — 5M+ downloads, used by Google DeepMind, Hugging Face, and LM Studio): what it actually takes to make consumer hardware run frontier models, why the server playbook fails on-device, and the hard-won engineering — quantization that holds accuracy, KV-cache compression, speculative decoding across chip generations — that closes the gap.
You'll leave with a new mental model for where AI workloads belong: which belong in the cloud, which belong on-device, and why the answer is about to flip for more of them than you think, for reasons of privacy, latency, cost, and sovereignty.
For every engineer and leader deciding where their AI features will live.
Prince Canuma
Prince Canuma is Founder and CEO of Neywa Labs, building the inference stack for Apple Silicon. He is the creator and primary maintainer of the MLX open-source ecosystem — mlx-vlm, mlx-audio, and mlx-embeddings — with over 4M combined downloads, foundational to on-device AI on Apple hardware. He previously led developer experience at Neptune AI and worked at Arcee on post-training and inference optimization. His work focuses on model optimization and making AI accessible on local hardware.