Conference50min
Build a private cluster for the office
A technical talk on building private, local AI systems for regulated organizations using open-weight models. It covers hardware choices, networking, model selection, private data access, and fast document retrieval, helping small teams reduce dependence on frontier model providers while ensuring privacy, governance, and long-term support.
talk.summaryAiDisclaimer
John DaviesIncept5
talkDetail.whenAndWhere
Thursday, October 8, 15:00-15:50
TBA 3
talks.roomOccupancytalks.noOccupancyInfo
For almost three years now John has been building private (local, shared server, private data room and private cloud) systems for his customers, mostly financial institutions, large banks, but also legal, medical and government work. Most of these organisations simply can't use the frontier models, even if hosted in the EU. It's not just cost, it's the need for long-term support, SLAs, total privacy, ownership, accountability, governance, repeatability, the list goes on... There's Bedrock and Vertex from Amazon and Google, and companies like RunPod with EU instances, but renting 24/7 costs roughly half the hardware's purchase price every month.
For an office of about 10–50 people, not coders, a small redundant cluster of Mac Minis or Studios, NVIDIA DGX Sparks or even both are perfect for the job, dedicated or shared. John will start with setting up the machines, what needs networking and what doesn't, which models work best, adding tools and access to local data feeds, even private ones. He'll also cover KV-cache loading from SSD for near-instant document access to replace RAG. These models can code but that's not the main goal here, it's for private internal and customer data.
This will be a technical talk but you will walk away with enough knowledge to start replacing your team's dependence on OpenAI, Anthropic and other frontier models. Open-weight models all the way.
For an office of about 10–50 people, not coders, a small redundant cluster of Mac Minis or Studios, NVIDIA DGX Sparks or even both are perfect for the job, dedicated or shared. John will start with setting up the machines, what needs networking and what doesn't, which models work best, adding tools and access to local data feeds, even private ones. He'll also cover KV-cache loading from SSD for near-instant document access to replace RAG. These models can code but that's not the main goal here, it's for private internal and customer data.
This will be a technical talk but you will walk away with enough knowledge to start replacing your team's dependence on OpenAI, Anthropic and other frontier models. Open-weight models all the way.
John Davies
After a degree in Astrophysics at UCL John started in hardware then assembler, C, C++ and later Java. Almost exclusively in finance he ran FX at Paribas was a global chief architect at JP Morgan, BNP Paribas and VISA. John has co-founded four successful startups since 2000, selling one of them twice to Nasdaq & LSE listed companies. After co-founding Velo Payments with the former president of VISA John has spun off the AI company Incept5. John has co-authored several Java books and is a frequent speaker at technical and banking conferences around the world. He is married to a French wife and has three boys in their 20s.