Lunch Talk40min
Measuring the Unmeasurable: Evaluating Generative AI in Practice
This session explores how to evaluate generative AI systems reliably despite their unpredictability. Using a RAG chatbot case study, it covers adapting metrics for accuracy, coherence, and business value; the limits of traditional and LLM-as-a-judge methods; the role of human review; and continuous monitoring for drift.
talk.summaryAiDisclaimer
Erin PacquetetSCIAM
talkDetail.whenAndWhere
Tuesday, October 6, 12:40-13:20
TBA 5
talks.roomOccupancytalks.noOccupancyInfo
Getting a generative AI prototype to impress a demo audience is easy. Getting it to behave reliably in production is another story. In production, you don't control what users ask or what the model answers. That's not a bug, that's the whole point of using LLMs. But it is also what makes evaluation genuinely hard, and what separates a convincing proof of concept from a trustworthy production system.
This session tackles the core paradox of building with large language models: how do you harness their creativity while maintaining control over their outputs? We will use a RAG-based chatbot as our running example, a real-world case where you have to guarantee output quality without ever controlling what users actually ask, and where linguistic fluency and factual accuracy constantly pull in opposite directions, to show what a complete, reproducible evaluation pipeline actually looks like in practice.
Along the way, we will cover why traditional metrics fall short, how to use LLMs as judges and where that approach quietly fails, when human evaluation is non-negotiable, and how to set up continuous monitoring to catch drift before your users do.
You will leave with practical frameworks and a clear methodology to make evaluation a constant driver of development, from the first prototype to the last production release.
This session tackles the core paradox of building with large language models: how do you harness their creativity while maintaining control over their outputs? We will use a RAG-based chatbot as our running example, a real-world case where you have to guarantee output quality without ever controlling what users actually ask, and where linguistic fluency and factual accuracy constantly pull in opposite directions, to show what a complete, reproducible evaluation pipeline actually looks like in practice.
Along the way, we will cover why traditional metrics fall short, how to use LLMs as judges and where that approach quietly fails, when human evaluation is non-negotiable, and how to set up continuous monitoring to catch drift before your users do.
You will leave with practical frameworks and a clear methodology to make evaluation a constant driver of development, from the first prototype to the last production release.
Erin Pacquetet
Erin Pacquetet is a Senior AI and Data Scientist at SCIAM, where she designs, develops, and deploys production-grade AI products.
At the intersection of linguistics and computer science, she brings a creative, human-centered approach to hard engineering problems, translating product challenges into robust, scalable AI architectures that work for both technical teams and business stakeholders.
Her expertise spans prompt design and optimization, advanced data pipeline management, model evaluation, and the integration of industrial-grade AI systems.
Beyond her industry work, Erin is an active researcher and community contributor. She regularly speaks at meetups and conferences on topics including AI model analysis, pragmatic AI integration in software development, and human–machine interaction through language.
At the intersection of linguistics and computer science, she brings a creative, human-centered approach to hard engineering problems, translating product challenges into robust, scalable AI architectures that work for both technical teams and business stakeholders.
Her expertise spans prompt design and optimization, advanced data pipeline management, model evaluation, and the integration of industrial-grade AI systems.
Beyond her industry work, Erin is an active researcher and community contributor. She regularly speaks at meetups and conferences on topics including AI model analysis, pragmatic AI integration in software development, and human–machine interaction through language.