Agentic Engineering & ToolingAgentic Engineering & Tooling
Deep Dive180min
BEGINNER

Solving Multimodal and Long-Context Challenges

This deep-dive explores extending LLMs with multimodal and long-context capabilities to solve real-world tasks. It presents step-by-step Python notebook solutions for video transcription with speaker identification, large-scale knowledge graph generation, image automation, and visual object detection, restoration, and editing.

talk.summaryAiDisclaimer

Laurent Picard
Laurent PicardGoogle Cloud

talkDetail.whenAndWhere

Monday, October 5, 09:30-12:30
TBA 2
talks.roomOccupancytalks.noOccupancyInfo
talks.description
Large Language Models have radically transformed how we solve problems, but how far can we go with real-world solutions?

In this deep-dive, we will enhance our problem-solving toolkit by integrating multimodal and long-context capabilities, and unlock an unprecedented range of new solutions. Step by step, we will architect solutions for four complex challenges:

  1. Multimodal Video Transcription: Transcribe videos and identify speakers in a single prompt.
  2. Knowledge Graph Generation: Extract entities and relationships from massive inputs (1M tokens) with a single request.
  3. Image Bank Automation: Set up a pipeline for consistent image generation and bring your visual archives back to life.
  4. Visual Object Detection & Edition: Identify, restore, and transform elements within your images.
long-context
knowledge graphs
large language models
multimodal
talks.speakers
Laurent Picard

Laurent Picard

Google Cloud

France

Laurent is a software engineer passionate about everything shaping the future. In a previous life, he co-founded Bookeen and spent 17 years pioneering the digital book industry. Today, as a Developer Advocate at Google Paris, he explores the frontiers of cloud technology and enjoys sharing a world of possibilities with the community.