Deep Dive180min
Solving Multimodal and Long-Context Challenges
This deep-dive explores extending LLMs with multimodal and long-context capabilities to solve real-world tasks. It presents step-by-step Python notebook solutions for video transcription with speaker identification, large-scale knowledge graph generation, image automation, and visual object detection, restoration, and editing.
talk.summaryAiDisclaimer
Laurent PicardGoogle Cloud
talkDetail.whenAndWhere
Monday, October 5, 09:30-12:30
TBA 4
talks.roomOccupancytalks.noOccupancyInfo
Large Language Models have radically transformed how we solve problems, but how far can we go with real-world solutions?
In this deep-dive, we will enhance our problem-solving toolkit by integrating multimodal and long-context capabilities, and unlock an unprecedented range of new solutions. Step by step, we will architect solutions for four complex challenges:
In this deep-dive, we will enhance our problem-solving toolkit by integrating multimodal and long-context capabilities, and unlock an unprecedented range of new solutions. Step by step, we will architect solutions for four complex challenges:
- Multimodal Video Transcription: Transcribe videos and identify speakers in a single prompt.
- Knowledge Graph Generation: Extract entities and relationships from massive inputs (1M tokens) with a single request.
- Image Bank Automation: Set up a pipeline for consistent image generation and bring your visual archives back to life.
- Visual Object Detection & Edition: Identify, restore, and transform elements within your images.
Laurent Picard
Laurent is a software engineer passionate about everything shaping the future. In a previous life, he co-founded Bookeen and spent 17 years pioneering the digital book industry. Today, as a Developer Advocate at Google Paris, he explores the frontiers of cloud technology and enjoys sharing a world of possibilities with the community.