Conference50min
Underneath the Black Marker: AI Techniques for Detecting and Reversing Redactions
This session demonstrates how redactions fail in PDFs and images, and how AI can recover hidden text or visual content. It also explains proper redaction practices, showing that true privacy requires removing data, not just visually obscuring it.
talk.summaryAiDisclaimer
David vonThenenNetApp
talkDetail.whenAndWhere
Thursday, October 8, 15:00-15:50
TBA 9
talks.roomOccupancytalks.noOccupancyInfo
Redaction is often treated like a final step. Add black boxes, blur a name, pixelate a screenshot, export the file, done. But recent headlines have reminded everyone that "hidden" is not the same as "gone." In practice, many redactions fail in predictable ways. Sometimes the underlying text remains in the PDF. Sometimes image-based redactions leave enough signal behind to detect what was covered. And sometimes the surrounding sentence context gives away exactly what was meant to stay private. Recent media coverage around heavily redacted files has exposed a hard truth: concealment is not the same thing as data removal.
This session is a live, technical walk-through of how modern redaction failures happen and how AI can expose them. We will demo detection methods for finding redacted text in documents, show how poorly redacted PDFs can be recovered, and end with examples of recovering information from redacted or pixelated images. More importantly, we will also cover what proper redaction looks like in practice. Along the way, we will connect document analysis, computer vision, and language models to a simple point: privacy controls fail when teams confuse visual hiding with real removal. Bring your curiosity and maybe a little paranoia.
This session is a live, technical walk-through of how modern redaction failures happen and how AI can expose them. We will demo detection methods for finding redacted text in documents, show how poorly redacted PDFs can be recovered, and end with examples of recovering information from redacted or pixelated images. More importantly, we will also cover what proper redaction looks like in practice. Along the way, we will connect document analysis, computer vision, and language models to a simple point: privacy controls fail when teams confuse visual hiding with real removal. Bring your curiosity and maybe a little paranoia.
David vonThenen
David vonThenen is an AI/ML Engineer where he focuses on production AI systems, enterprise AI strategy, and architectures for explainable, governable, and reliable generative AI. His work spans Agentic AI, Graph RAG, Document RAG, AI memory systems, multi-agent architectures, OpenSearch, Neo4j, data lineage, provenance, and AI governance.
David has more than 20 years of experience building production software across AI/ML, speech and NLP, Kubernetes, cloud-native platforms, storage, virtualization, and backup/recovery. He combines hands-on engineering with technical strategy, open-source leadership, developer advocacy, and customer-facing architecture work.
David has more than 20 years of experience building production software across AI/ML, speech and NLP, Kubernetes, cloud-native platforms, storage, virtualization, and backup/recovery. He combines hands-on engineering with technical strategy, open-source leadership, developer advocacy, and customer-facing architecture work.