Security, Trust & ComplianceSecurity, Trust & Compliance
Conference50min
INTERMEDIATE

Underneath the Black Marker: AI Techniques for Detecting and Reversing Redactions

This session demonstrates how redactions fail in PDFs and images, and how AI can recover hidden text or visual content. It also explains proper redaction practices, showing that true privacy requires removing data, not just visually obscuring it.

talk.summaryAiDisclaimer

David vonThenen
David vonThenenNetApp

talkDetail.whenAndWhere

Thursday, October 8, 15:00-15:50
TBA 9
talks.roomOccupancytalks.noOccupancyInfo
talks.description
Redaction is often treated like a final step. Add black boxes, blur a name, pixelate a screenshot, export the file, done. But recent headlines have reminded everyone that "hidden" is not the same as "gone." In practice, many redactions fail in predictable ways. Sometimes the underlying text remains in the PDF. Sometimes image-based redactions leave enough signal behind to detect what was covered. And sometimes the surrounding sentence context gives away exactly what was meant to stay private. Recent media coverage around heavily redacted files has exposed a hard truth: concealment is not the same thing as data removal.

This session is a live, technical walk-through of how modern redaction failures happen and how AI can expose them. We will demo detection methods for finding redacted text in documents, show how poorly redacted PDFs can be recovered, and end with examples of recovering information from redacted or pixelated images. More importantly, we will also cover what proper redaction looks like in practice. Along the way, we will connect document analysis, computer vision, and language models to a simple point: privacy controls fail when teams confuse visual hiding with real removal. Bring your curiosity and maybe a little paranoia.
computervision
pdf
privacy
redaction
talks.speakers
David vonThenen

David vonThenen

NetApp

United States of America

David vonThenen is an AI/ML Engineer where he focuses on production AI systems, enterprise AI strategy, and architectures for explainable, governable, and reliable generative AI. His work spans Agentic AI, Graph RAG, Document RAG, AI memory systems, multi-agent architectures, OpenSearch, Neo4j, data lineage, provenance, and AI governance.

David has more than 20 years of experience building production software across AI/ML, speech and NLP, Kubernetes, cloud-native platforms, storage, virtualization, and backup/recovery. He combines hands-on engineering with technical strategy, open-source leadership, developer advocacy, and customer-facing architecture work.