This session is for members.

Subscribe or log in to watch every GoSec session.

Subscribe Log in

This recording is not available yet.

Improving Data Classification in the Modern Age

Download resources

About this session

John Loya, VP of Sales Engineering at Cyberhaven, argues that data classification keeps failing because it leans on content alone, and walks through how AI and data lineage change that. He surveys traditional content inspection (keywords, regex, exact data matching, OCR) and shows how easily each is defeated: obfuscated social security numbers, cursive or rotated text that beats OCR, password-protected archives no scanner can open. He contrasts this with user-applied labels, Microsoft's preferred approach, which fail because employees mislabel files or skip formats that cannot carry a label, and with contextual signals like file location or author, which are weak alone. His core pitch is lineage: logging every copy, move, upload, download and paste across endpoints, browsers and cloud connectors to build a history for each file, so classification can rely on where data originated rather than pattern-matching. He describes Cyberhaven's lineage model, which predicts unusual next actions, auto-summarizes incidents for analysts and scores severity to cut false positives. Audience questions cover on-premises content inspection, where graph-database data is hosted, and event volume at scale.

Legacy data classification methods have long struggled with accuracy and scalability, often plagued by high rates of false positives and false negatives. The landscape is shifting with the rise of Artificial Intelligence and greater contextual visibility into how data is created, accessed, and used, organizations now have powerful new tools to dramatically improve classification outcomes. This session explores the next generation of data classification and how modern approaches are overcoming legacy limitations and enabling more precise, automated, and scalable data protectio

Key takeaways

  • Do not rely on content inspection alone: regex and OCR are easily defeated with obfuscated characters, unusual fonts or rotated images.
  • Treat user-applied sensitivity labels (Microsoft-style) as a starting point, not a control: unsupported file types and mislabeling leave real gaps.
  • Log every file event (copy, move, upload, download, paste) across endpoint, browser and cloud connectors to build lineage, not just point-in-time scans.
  • Use lineage context (e.g., personal vs. corporate Gmail, source application) to cut false positives before adding content inspection on top.
  • When exporting to a SIEM or SOAR, scope the feed to incidents rather than every logged event unless you truly need full volume.

Speakers

John Loya
John Loya
Vice President of Sales Engineering · Cyberhaven
John Loya is the Vice President of Sales Engineering at Cyberhaven. He has previously held roles at Microsoft, McAfee, and Digital Guardian. In his current role he assists customers and prospects globally with their data protection needs in terms of… Read moreRead less

John Loya is the Vice President of Sales Engineering at Cyberhaven. He has previously held roles at Microsoft, McAfee, and Digital Guardian. In his current role he assists customers and prospects globally with their data protection needs in terms of compliance, governance, privacy, classification, and security. He has worked previously as a developer, quality assurance engineer, and an automation engineer. He has been in the security space for over 20 years with a strong focus on Data Loss Prevention and Insider Risk Management.

Resources

Tags

More from GoSec 2025

Also from John Loya

On the same topic

This site is registered on wpml.org as a development site. Switch to a production site key to remove this banner.