This session is for members.

Subscribe or log in to watch every GoSec session.

Subscribe Log in

This recording is not available yet.

A Dumb AI Talk

Download resources

About this session

Martin Holste, Trellix's chief product architect, delivers a live technical talk (rendered as HTML rather than slides) arguing that small, on-device 'dumb' language models, not frontier models, are what make AI SOC investigation trustworthy at scale. He opens with the 2023 Hugging Face incident, where alert volume overwhelmed human triage but cloud frontier models could not be used because malicious code embedded in prompts had to stay air-gapped, forcing small, private models instead. He walks through Trellix's confidence-scoring pipeline: a 'road to 70 percent' threshold where an investigation earns analyst attention, a 90 percent threshold for automated remediation such as auto-quarantining a device, and a stack-ranking exercise (borrowed from baseball's wins-above-replacement statistic) that measures how much confidence each data source, threat intel, SOPs, red-team output, EDR, actually adds. He explains why small models over-weight alarming words like 'threat' and describes the fix: separating extraction, verification and classification into different agents, and computing the final score outside the model in Python so the LLM cannot simply assert a number. He closes on why small models struggle with MCP tool orchestration and need tiered, need-to-know tool disclosure, then takes live audience questions.

Yeah, yeah, AI SOC, we get it. But what really goes into getting true wins out of using AI in security vs saving a couple of clicks? This talk will show the importance of small (dumb) LLM's and how they create smart opportunities for secops by allowing you to use more tokens, anywhere you want. It will demo how you can go beyond enriching a few alerts to stack rank your most valuable security tools in an ROI thunderdome. You'll learn the AI confidence guardrail framework you need so you can let it actually start taking actions on your behalf without accidentally deleting all your VM's when you asked it to 'eliminate all vulnerabilities.' Seeing is believing, so there will be real-world demos of where AI works, where it doesn't, and why--bring your 3D glasses.

Key takeaways

  • Do not let an AI model self-report its own confidence; the speaker found frontier models converge on a fixed number (around 85%) regardless of accuracy, so compute the score outside the model from verified extractions instead.
  • Set two separate thresholds: one where an AI SOC finding earns an analyst's attention (around 70% in Trellix's framework) and a stricter one before you allow automated remediation (around 90%), because the two decisions carry very different risk.
  • Before auto-remediating (e.g. quarantining a device), require confidence in the asset type as well as the event, not just that something bad happened; confusing a laptop for a server is the kind of mistake that makes automation dangerous.
  • Stack-rank your security data sources by how much they actually add to investigation confidence, not by cost or reputation alone; in this talk, threat intel and EDR carried outsized weight while some custom-written detection rules were the noisiest.
  • When deploying small, on-device language models for tool use (MCP), load only the tools relevant to the current step in a tiered, need-to-know structure; dumping every tool description into context overwhelms a small model's context window.

Speakers

Martin Holste
Martin Holste
Chief Product Architect · Trellix
Martin Holste is the Chief Product Architect for Trellix where he oversees the overall vision for product capabilities. He joined Mandiant in 2013 after running the Security Operations Center and Incident Response Team for the State of Wisconsin for… Read moreRead less

Martin Holste is the Chief Product Architect for Trellix where he oversees the overall vision for product capabilities. He joined Mandiant in 2013 after running the Security Operations Center and Incident Response Team for the State of Wisconsin for seven years. He founded and developed several Trellix products, including Helix, Cloud Security, and Trellix Wise, and has presented his work at many conferences, including RSA Conference and AWS Re:Invent.

Resources

Photos

Tags

More from GoSec 2026

Also from Martin Holste

On the same topic