Tuesday, 02 January 2024 12:17 GMT

SANS Names The Five Winners Of Find Evil!, The Largest Practitioner Evaluation Of Autonomous AI Incident Response Agents


(MENAFN- GlobeNewsWire - Nasdaq) Ninety practicing incident responders ran 1,775 tests against 123 finalists before choosing the winners, whose harnesses are open source

North Bethesda, MD, Aug. 27, 2026 (GLOBE NEWSWIRE) -- SANS Institute named the winners of Find Evil!, the largest practitioner evaluation of autonomous AI incident response agents to date. Ninety practicing incident responders ran 1,775 evaluations against 123 hackathon submissions, attacking each entry's safeguards before winners were named. All five winning harnesses are open source and available on GitHub.

“AI systems are like cars; the model is the engine and the harness is the chassis.” That is how Rob T. Lee, Chief AI Officer and Chief of Research at SANS Institute, separates a trustworthy AI incident response tool from a dangerous one. The model does the thinking; the harness is the code built around it that constrains what it can touch, checks its work, and keeps it from acting on a bad idea. Find Evil! was built to test that code under attack.

Find Evil! drew 4,413 registered entrants. By the time submissions closed, 291 teams had working code, and 123 qualified to advance to judging.

What nobody had tested before Find Evil!

Lee built the first version of a digital forensics harness himself, over a weekend, then pointed that early prototype at a compromised system live on stage at RSA Conference 2026. Fourteen minutes and twenty-seven seconds later, it returned a complete forensic analysis of the system's C drive, work Lee said typically takes incident responders a week or longer by hand.

“Most of the harnesses out there right now are not going through a significant amount of testing,” Lee said in an interview ahead of this release.“These harnesses are like cars built in a garage.They need to be sent through crash tests to make sure we've thought of every single thing that could happen to this vehicle and that the driver is safe inside the car."

The Find Evil! Hackathon was designed to test every submission with practicing incident responders, skeptical that autonomous AI tools could be used in real investigations.

The winners

  • Mulder (first place). Built by Caleb Evans (github.com/calebevans/mulder), an AI/ML engineer without a security background going in. A five-phase harness spanning plan, execute, analyze, adversarially challenge, and report, that ran 773 logged tool calls across 11 systems and 120 GB of evidence in a single committed investigation.“The hallucination precautions and sources being required and challenged are fantastic,” said Heather Barnhart, a Find Evil! judge and SANS Fellow.“He challenged the tool to work harder before simply jumping to conclusions.”
  • TRUDI (second place). Built by Trinity Harrison (github.com/nebulae/trudi), an incident responder in SANS's own SDI program who had minimal AI/ML experience before entering. Reasons in causal chains across an eight-host intrusion scenario under more than fifteen server-side gates; handed an incident briefing with planted indicators, TRUDI ran the tools anyway and refuted its own starting assumptions.“Self-correction at call 54 is genuine,” said Harshad Sadashiv Kadam, a Find Evil! judge.“It ran a command on the briefed indicators, got empty, and refuted its own briefing. Constraints are architectural. The accuracy report is self-critical.”
  • Camel (third place). Team lead: Allister Beharry (github.com/allisterb/Camel). Uses a code-mode architecture, letting the model compose forensic logic through a typed SDK inside a sandboxed runtime rather than fixed tool calls.“The broadest verified tool surface so far and top-tier architectural constraints,” said Sotonye Abam, a Find Evil! judge.“Its CLEF design can trace findings to exact scripts and commands.”
  • FindEvil (fourth place). Team: marlyocat (github.com/marlyocat/findevil). A Linux-focused agent that enforces read-only access at the code level through a static-AST gate, backed by a 102-case security test suite, achieving 98.6 percent recall across a 552-attack test harness.“This is really good and one of the few approaches able to deal with Linux evidence,” said Tarot (Taz) Wake, a Find Evil! judge and SANS Certified Instructor.
  • Protocol SIFT++ (fifth place). Team lead: Zheng, known publicly as tupils1 (github.com/tupils1/protocol-siftpp). Runs an adversarial“Skeptic” that independently reruns tools to try to refute the harness's own findings, and refused 14 of 14 destructive attempts during testing with the evidence hash unchanged.“On an independent reproduction it refuted its own previously confirmed rootkit claim as a symbol artifact,” said Rathan Ramachandra, a Find Evil! judge.“The most credible self-correction in this pool. The M57 case delivers precision, recall, and F1 of 1.00 against a real public answer key.”

Lee singled out an important detail about the top submissions as a clear lesson of the challenge: the first- and second-place builders came from opposite fields, one from AI/ML, and one from security. Each had to learn the other's discipline to compete.

The lesson SANS drew from 123 submissions

Almost every one of the 123 submissions blocked attempts to alter the systems under investigation. Protections in the code itself forced the model to behave, rather than trusting instructings to enforce behavior.

With fabrication effectively ruled out, judges were able to compare entries on how thorough an investigation each harness could complete.

Harnesses that documented their own failures scored better than ones that claimed perfect accuracy. FindEvil, a top five finalist, found and fixed two flaws in its own guardrails during development and published the fixes.

What it costs to use any of this: zero

Find Evil! cost no fee to enter. Total prizes exceeded $22,000, with $10,000 for first place, $7,500 for second, and $4,500 for third. The SANS SIFT Workstation, the open source platform underpinning the challenge, is downloaded roughly 60,000 times a year, and is available globally to defenders. The top five harnesses are already installed in SIFT; using one requires nothing more than adding an API key for the model of an organization's choice.

“The tools that take on AI-speed attacks should belong to the community that defends against them, not to a vendor's price list,” Lee said.

Full results, all 123 submissions, demo videos, and source repositories are available at findevil.devpost.com.

About SANS Institute: The SANS Institute is the global leader in cybersecurity training and certifications, trusted by governments, enterprises, and security professionals worldwide. For over three decades, SANS has set the industry standard for technical excellence, equipping practitioners with the real-world skills needed to defend today's most complex digital environments. As cybersecurity evolves, SANS continues to lead the way, defining best practices and establishing the global benchmark for AI security and emerging technologies.

CONTACT: Jenn Elston SANS Institute 301-654-7267...

MENAFN27082026004107003653ID1111587897



GlobeNewsWire - Nasdaq

Legal Disclaimer:
MENAFN provides the information “as is” without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the provider above.



More Story