Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
BigSleep - Using Large Language Models To Catch Vulnerabilities In Real-World Code
The talk shows how an autonomous LLM‑driven agent explores code, identifies vulnerabilities, and generates exploit inputs using a code search tool, debugger, and sandbox.
BigSleep is an autonomous security agent that finds vulnerabilities in code.
As the code comprehension and general reasoning ability of Large Language Models (LLMs) has improved, we have been exploring how these models can reproduce the systematic approach of a human security researcher when identifying and demonstrating security vulnerabilities. We hope that in the future, this can close some of the blind spots of current automated vulnerability discovery approaches, and enable automated detection of “unfuzzable” vulnerabilities.
The BigSleep Agent has access to a code search tool, a debugger and a python sandbox. It dynamically explores the codebase and figures out potential program points where a vulnerability exists. It then generates an input to demonstrate the crash, as proof.
Big Sleep LLM autonomously finds exploitable memory-safety vulnerabilities in real-world code.
Compose Email
Loading recent emails...