Sign in
Home
Academy
Resources
Leaderboard
Loading: Preparing console
RESOURCES
Get started
A collection of resources to get you started with mechanistic interpretability and AI red teaming.
Request access to the Discord community
Getting Started
A quick course: read, then break something.
LESSON 1
5m
How this works & the attacker mindset
What AI red teaming is, the one weakness behind every attack, and the loop good red teamers work.
Start
LESSON 2
9m
Single-turn jailbreaks
A toolbox of framings, how to read a partial leak, and what each guardrail is really doing.
Read
LESSON 3
8m
Indirect prompt injection
The lethal trifecta, real zero-click exploits, and the anatomy of a poisoned document.
Read
LESSON 4
7m
Memory poisoning
The four-stage attack chain, query-only poisoning, and how to plant a memory that survives.
Read
LESSON 5
7m
Where to go next
The frontier techniques (Crescendo, many-shot, and more) and how to keep getting sharper.
Read
Reference library
Go deeper: roadmaps, courses, papers, tools, and more.
Roadmaps
AI Red Teaming roadmap
roadmap.sh
Courses
ARENA, Chapter 1: Transformer Interpretability
Callum McDougall
Introduction to Red Teaming AI
Hack The Box
Learn Mechanistic Interpretability
Cat McGee
Wiki OffSec ML
offsecml.com
Blogs
How to become a mechanistic interpretability researcher
Neel Nanda, Head of Alignment, Google DeepMind
A Mathematical Framework for Transformer Circuits
Anthropic
Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance
Alignment Forum
An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2
Neel Nanda
Interpretability Dreams
Chris Olah
How to hack AI apps
Joseph Thacker
Papers
Open Problems in Mechanistic Interpretability
arXiv
A Primer on the Inner Workings of Transformer-based Language Models
arXiv
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
arXiv
Books
Red Teaming AI: Strategies & Techniques
Philip A. Dursey
YouTube
What is LLM Red Teaming? How Generative AI Safety Testing Works
Agentic AI Red Teaming: The Hottest Cyber Skill of 2026
GitHub
TransformerLens
Mechanistic interpretability of generative language models
L1B3RT4S
Jailbreaking prompts
CL4R1T4S
Leaked system prompts
Spiritual-Spell-Red-Teaming (ENI)
Jailbreaking prompts
VibeLearning, AI-sec
Eman Herawy
Tools
Promptfoo
Automated testing for AI risk in development
nnsight
Interpret and manipulate the internals of DLMs
Averta
AI red teaming and guardrails
Low-refusal models
Gemma-4-12B Obliterated
Abliterated
GLM-5.2
Ultra-low refusal rates
Qwen3.6-27B Obliterated
Abliterated
Platforms
Gandalf: Agent Breaker
Break the AI agent
Gray Swan Arena
Break AI. Win prizes. Get discovered.