Reading
Published work I am reading for each of my research interests. None of it is mine. Every entry is the peer-reviewed version, from a conference or a journal.
AI-augmented adversarial attack and defense
How automated offense changes attacker cost, and what detection has to do once reconnaissance and evasion are cheap.
Most of the attention here goes to attack demonstrations. I think the benchmarks matter more. Cybench, NYU CTF Bench, and AgentDojo turn "could an agent do this" into a number, and a number is what lets a defender argue about coverage instead of intuition. I am more optimistic than most about which side gains from this. The same automation that produces an exploit also produces the detection for it, and the defender gets to run it against their own environment first.
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
- PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing
- Jailbroken: How Does LLM Safety Training Fail?
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
SOC optimization
Where AI assistance lowers analyst load, and where it moves the bottleneck instead of removing it.
Both surveys reach the conclusion I reached in production: alert volume, not detection logic, is the binding constraint. That is where I think AI earns its place first, filtering noise rather than trying to replace the analyst. Rule generation is the other half of the problem. If writing a detection gets cheap, coverage expands, and the filtering has to improve at the same rate or the analyst ends up worse off than before.
- From Texts to Rules: Generating Sigma Rules with Large Language Models from Cyber Threat Reports
- Alert Fatigue in Security Operations Centres: Research Challenges and Opportunities
- Alert Prioritisation in Security Operations Centres: A Systematic Survey on Criteria and Methods
- A Human Capital Model for Mitigating Security Analyst Burnout
Threat hunting and adversary intelligence
Hunting, OPSEC, and tracking adversary infrastructure over time, including what that tradecraft costs to sustain.
Provenance-based detection answers the attribution question and creates a volume question. HOLMES, UNICORN, and MAGIC each produce a graph that an analyst still has to read. The Tea Leaves result is the uncomfortable one, because threat intelligence feeds agree with each other less than most teams assume. Given the choice, I would rather expand rule coverage and filter hard than trust any single feed’s precision.
- MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation Learning
- HOLMES: Real-Time APT Detection through Correlation of Suspicious Information Flows
- UNICORN: Runtime Provenance-Based Detector for Advanced Persistent Threats
- ATLAS: A Sequence-based Learning Approach for Attack Investigation
- Reading the Tea Leaves: A Comparative Analysis of Threat Intelligence