Jan Wehner

About

Technical governance to reduce risk from AI

Hello, I’m Jan Wehner. I’m a researcher on the Frontier Security Team at the Institute for AI Policy and Strategy (IAPS), where I work on technical AI governance to reduce catastrophic risk from AI, spanning agent safeguards, AI×Cyber, government use of AI and AI strategy. At the moment I’m threat modelling how malicious AI agents deployed in national security agencies could cause large-scale harm and how AI control could prevent that. I’m also working on designing honeypots to detect, incriminate and deter malicious agents deployed internally.

Previously, I was a Winter Fellow at GovAI under Alan Chan, and I researched AI Safety and Interpretability at the CISPA Helmholtz Center for Information Security, supervised by Prof. Mario Fritz and Prof. David Krueger. My past work spans Representation Engineering, harmful fine-tuning attacks and Inverse Reinforcement Learning. During my MSc I also founded the Delft AI Safety Initiative.

I love mentoring research projects — see my mentoring page for past projects and opportunities to work with me.

Please don’t hesitate to reach out to me at jan(at)iaps.ai!

Recent publications

All publications
IAPS Report2026

Detecting Offensive Cyber Agents: A Detection-in-Depth Approach

Matt Mittelsteadt, Jam Kraprayoon, Robin Staes-Polet, Oskar Galeev, Jan Wehner, Christopher Covino, Shaun Ee

AI agents can now orchestrate cyberattacks, increasing their speed, scale, and autonomy while lowering costs. This report frames the offensive cyber agent detection challenge, introduces detection-in-depth as a strategic framework for policymakers and defenders, and proposes five concrete detection mechanisms.

Recent writing

All posts

How Durable is Our Visibility into AI Cyberattacks?

Our ability to monitor AI-enabled cyberattacks may be far more fragile than it looks. Today we get useful signal about how attackers use AI, but much of that visibility could erode as adversaries grow more sophisticated and move off detectable channels.

  • AI Safety
  • AI Governance
  • Cybersecurity

Will we get automated alignment research before an AI Takeoff?

AI may automate large parts of AI R&D within the next decade, dramatically accelerating progress. A crucial question for existential risk is the ordering: will automation speed up capabilities research or safety research first? If capabilities race ahead while safety lags, we could find ourselves with very powerful...

  • AI Safety
  • AI Governance