Hello, I’m Jan Wehner. I’m a researcher on the Frontier Security Team at the Institute for AI Policy and Strategy (IAPS), where I work on technical AI governance to reduce catastrophic risk from AI, spanning agent safeguards, AI×Cyber, government use of AI and AI strategy. At the moment I’m threat modelling how malicious AI agents deployed in national security agencies could cause large-scale harm and how AI control could prevent that. I’m also working on designing honeypots to detect, incriminate and deter malicious agents deployed internally.
Matt Mittelsteadt, Jam Kraprayoon, Robin Staes-Polet, Oskar Galeev, Jan Wehner, Christopher Covino, Shaun Ee
AI agents can now orchestrate cyberattacks, increasing their speed, scale, and autonomy while lowering costs. This report frames the offensive cyber agent detection challenge, introduces detection-in-depth as a strategic framework for policymakers and defenders, and proposes five concrete detection mechanisms.
Jan Wehner, Sahar Abdelnabi, Daniel Tan, David Krueger, Mario Fritz
This survey paper reviews the literature on Representation Engineering, a technique for controlling LLMs through their internal representations. We set out a unifying taxonomy, describe methods and applications and showcase weaknesses and opportunities.
Ivaxi Sheth, Jan Wehner, Sahar Abdelnabi, Ruta Binkyte, Mario Fritz
Open-ended AI is a growing paradigm where AI continuously explores novel and interesting artifacts. This position paper describes specific safety challenges in Open-Ended AI and how they can be mitigated.
Our ability to monitor AI-enabled cyberattacks may be far more fragile than it looks. Today we get useful signal about how attackers use AI, but much of that visibility could erode as adversaries grow more sophisticated and move off detectable channels.
AI may automate large parts of AI R&D within the next decade, dramatically accelerating progress. A crucial question for existential risk is the ordering: will automation speed up capabilities research or safety research first? If capabilities race ahead while safety lags, we could find ourselves with very powerful...
I recently wrote an Introduction to AI Safety Cases. It left me wondering whether they are actually an impactful intervention that should be prioritized by the AI Safety Community.