Portfolio item number 1
Short description of portfolio item number 1
Short description of portfolio item number 1
Short description of portfolio item number 2
This paper introduces the notion of Robustness to Qualitative Constraint Networks and finds a tradeoff between speed and robustness in heuristics for solving QCNs.
We propose a method for explaining reward functions by showing the rewards given to counterfactual trajectories.
LLMs can be fine-tuned with harmful data to remove their safeguards. We formalize the problem and set out conditions for a solution.
We propose Representation Noising which prevents harmful fine-tuning by removing harmful representations.
Open-ended AI is a growing paradigm where AI continuously explores novel and interesting artifacts. This position paper describes specific safety challenges in Open-Ended AI and how they can be mitigated.
This survey paper reviews the literature on Representation Engineering, a technique for controlling LLMs through their internal representations. We set out a unifying taxonomy, describe methods and applications and showcase weaknesses and opportunities.
AI agents can now orchestrate cyberattacks, increasing their speed, scale, and autonomy while lowering costs. This report frames the offensive cyber agent detection challenge, introduces detection-in-depth as a strategic framework for policymakers and defenders, and proposes five concrete detection mechanisms.
This is a description of your talk, which is a markdown file that can be all markdown-ified like any other post. Yay markdown!
More information here
More information here
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
This is a description of a teaching experience. You can use markdown like any other post.
This is a description of a teaching experience. You can use markdown like any other post.