Pip Foweraker
A nightmarishly hard AI safety strategy game about holding p(Doom) down. You can't win; you can only buy time.
Taehyun Cho
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide rather than as a reward to maximize.
Felix Harder
A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization.
Christopher Leet
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains t
Florian Dietz
Clearing barrieers to adoption for an existing ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms
Deepanshu Goyal
One training-free geometry fitted to a model's residual-stream activations that reads a state, moves it, and tests whether the behaviour follows
Michail Patsakis
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, prote
Eitan Sprejer
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries.
Jai Dhyani
Creating conditions for cooperative strategies to dominate adversarial ones among near-future AIs while we still can
Nickola Horozov
Karthik Viswanathan
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how man
lucas.irwin
A policy memo, co-authored with the Institute for Public Policy Research, resolving the open technical, economic, and legal questions blocking real-world implem
Yunika Bajracharya
Five-week AI safety fellowship + 3-month project mentorship
Ari Spiesberger
Perform research to rigorously elucidate and quantify generalization versus memorization, and examine evidence of originality in LLMS.
Sergey
Fund demonstrated/rigorous quantitative researcher (already run reproduction/audit pipelines on published economics) for 6-month AI safety transition, shipping
Aryo Pradipta Gema
Measuring whether CoT monitoring fails when an influence reaches an agent through a tool return rather than the user message.
Joshua Reiners
Cedric Potvliege
Yunus Emre Tapan
A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.
Octavian Untilă
The same grounding instruction leaves 0–65.5% leak depending on the judge; forcing per-claim counting closes it to 0.5%. 4 model families, 2,000 sealed calls.