Abhir Mehra
AI companies change the models behind the names you build on. Sometimes they say so. Sometimes they do not. Either way we write it down.
IBBIS
Defining which sequences are dangerous enough to screen for, so providers and regulators screen consistently
Next year of Commec (the Common Mechanism), the free, open-source, globally-available DNA synthesis screening tool hosted by IBBIS
Constance Li
Animal Welfare Midtraining Data Creation
Christopher Head
Reproducible detector that reads whether a manipulation is still active in an LLMs stream. xfers across 6 model families.
Salvatore Barbera
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
Stewy Slocum
Abeer Sharma
Nikhil Maturi
An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona cha
Raffaello Fornasiere
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Gaurav
Accepted paper at Mechanistic Interpretability for Foundation Models workshop.no travel funding. Early Career Researcher
Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work
Adrian St. Vaughan
Published. Validated on 1,200 cases (97.9–99.5%). The reasoning layer has a bug - we proved the fix works. $9,800 / 90 days to ship the open-source toolkit.
陳鈺澔
An AI platform for crypto and stock analysis, news verification, scam detection and wallet safety, with an open-source Safety Kernel tested on TON.
Tomáš Gavenčiak
An online overview of the AI safety techniques we use in the defense-in-depth stack
Justin Shenk
Increasing public awareness of AI risks and benefits through in-person, interactive experiences
Jordyn Harland-Graham
Persistent memory in AI using LoRAs
Pete Wolfendale
Integrating Selves and Research on Selves in the AI Community
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.
Ankur Pandey
AI safety for builder hackathon / fellowship - to build tools, products, etc.