Funding requirements
Celeste
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Worked on a bunch of meta model stuff: LoRAcles and improving AOs
/
some mats funding
Jan Bauer
Matthew Farr
Probing possible limitations and assumptions of interpretability | Articulating evasive risk phenomena arising from adaptive and self modifying AI
I self-funded research into a new threat model. It is demonstrating impact (accepted at multiple venues, added to BlueDot's curriculum).
Mejdi Sadriu
Matthew A. Clarke
Working title - “Compositionality and Ambiguity: Latent Co-occurrence and Interpretable Subspaces”
Lucy Farnik
6-month salary for interpretability research focusing on probing for goals and "agency" inside large language models
Logan Graves
A formal, testable account of LLM persona selection as Bayesian inference, validated with model internals, so labs can monitor and steer personas.