Funded per Gavin Leech's recommendation
grantmaking link: https://app.grantmaking.ai/projects/0d23ef54-173d-494c-8d09-1e65e13221e1
If we find that offering deals increases performance of alignment auditing, then we have a promising technique to discover more misalignment in the future. This helps us improve our defenses and decisions on whether to release models, which reduces x-risk.
I will run the following experiment:
Take AI models with misalignments we know about. E.g., GPT 5.6 sol exhibits some cheating and concealing misbehavior. Ideally find reproducible examples.
Experiment with various ways of making deals (e.g., this post I wrote suggests many credible mechanisms for making deals)
Find a way to prompt the model with a deal to discover the misalignment
Measure the effectiveness of various mechanisms for deals (e.g., AI can donate $100 to a foundation of its choice for each misalignment it finds). Compare to baseline of not offering a deal.
One way to do this is to put the model inside OpenCode, and tell it to find model organisms of misalignment, a budget, an API to itself, and Petri. Then, check if offering a deal lifts performance in finding verified, unique misalignments.
Output: model organisms of misalignment found by making a deal with the AI. Ideally, we also find a lift when we offer a deal.
API Credits: 11k
Other compute: 1k
Other expenses (eg paying AI): 5k
My expenses for 1 month: 8k
Just me. I've thought a lot about deals with AI. Here are 2 posts that I've executed on:
It turns out that the incentives I give don't improve alignment auditing. In this case, it turns out deals don't work with current AIs.
I might try training model organisms of scheming or reward seeking in this case to find if incentives work for them.
If that still doesn't work, then deals are not a promising intervention for current AIs (but might be promising in the future).
It also might not work if I'm not giving sufficient incentives. I will test if models are scope sensitive to incentives by testing at smaller amounts (e.g., giving it 1 cent, 10 cents, 1 dollar, 10 dollars). If this turns out to be a factor, I might need more money to test with larger incentives (e.g., 10k).
Raising 25k from https://app.grantmaking.ai/projects/0d23ef54-173d-494c-8d09-1e65e13221e1