You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Compile a list of all the video game evals for AI that I can find, and stare at them to see what they tell us.
To better understand the limits of LLMs capabilities by looking at how they perform on video games. Especially video games with high skill ceilings and rich dynamics.
AIs are still pretty bad at a bunch of games, so looking at how they perform at them gives us meaningful signal on what the models lack. Which is why a bunch of people have already made video game evals for llms. However, no one has collated these evals and comprehensively analysed them to see what they have to tell us.
Hence this project: I will hunt down and store as many video game evals as I can find and analyse as many of them as I can stomach or until I stop learning new things.
As a stretch goal, I'd like to modify some of the evals to stress test AI's ability to generalize by forcing the model to pursue different objectives than the norm e.g. challenge runs like "do not pick up any new wands in Noita".
To pay for 1 month of working on this project.
Just me. I have written literature reviews on different things in the past (e.g. ambitious methods in education, strategic reasoning in LLMs). I have also reviewed, corrected and improved evals as a contractor for Equistamp for 6 months.
Or it turns out to be hard to come up with new objectives in a game that force models to use different strategies.
None for myself.