You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
This is the security arm of my company Small Mind. I red-team agentic AI products launching all over the place that you put data into and they act on it. I also test whether the benchmarks we use to call these products secure actually catch the real attacks. I find the holes in real agent products through prompt injection. Then I prove that the leaderboards vendors point to when they say "enterprise-grade security" are actually evaling the wrong thing.
2 goals. One is developing a red team harness for direct and indirect prompt injection in agentic systems to protect against confused deputy class attacks. These attacks target agents with tool access and permissions that read from text, emails, documents or files. The attack hides prompts and gives instructions that the agent then acts on. Every "AI agent platform I've seen is still vulnerable to these attacks, and too many companies don't have an answer for it. They just take the models at face value and assume they work and the benchmark stamp supports that. I run attacks against local open weight models, and deployed commercial agent APIs. I write it up and either give it back to the company or publish the findings. This way providers can patch the loop holes and independents can expand on the work. The second part is expanding on my eval integrity work. I took the Jailbreakbench leaderboard and every attack they run and changed nothing but the judge model they use for scoring. Four of the five alternate judges I put in reordered the leaderboard. Minimum Kendall tau was 0.115. The ranking is almost 50/50 depending on which judge you pick. A parser defect alone in a reference judge flipped 534 of 1637 rows on one model. The paper was pre-registered and commited on GitHub (GitHub.com/Threadborne/eval-sensitivity) With this funding I can run the 70B reference judges locally. This is where I got pushback from the community because I used "smaller" models. Their argument is that I'm not comparing apples to apples. Running the 70B models would put the debate to rest. The way I achieve these goals is with unlimited local runs on owned hardware. Then I rent burst compute for the frontier scale open models and API spend to test the deployed commercial agents. I will keep using my pre-registered methodology on the eval side so the results can't be dismissed out of pocket.
Down to the penny. Hardware and compute only:
Local runs:
2x Lenovo ThinkStation PGX workstations (NVIDIA GB10 Grace Blackwell, 128gb unified memory, 4tb Gen5 self-encrypting NVME, DGX OS) @ $7,989.00 each = $15,978.00 total
UPS + metered PDU (multi day evals will survive brownouts) = $329.98
Rented compute & model access:
GPU burst for frontier open weight models (vast/runpod) = $2,000
Commercial agent API credits (the deployed agents I'm testing) = $3,000
Total: $21,307.98
The 2 ThinkStations are the backbone. I need thousands of generations per attack and they have to be uncensored and not metered.I also have to run 70B models. That's why I need local. It was something like $30k just in rented server time to run all this. I spent a great deal on rented servers just for the first judge. This is in the literal millions of tokens. Its taking almost 16 days per judge on my current setup to generate. The self encrypting disks are also a must, I'm running client adjacent data and these encrypt at rest. The rented side covers the models too big to run locally, period. And the commercial agents that are only available behind API.
Just me. Small Mind LLC
I was raised in forums and the piracy scene moving into ethical hacking, red teaming and pen testing. I've had a computer since 2nd grade and I've been online through every platform generation watching how they've been broken and fixed since yo-yos were hype. Right now I'm red teaming models on Gray Swan Arena and active on HTB. I also run AI security assessments, mostly on AI start ups to try and close some of the big holes. The harness should make that a ton easier. Before any of the AI work I was in the 101st Airborne, company RTO and heavy weapons team leader. Deployed in OEF 10 and held secret clearances.
Most likely causes of it failing are the vendors just straight up don't want to see the results. This is all still new and even though everyone talks about security with AI its usually about them breaking out of sandboxes. Every day I see companies and vendors just straight up ignoring security where it involves people's sensitive data. And the second part of that is the eval-sensitivity results being dismissed. For the same reason. All the benchmarks are using AI to grade and judge AI. Even preregistered they want to nitpick results because theyre so heavily invested in the leaderboards.
Even if the project fails it doesn't fail hard. The results will be published, the hardware this grant would provide doesn't get shut off with a pat and a "we tried". I keep going and keep publishing and creating open tooling. If they want to keep arguing the eval work and dismissing it. Thats the point. The conversation drives the change. The worst case is the money provides local research and a mountain of working attacks anyone can run themselves, out in the open. There is no version of this where the computer sits idle. Breaking things and pushing the security industry is what I do. It doesn't matter if anyone pays me for it or not.
Nothing yet. Entirely out of my own pocket. I have one other fundraiser live on Manifund right now for the research arm of Small Mind. "Persistent maybe conscious AI and whether it chooses for itself" ($15k goal, closes in mid nov) which as of writing this is unfunded. This would be the first outside money my work will have received.