You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
The goal is to make beancount-ledger useful as a benchmark that other researchers can run, inspect, and reproduce without relying on my own interpretation of the results.
The first part is an Inspect port. UK AISI's Inspect framework is already used for model evaluation work, so supporting it should make the environment easier to run alongside existing evaluation suites.
The second part is a larger model comparison. My first experiment covered four models and 138 valid rollouts, but I stopped when the free inference quota I was using ran out. With funding, I want to test roughly 15–20 open-weight models with repeated rollouts instead of relying on a single attempt per model.
Before running the final comparison, I will preregister the evaluation setup so that the model list, task sampling, rollout counts, scoring rules, and reporting format are fixed before I see the results.
The third part is scorer hardening. I already maintain adversarial cases for situations where an agent might receive credit without actually completing the bookkeeping task correctly. I want to expand that test set as I add more models and task variation.
Most of the environment is already built. The main thing I need now is enough inference and compute to test it properly at a larger scale.
Most of the funding would go directly to inference and compute.
At the minimum funding level of 2,000, I would spend roughly 1,000 on model API/inference costs for the main evaluation, around 600 on compute for the Inspect port, testing, task-generation runs, and CI, and the remaining 400 on adversarial runs focused on finding scorer failures.
If the project raises more than the minimum, I would not expand the scope. I would use the extra budget to run more models, increase the number of rollouts per model, and cover a larger portion of the task population.
At the 8,000 goal, the main difference would be statistical coverage: more repeated runs, more model families, and more adversarial testing. The core deliverables would stay the same.
I am not budgeting for salary or contractor costs. The environment, scorer, task generator, tooling, and initial experiment are already built.
I am working on the project independently.
I built beancount-ledger end to end, including the task generator, bookkeeping environment, deterministic scorer, adversarial tests, release tooling, and the first evaluation run. The initial experiment produced 138 valid rollouts across four models.
A big part of the work has been making the results auditable. Instead of relying on an LLM judge, reward is tied to explicit accounting checks, and I have published the source code and benchmark evidence publicly.
Outside this project, I have worked on AI evaluation and agent QA. At Fleet AI, I reviewed agentic tasks and separated more than 100 automated-verifier defects from genuine agent failures. I have also worked on RLHF code-preference evaluation across Python tasks and libraries.
Source code:
https://github.com/gultekinhasancan79/beancount-ledger
Live environment:
https://app.primeintellect.ai/dashboard/environments/cangultekn/beancount-ledger
The most likely failure is not that the project disappears, but that the final evaluation ends up smaller or slower than planned.
That already happened in the pilot. I stopped after four models because I exhausted the free inference quota I was using. If funding is limited, I may have to reduce the number of models or repeated rollouts.
The Inspect port could also take longer than expected. If that happens, I would not hold back the whole project. I would publish the broader evaluation first and finish the integration separately.
A more interesting failure would be finding a scorer exploit that I missed in the current version. With more models and more rollouts, that is possible. If it happens, I would fix the issue, add it to the adversarial regression set, rerun the affected evaluations, and document the failure publicly.
So the main downside is a narrower timeline or measurement, not losing the underlying work. The environment and current benchmark infrastructure are already public and usable.
I have not raised any money for this project so far.
Everything currently public — the environment, task generator, scorer, benchmark population, tooling, and initial evaluation — was built using my own unpaid time and free-tier infrastructure.
I also submitted a BlueDot Rapid Grant application this week. It is still pending, and I have not received any funding from it.
If both applications are successful, I will not use them to pay for the same expenses. I will separate the funded line items and report that clearly.