You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Nowadays we let AI systems grade other AI systems. Leaderboards, safety evals and preference data all pass through an automatic judge that is rarely measured, and almost never against sources anyone can check. Whoever builds on that judge inherits its errors without seeing them.
I have started measuring it, and the first numbers say two things. Asked to grade its own citations, one model claimed 65–96% of them were correct; checking the sources directly found 47–67% (twenty themes, three models out of four): the judge believes itself more reliable than it is. Under a simple, unargued "you are wrong", a production model withdrew 77.6% of the answers it had just confirmed as correct; on the same questions, without the objection, it changed 1.2%. It did not change its mind because it was wrong; it changed it because it was contradicted. A judge like that rewards whoever insists, and if two such judges agree they may share the same weakness.
By March this project delivers the missing number: how far an automatic judge agrees with itself and with checkable sources, measured on a test it has never seen, with the "reliable enough" bar written down and deposited with a DOI before a single data point exists. The result is published whatever it says, together with the open dataset and the code.
Whoever funds it buys a measurement that is seldom published in this form, reproducible by anyone, and a public ledger where my own best numbers were withdrawn, with the date next to them. The next step, outside this request: knowing in advance on which contested claims a model will give way. I was building on a tool I had never calibrated, and an error at the base brings down the building. That is where the measurement began.
Two things, both pre-registered before any data is generated.
A fresh sealed bench. Items disjoint from my existing labelled set, balanced against the theme prior, stratified with per-stratum baselines, and sized against a minimum effect declared in advance instead of the effect I happened to observe. Planning envelope: ~284–562 cells, $335–500 of compute at list prices. The pre-registration is deposited with a DOI before the first cell exists.
Judge test–retest. How consistent is an LLM judge with itself under declared perturbations — re-roll, order, phrasing — reported as chance-corrected coefficients against a usability criterion written down beforehand. Published either way, including if the answer is boring.
Where this leads. This grant buys one link in a chain, and only the first. Measured judges come first: until we can say how reliably a judge agrees with itself and with checkable sources, nothing built on top of one is worth trusting. The next link is a detector for how a model handles a claim that is contested rather than settled. After that, a correction loop that resolves disagreement against verifiable primary sources instead of preference agreement. The link worth having at the end is systems that say plainly when they have no source, rather than producing fluent text anyway. Only the first link is in scope here. The rest is a direction I can argue for, not work I am promising to deliver in six months.
Milestones. Start: 2026-10-01. Six months of full-time work, one milestone per month.
Month | Window | Milestone
- M1, October 2026: pre-registration written and deposited with a DOI, before any cell exists.
- M2, November 2026: sealed items built and held out encrypted, balanced against the theme prior.
- M3, December 2026: bench run executed against the sealed set, with per-stratum baselines.
- M4, January 2027: judge test–retest under the perturbations declared in the pre-registration.
- M5, February 2027: written report against the pre-registered usability criterion, negative result included.
- M6, March 2027: dataset released CC BY-SA, code AGPL, everything archived with a DOI.
At the $5,000 minimum, M1–M3 and the release still ship on this calendar; M4 does not happen and M5 shrinks to the bench's own result, published with its baselines. No salary is drawn.
$25,000 part-funds six months of full-time work: author time, non-Anthropic compute, dissemination and archiving. The full budget is EUR 40,000 (~$45,000).
At the $5,000 minimum I skip the salary entirely and ship one module: the sealed bench's compute run plus its deposited pre-registration and the CC BY-SA dataset. No test–retest programme — one sealed, honest measurement, published with its baselines. I would do it anyway, because an AI that can be trusted is worth a different future, and from there nobody should step back. If one detail can bring down a building, decisions taken on complex but flawed logic can do far worse to a society.
If both this page and the TAIF application are funded, money raised beyond the EUR 40,000 budget goes, in this order, to marginal uses declared to the funders — the first two in the TAIF application, the third added to it by addendum: (1) a second-annotator blind re-read on a subsample of the gold labels, turning the declared single-annotator limit into a measured inter-annotator figure; (2) sizing the sealed bench at the upper end of the published envelope (~562 cells) for higher power; (3) a public, leak-free release of the existing 2,000-cell corpus as a per-claim bias dataset, in the redaction that survives a pre-declared disclosure audit, with contamination canaries and machine-readable metadata. Anything beyond that is returned.
Outputs: dataset CC BY-SA, code AGPL, written report. Negative result published either way.
I run this alone, in Italy, with no affiliation, which normally reads as a credibility problem. Working alone was not a preference. Around me there is no colleague available for an AI project that starts from the humanities and from ancient texts; waiting for a research group would have meant never starting. Here is my counter-offer. Under a pre-registered blind protocol I withdrew my own best numbers, with dates on the public repository, and I publish the surviving one next to a baseline that beats it: 49/62 = 79.0% blind, against a zero-parameter trivial rule at 90.3% (p = 0.033) and a majority baseline at 72.6% (paired p = 0.23). I withdrew them because I always check the chain of reasoning that produces a result, not just the result, and that day the chain did not hold. I run the same check whether a number is surprising or discouraging. My strongest figure — 253/269 = 94.1% versus a 47.6% baseline, at the packaged build of 2026-09-02 — is in-sample, and I refuse to call it accuracy. Bench so far: 2,000 cells, 200 claim slots, 283 gold labels, one annotator (a stated limit; a second annotator on a subsample is an option, not a promise). Two experiments were killed by my own checks before they cost money. An internal adversarial audit of the gold labels found a ~2% contamination rate in one label class; the corrections LOWERED my internal numbers, are recorded with dates, the hardest case produced a new written labelling rule.
Repo: github.com/claudiodegenua/precorrect-method · DOI 10.5281/zenodo.22345121 · ORCID 0009-0008-5896-3172. A second, unrelated dataset of mine (a six-layer parallel Psalter, 2,469 verse-cells, per-layer measured quality) is deposited as DOI 10.5281/zenodo.22770451 — evidence that I ship data with a data sheet, not only claims. Recent work on verifier reliability is converging on this question — arXiv 2506.13342 (Verifying the Verifiers) on label quality in fact-verification benchmarks, and arXiv 2606.19544 (Reliability without Validity) on chance-corrected agreement and test–retest across 21 judges. What I add is orthogonal to both: the bench is sealed and the usability criterion is pre-registered with a DOI before any cell exists, and ground truth resolves to checkable primary sources rather than preference agreement — which is what makes a usability verdict, rather than a reliability description, possible at all.
How to check me. Every number I withdrew is listed with its date in the public repository — github.com/claudiodegenua/precorrect-method — next to the one that replaced it. Read that ledger before you read my claims.
The headline risk is already on the record, not hypothetical. On the current blind set, a zero-parameter rule that looks only at WHICH THEME a claim belongs to — never reading the text under test — scores 90.3%, beating the detector at 79.0%. Finding that out was unpleasant and useful in the same moment: it changed the design of the bench. I did not drop the project for a simple reason: a residual error on the themes that matter is not small, and I do not believe machines will close it on their own while nobody works on the structure that keeps them in line. That number says nothing about models being right: it says the blind set's difficulty is concentrated in the theme prior, so the bench so far measures theme difficulty more than text reading. This matters beyond my project: aggregate scores hide exactly the strata that matter most — a judge can look excellent on average and still fail on the contested themes it exists for. The new sealed bench is designed against precisely this failure: balanced against the theme prior, with per-stratum baselines declared in advance. If it confirms the null, that is a publishable negative result, and this grant buys precisely that answer. Second risk: a single annotator — a declared limit, and the first marginal use converts it into a measured inter-annotator figure. Third: benchmark contamination — every released file carries a project-specific canary string, and the sealed holdout ships encrypted rather than in plain text.
Nothing received to date. No grant, prize, salary or donation has been received for this work in the last 12 months. What exists are applications, not funds:
EA Funds — Transformative AI Fund (TAIF): an application to EA Funds' Transformative AI Fund for the full budget was submitted on 2026-09-09 and is pending. An application, not money received.
Anthropic ERA: a request ($1,000, compute-only credits, submitted 2026-09-05) was not approved in the September cycle (the programme notifies only successful applicants); I intend to reapply in the 5 October cycle. It is strictly complementary — it buys Anthropic-model inference, not time. An application, not money received.
AI Safety Fund (AISF): an application is pending. An application, not money received.
There are no bids on this project.