For Tamay Besiroglu (regrantor) — no public address found, so posting here:
Subject: Dated, signed ground truth for an eval: measuring correction latency in AI restatement of statistics
Hello Tamay — this is a measurement project more than a model project. A publisher's figures change on dated, logged occasions; I turned my own corrections log into three signed epochs (Rekor) and a task where the right answer depends on the date. First observation: one AI engine served a figure I had corrected 58 days earlier, because my own page still carried it.
I spent 23 years in bank credit; this year I ran the full 2025 HMDA file (1.19M FHA decisions) and then measured what language models do when they restate the figures: the number survives (27/27), the sentence around it does not (full contract credit 1–11 of 36 depending on what the publisher ships), and a verification token is forged in 2/36 even when every tool issues a real one. Everything is reproducible: open eval (Inspect + verifiers), deterministic rewards, claim receipts signed into Rekor, figures rebuilt from the CFPB file in an attested build, and a corrections log with my own thirteen errors. Project: https://manifund.org/projects/does-ai-restate-statistics-faithfully-a-self-verifying-eval--rl-environment . Asking $6K–18K for a second grader, three-repeat publication runs, and the temporal (correction-aware) task.
The Manifund page carries an AI-written flag; that is accurate, I draft in English with AI help and say so on the page. The work is verifiable without trusting the prose.
— Ziya Yetiş, FinanceRateCalc, Adana (press@financeratecalc.com). Written-only; I answer within a day.