You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI can make a research claim sound finished before anyone checks it. I made a one-page Falsifiability Sheet to slow that down. It puts one claim beside the exact sources used for it. A person checks the claim first. Then AI checks only the same sources. The result is Green, Yellow, Red, or STOP. STOP means the needed evidence is missing. The sheet does not decide what is true, and it does not replace a trained scholar.
This grant would test whether the sheet holds up beyond a demo. Over three to four months, I will run eight cases with answers set before the work begins, plus one real research case from Alexandria. Each case gets a human check. The AI step will run 54 times across three model families and fresh sessions to see whether the result changes. Every result, disagreement, error, refusal, and removed case will be posted. If the sheet fails, I will post that too.
The real-use case comes from Mahmoud El Tahan, an Egyptology master's researcher at Alexandria University, who has agreed to try the Arabic sheet on one small part of his dissertation work. Alexandria University has approved a guide in Arabic and English for using AI in research and has discussed new AI research tools. This is work between two researchers, not an official university project.
The Falsifiability Sheet: one claim, one set of sources, two separate checks, and one saved result.
Why fund this now - and why it fits Manifund
I do not need money to come up with the idea. I need money to test what I already built. A $12,000 grant covers the basic study, outside review, the Arabic case, and a public record. The main study can be finished within four months. A bad result still matters because it will show where the sheet fails or changes too much.
What I have already built and what the grant pays for
Version 6A is online as a full example. It includes the sources, human check, AI check, final result, and test limits. I have also tried the steps with ChatGPT, Claude, and Gemini, and an Arabic draft is ready. That shows the steps can be followed. It does not prove the sheet is reliable.
The Falsifiability Sheet began with V1 and has been developed through successive versions, public examples, testing, and outside feedback. A University of Dayton discussion helped sharpen questions around evidence isolation, reviewer separation, and what happens when the result is not clearly Green or Red. That feedback helped shape what this pilot now needs to test rather than serving as an endorsement of the project.
The grant pays for what is still missing: test cases with known answers, outside reviewers, new AI sessions, one real Alexandria case, an Arabic check, and a public record of the full test.
Working V6A record — Zenodo | Open project repository
Why start with ancient writing?
Egyptian hieroglyphs are already understood, so I can test the sheet against known answers. Ancient writing is a good challenge because small details matter and a sure-sounding claim can go too far. This study will not prove the sheet works in every field. I am asking whether it works here and whether a larger test is worth doing.
What are this project's goals? How will you achieve them?
Ancient writing is the controlled starting point because Egyptian hieroglyphs give the study known answers against which the process can be tested. It is a test environment for the process, not a claim that the process already works in other fields.
If the process stays stable under these controlled conditions, the result will tell us whether testing the same evidence-bound approach in other research settings is worth the effort.
What the $12,000 study includes
Eight planned test cases: Before the study starts, I will post each case and its expected answer. There will be two Green, two Yellow, two Red, and two STOP cases.
One real Alexandria case: Mahmoud will use the Arabic sheet on one small part of his dissertation work. Private details will be removed when needed. If the case cannot be used, I will use a public backup and say clearly that the real-use test did not happen.
One outside human check for each case: The reviewer will not be told the expected answer. Mahmoud will not be the only reviewer of his own case.
Fifty-four AI runs: 9 cases x 3 model families x 2 fresh sessions. This tests whether the result changes across models or new sessions. It is still nine cases, not 54 separate cases.
Everything will be shared: the English and Arabic sheets, prompts, data, and final run record. Source packets will be shared when I have permission. The release will show every run, the reviews, removed cases, and what failed.
What the $18,000 goal adds
Twelve cases in all and 72 AI runs.
Two more real or public cases.
At least one person outside my project will run a full case from start to finish.
Some cases will get a second human check. Part of the study will also be run again to see if the results hold.
How the test works
1. Set the rules first. Before the first run, I will post the claims, sources, prompts, expected results, removal rules, and final decision rules. Digital hashes will show if a file changes later.
2. Keep the two checks apart. R1 is the human check. R2 is the AI check. R2 sees only the claim, sources, and blank form - not R1's answer and not the web. If the needed evidence is missing, R2 must say STOP.
3. Start with a clean session each time. Each case will run twice in ChatGPT, Claude, and Gemini. I will save the model name, date, settings, files, prompt, refusal, and full answer.
4. Compare the results. I will count every time a known Red or STOP case becomes Green, check how often the AI runs agree, show human-AI disagreements, and record whether the Arabic sheet is usable from start to finish.
5. Post the result even if it is bad. I will share the rules, data, removed cases, errors, and limits. Only material I cannot share or details that could identify someone will be held back.
What counts as failure, and what still helps
The worst error is a known Red or STOP case ending as Green. Large changes between new sessions or model families also count against the sheet.
The study does not need a good score to matter. A useful result can show where the sheet works, where it changes too much, where it takes too long, or when it should not be used.
Timeline
Month-by-month plan
Month 1: Set the rules, have a professional Arabic proofreader check the sheet, test the QR code, hire outside reviewers, and post the plan before the first run.
Month 2: Run the eight planned cases. Save every file and answer. Fix only small process problems allowed by the posted plan.
Month 3: Run the Alexandria case and finish the new AI sessions. Record what was easy, what was hard, where the evidence stopped, and where reviewers disagreed.
Month 4: Review the results and get an outside check on how I ran the study. If the $18,000 goal is reached, repeat part of the work. Post the data, English-Arabic kit, guide, and final report.
How will this funding be used?
Budget
The $12,000 minimum pays for the whole basic study. The $18,000 goal adds more cases and more outside checking.
Where the money goes
My work: set the rules, prepare cases, run tests, study results, and post the files. Minimum $5,000 | Goal $8,000
Outside R1 reviewers and second checks: Minimum $3,500 | Goal $5,000
Alexandria Arabic test: help with the case, prepare the evidence, and remove private details: Minimum $1,000 | Goal $1,000
Outside check of the study and a full run by someone else: Minimum $750 | Goal $1,500
AI model access, saved records, file storage, and online archive: Minimum $750 | Goal $1,000
Public data, notes, English-Arabic kit, final check, and sharing: Minimum $1,000 | Goal $1,500
TOTAL: Minimum $12,000 | Goal $18,000
I plan to take this as a direct Manifund grant. Echoes of the Script also has a fiscal sponsor, Fractured Atlas, if Manifund needs that route. I will list any fee or budget change before I take the money.
You can support one part of the work
A donor can use Manifund to support the Arabic test, outside review, a full repeat by someone else, or one more public case. Researchers can also suggest a small case or volunteer to review one.
Every case follows the same posted rules. Money does not buy a Green result. A sponsor cannot pick the answer, change the sources after the test starts, or stop me from posting a bad result.
Who is on your team? What's your track record on similar projects?
Team
Michael Grasa - project lead and creator of the sheet
I made the Falsifiability Sheet and will lead the study: set the rules, prepare cases, run the AI sessions, compare results, and post the final record. Emergent Ventures gave me about $12,000 for earlier work, and I shared an early form of this work at the Archaeological Survey of India in New Delhi in September 2025. Version 6A is a full public example. I am not the expert who decides the Egyptian answers; trained reviewers will do that part, and disagreements will be shown.
Mahmoud El Tahan - Alexandria research partner
Mahmoud is an Egyptian master's researcher in Egyptology at Alexandria University who studies Late Egyptian hieratic and hieroglyphic writing. He has a bachelor's degree in archaeology and about five years of experience as a senior archaeologist in Alexandria. He will try the Arabic sheet on one small part of his dissertation work and note what is clear, what is not, and where the evidence runs out. He is helping as an individual researcher, not speaking for Alexandria University, and he will not be the only R1 reviewer for his own case.
Outside reviewers
The grant pays for outside R1 reviewers who know the subject and for one person to check how I ran the study. Reviewers will not be told the expected result. I will name each reviewer, list relevant background and conflicts, and say exactly what each person did.
What are the most likely causes and outcomes if this project fails?
What could go wrong, and what I will do
A Red or STOP case turns Green, or answers keep changing: I will post the failure and fix or stop using the sheet. I will not hide the problem by averaging answers.
My own choices could push the result: Claims, files, expected answers, prompts, and rules will be set and posted before the study. Reviewers will not see the expected answers.
The Alexandria case is delayed: I will use the public backup case and say clearly that the real-use test did not happen.
AI models change: I will save the model name, date, settings, files, prompts, and full answers. I will not claim that a private model can be copied exactly years from now.
I made the sheet, so I could be biased: Outside reviewers will check the cases, another person will check how I ran the study, and if the goal is reached, someone outside the project will run a full case.
Few people use it: I will not pretend the sheet has been adopted. I will post the English-Arabic kit and study notes so others can decide whether another test is worth doing.
How much money have you raised in the last 12 months, and from where?
Prior funding
Emergent Ventures gave me about $12,000 during the last 12 months. That money paid for earlier work on the Indus script and the Falsifiability Sheet. The award was announced on April 18, 2026. It did not pay for this new study. I have raised no money yet for the Alexandria test.
Public links and proof
Alexandria University AI research guide announcement
Alexandria University research-infrastructure announcement
Mahmoud El Tahan — public professional profile
British Museum EA1006 — example source object
Emergent Ventures award announcement
Fiscal sponsor page, if routing is needed
Important note
I am not claiming that Alexandria University backs this project. The university links only show what it has said in public. Mahmoud is helping as an individual researcher. Before the study starts, a professional Arabic proofreader will check the sheet, link, and QR code.
There are no bids on this project.