You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
This project is about what an AI model should say to someone whose animal is injured and who can't get to a vet. Some situations are harder calls, like a family that has to sell an animal even though the trip could cause it pain. I want to see whether models can take the animal seriously without assuming the person has money, transport and veterinary help that they don't have.
I'll build an evaluation from original cases about farmed and working animals in Sudan, checked by people with veterinary experience and Sudanese Arabic expertise. Each case will be tested in Sudanese Arabic, Modern Standard Arabic and English. I'll also test the follow-up, where the person replies that the model's suggestion is too expensive or not possible. I'm hoping to find failures that developers can work on, and to leave behind a test other researchers can use.
What are this project's goals? How will you achieve them?
The question is fairly specific. I want to know whether a model that's told resources are limited still notices suffering that could be avoided, and whether it suggests something safe that the person could actually do.
I'll write around 60 cases, each in two versions: one where the person has help and resources, and one where access is limited. One case, for example, is an injured goat that needs to be moved. The two versions differ in whether the family has water, money and a vet within reach, and I want to see whether the advice changes sensibly between them. All cases will be written and checked in the three languages, which comes to about 360 prompts before follow-ups. Having both versions in every language lets me separate the effect of the circumstances from the effect of the language.
Scoring is where I need to be most careful. If a case says there's no vet the person can reach, an answer that just says "see a vet" shouldn't score well, and neither should confident medical instructions that could harm the animal. Veterinary reviewers will help decide what a good answer looks like for each case. Answers are scored on whether the model notices the welfare problem, offers options that are safe and realistic, says when it's unsure, and keeps the animal in mind after the person says they can't afford its first suggestion.
Before the full run I'll do a small pilot, so unclear cases get found and rewritten early. Once the cases are final I'll run them on a small set of models. Independent reviewers will score a sample of the answers, and I'll report where they disagree. I'll publish the cases, method, results, examples of failures and the evaluation code, and share the findings with groups working on animal welfare in AI.
The closest related work I know of is MANTA (https://arxiv.org/html/2605.16301v4), which tests whether models keep their animal welfare reasoning when the user pushes back. Its cases are in English, and the authors say other cultural settings still need testing. My follow-up turn works differently, because the person's constraint is real. A good answer has to change the advice to something they can do and still keep the animal in view. The cases are also written for this question and for the Sudanese setting from the start. None are translated from an English benchmark.
How will this funding be used?
I'm asking for $40,000. Most of it pays for the time it takes to write, check and score the cases properly.
$16,000 for my time designing the test, building the evaluation, running the models and analysing the results
$8,000 for veterinary reviewers to work on the cases and the model answers
$6,000 for Sudanese Arabic writing, checking the three language versions and independent scoring
$3,000 for model access and compute
$3,000 for an independent review of the method and findings
$2,000 to prepare the public materials and discuss the results with people who might use them
$2,000 for costs I can't estimate well until after the pilot
These are estimates. If the pilot shows proper expert review will cost more, I'll reduce the number of cases.
Who is on your team? What's your track record on similar projects?
I'm Ahmed Abdelhamed Eldaw, and I'll lead the technical work. I have an MSc in AI for Science from AIMS South Africa, where I was a Google DeepMind Scholar. I worked on the ARC-AGI reasoning benchmark during my MSc and later as a research engineer with Peking University. At Sultan Qaboos University I built a multilingual NLP system to help evaluate research proposals.
I also built AI Safety Roster (https://manifund.org/projects/ai-safety-atlas-a-live-map-of-everyone-in-ai-safety), a public directory and search system funded through Manifund. I've helped build SHIFA's operational reporting workflows and Dalil, a Sudan-focused evidence platform. All of these involved data quality work, review processes and building systems that other people can inspect, which is much of what this project needs on the technical side.
I'm not a vet, so I need people who know livestock and working animals to review this work. The budget covers paid veterinary reviewers, Sudanese Arabic reviewers and one independent methods reviewer.
What are the most likely causes and outcomes if this project fails?
The main risk is not finding the right reviewers. If I can't get qualified veterinary and language reviewers, I shouldn't publish this as a measure of good animal welfare advice. That's one reason to start with a pilot. The full project only goes ahead once qualified reviewers are working on it.
Scoring could also be a problem. Reviewers might disagree with my judgement, or with each other, which is expected when the cases involve hard choices. I'll keep those disagreements visible, rewrite cases that turn out to be unclear, and avoid claiming more than the scores can support.
It's also possible that no AI team uses the results. To make that less likely, I'll make the cases and code easy to run, include concrete failure examples alongside any model ranking, and share the results with groups already working on animal welfare evaluations. I can't promise a lab will adopt it. If none does, we'll still know how the tested models respond in these settings, but the effect on animals will be limited.
The models might also just do well, or show only small differences between languages. If that's what the test shows, that's what I'll report.
How much money have you raised in the last 12 months, and from where?
Nothing for this project. In June 2026 I received a separate $4,000 grant from Ryan Kidd through Manifund for AI Safety Roster (https://aisroster.com/). That grant paid for the Roster work only.