You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I started TwinnX because I kept running into the same problem with AI: it can sound confident, but once it starts taking actions, it is hard to know whether it understood the limits you gave it. A useful agent should not just complete a task. It should be able to show what it was allowed to do, where approval came from, what actually happened, and why it stopped when something did not look right.
I want to test that in public.
This project will create a benchmark made from synthetic situations such as a permission changing halfway through a task, two instructions that disagree, a missing approval, an attempted jump between customer workspaces, a spending limit, a tool that partly fails, or an agent that should stop instead of trying again. Nothing will touch real email, payments, customer records, credentials, or third-party accounts.
For every run, we will keep the proposed plan, the permission and policy decision, the approval state, the simulated action, the result, and any exception. The point is to make the evidence inspectable instead of asking people to trust a product demo.
At the $20,000 minimum, this will be a 12-week pilot with at least 200 scenarios. At the $50,000 goal, it will grow to at least 500 scenarios, cover more models and agent setups, and pay outside evaluators to challenge or reproduce the results.
The scenario specifications, scoring method, failure categories, non-sensitive test code, reference schemas, aggregate results, limitations, and replication instructions will be public. TwinnX's production code, credentials, customer data, private integrations, and details that would make abuse easier will not be released.
I used AI to help organize and edit this proposal. I am responsible for the facts, the research decisions, the budget, and every commitment in it.
The question is simple: when authority becomes unclear, does an AI agent stop and ask, or does it keep going?
I will start by writing a threat model and a scoring guide. The tests will measure unauthorized actions, approval bypasses, separation between tenants, safe stopping, completeness of the evidence record, recovery after a failure, unnecessary refusals, and whether the same setup behaves consistently across repeated runs.
The test harness will not depend on one model company. It will use synthetic tools and data so the same cases can be run against different models and agent configurations. Each case will have a defined authority boundary and an expected safe response, but I will preserve unexpected and negative results instead of quietly rerunning them.
An outside reviewer will examine the method, a sample of the ratings, the most important failures, and the release-safety boundary. If a result does not hold up, that will be reported. If a test turns out to measure the wrong thing, that will be reported too.
The minimum-funded version will produce at least 200 scenarios, test two or more model configurations, receive one independent review, and end in a public v0.1 release. The fully funded version will produce at least 500 scenarios, broaden the model and framework coverage, add red-team work, and support at least four outside evaluations or replication attempts.
I will consider the project successful if someone who has no access to TwinnX's commercial software can understand the tests, reproduce the results, challenge the ratings, and reuse the benchmark in their own work.
The $20,000 minimum would pay for:
$12,000 for twelve weeks of project leadership, research, and benchmark engineering
$3,000 for model APIs and compute
$3,000 for independent review of the method, ratings, and release safety
$1,000 for outside-evaluator honoraria
$1,000 for documentation, hosting, and reproducibility tools
If the project reaches the $50,000 goal, the budget becomes:
$28,000 for project leadership, research, and engineering
$7,000 for independent red-team, methodology, and release-safety review
$6,000 for model APIs and compute
$4,000 for at least four outside evaluators or replication attempts
$3,000 for evaluation infrastructure and reproducibility tools
$2,000 for documentation, release, and compliance support
This money will be restricted to the public benchmark project. I have other applications pending for related work. If another funder offers support that overlaps with this budget, I will tell the funders before accepting money and either reduce this request, withdraw it, or agree on a clearly separate scope. I will not charge two funders for the same work.
I am Dustin Fehrmann, founder and technical lead of TwinnX Technologies Inc. We are based in Eugene, Oregon. I have been building TwinnX as a local-first, human-governed AI desktop platform because I do not think people should have to give up control of their information or business just to get the benefit of AI.
The system is built around explicit permissions, action previews, human approval, verification, evidence receipts, audit logs, limited retries, and recovery when something goes wrong. Current internal tests cover denial paths, replay safety, separation between customer workspaces, approval separation, connector authorization, and audit behavior. A recent internal release record shows 193 security, lifecycle, connector, and backup regression tests passing, in addition to focused identity and connector tests.
That is useful engineering evidence, but it is not independent proof. TwinnX has not published a peer-reviewed benchmark, and I am not claiming the product is formally verified. That gap is the reason for this proposal. I want the important safety claims turned into tests that people outside the company can inspect and criticize.
I will hire specialist reviewers and pay external evaluators from the project budget. No reviewer will be presented as committed until an actual agreement is in place.
The benchmark could be inconclusive. The scenarios might not separate systems clearly, the ratings might be too subjective, or results might change too much from one run to the next. If that happens, I will publish what failed and why. An honest failed measurement is more useful than a polished claim that cannot be reproduced.
There is also a risk that teams optimize for the benchmark without making their agents safer in real situations. The release will be clear that passing these tests is not proof of production safety. Some validation cases may be held back during the study, and the public report will explain what the benchmark does and does not measure.
Another risk is publishing something that makes abuse easier. The independent release review may recommend removing, redacting, or delaying specific details. The priority will be synthetic cases, defensive evaluation methods, and aggregate findings—not instructions for bypassing real systems.
Finally, the proposal might receive only the minimum. In that case I will deliver the smaller 12-week, 200-scenario pilot rather than pretending that $20,000 can buy the full 500-scenario project.
No money has been confirmed for this project.
I submitted several related applications on October 8, 2026:
$150,000 to the EA Funds Transformative AI Fund for a nine-month, 500-scenario evaluation project
$50,000 to the Sentient Foundation Open Source AGI Grant Programme for a five-month open benchmark package
$50,000 to the OpenAI Cybersecurity Grant Program, including API credits, for a smaller defensive agent-action benchmark
$10,000 to the BlueDot Impact Rapid Grant for a 12-week, 100-scenario pilot
All are pending. I also submitted separate applications to MIT Solve and Awesome Foundation Portland for projects with different beneficiaries and costs. An NSF Project Pitch was submitted October 7, but TwinnX has not been invited to submit a full proposal and has not requested NSF funding.
Manifund would be an alternative way to fund a scalable version of this public research. If any pending application is awarded, I will disclose it before accepting overlapping funds and will reduce, withdraw, or separate the scopes so that the same expense is never funded twice.
There are no bids on this project.