You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project Description
Over the last ten months, I have been developing Aelira, and I have funded my development of Aelira with my own money. As a full-time caregiver and owner of a small IT consulting firm, I have managed to continue working on Aelira, but I am now at the limit of my ability to fund this project.
I am seeking up to $20,000 to continue developing the free and open source core of Aelira, and to conduct an in-depth analysis of the potential for AI to aid in accessibility remediation.
The immediate requirement is for funded time and independent review. This will enable me to convert the work I have already completed into something that others can evaluate and utilize. The area that I wish to examine is whether a model maintains the meaning of the original document.
For example, a description of a graph may appear to provide an accurate representation of the data, however, the graph may illustrate a trend that is inaccurate. Similarly, repairing a table may result in altering a value. These are two examples of the types of failures that the study will investigate; however, the frequency of these failures is yet to be determined. A student utilizing a screen reader requires the information provided within the document to be maintained during the repair process. The successful completion of an automated check is only a portion of achieving this. As explained by the W3C's Evaluation Guidelines (https://www.w3.org/WAI/test-evaluate/tools/selecting/), human judgment continues to be required.
The grant will cover the expenses associated with creating a publicly available test set, creating software for evaluating the results, hiring independent reviewers, and improving the open-source core. I will publish the results, including failures and limitations. The results will be available for individuals to independently assess, without reliance upon Aelira's paid hosting services.
What are the objectives of this project? How do you intend to achieve these objectives?
My objective is to determine where AI can contribute positively to this area, and where the software needs to pause for a person to review it. The primary objective is to maintain the source information.
At full funding, my goal is to assess 150 cases, including 30 STEM cases reserved for the final assessment. I will utilize either newly authored content or previously authorized content. The cases will consist of images, graphs, and tables, as well as omitted evidence and instructions included within source files which could potentially confuse a model. This can be accomplished with content authorized for public use, while excluding confidential university documents and student records from the scope of this study.
I will compare three methods: normal prompting; prompting that specifically mandates preservation of the meaning; and a workflow that verifies the source evidence and has the capability of pausing for human review.
Utilizing four model configurations and three repetitions per method will produce approximately 5,400 outputs. Initially, I will verify compatibility and costs. Once verified, I will retain the configurations for comparison. The verifications may minimize errors, or they may primarily lead to increased refusals. Alternatively, they may generate additional expenses and review efforts for minimal benefit.
Collectively, I will quantify the following outcomes: invented or erroneous assertions, loss of information, successful task completion, refusals, review duration, response duration, and costs. A system that declines most tasks would possess minimal practical applicability.
A sample of 200 outputs selected prior to commencement of the study will be evaluated by independent reviewers. A second reviewer will score 50 of those outputs to establish inter-rater reliability. Reviewers will receive access to the source material with the model and method identifiers concealed. Each reviewer will evaluate whether the output preserves the source independently of whether it is beneficial for accessibility.
Prior to examining the final outputs, I will define the sampling and scoring criteria and consider repeated runs of similar cases in my analysis.
The twelve-week timeframe will commence with two weeks dedicated to finalizing the scoring criteria, permission requirements for cases, and realistic workload expectations for reviewers.
- Weeks 3-5 will be utilized for generating cases and developing evaluation software.
- Weeks 6-9 will be utilized for conducting runs and performing independent reviews.
- Weeks 10-12 will be utilized for analyzing data and making public disclosures.
A small pilot study will evaluate the workload prior to conducting the full-scale study. Any modifications to the scope of materials will be agreed upon prior to initiating development.
I will publish the stand-alone evaluation software under MIT license, newly authored cases and labels under CC BY 4.0 license when applicable, and enhancements to the open-source core under the same license as the open-source core. The publication will include model versions, configurations, API costs, and procedures for reproducing the work.
All of these outputs may be generated using existing models. This is an exploratory study. Conclusions drawn from this study will be limited to cases and configurations tested. Additional testing will be necessary prior to making broad statements regarding safety or accessibility.
How will this funding be used?
Funding provided will enable me to devote time to research and involve individuals who can independently verify my work.
I have previously financed the purchase of a Founders Edition DGX Spark for testing and benchmarking. This hardware is available for use in this project.
All monetary values presented below are in United States Dollars. Hourly rates are estimated; I will confirm availability of reviewers and obtain quotes prior to defining the scope of the study.
Minimum Funding Level - $5,000
- $2,000: Research & Engineering (50 hours @ $40/hour).
- $2,000: Independent Accessibility & Technical Review (20 hours @ $100/hour).
- $500: Model APIs & Supplementary Evaluation Compute.
- $250: Project Hosting, Storage & Research Artifacts.
- $250: Independent Release & Reproducibility Review (2.5 hours @ $100/hour).
Total: $5,000.
Intermediate Funding Level - $10,000
- $4,500: Research & Engineering (112.5 hours @ $40/hour).
- $3,500: Independent Accessibility & Technical Review (35 hours @ $100/hour).
- $1,000: Model APIs & Supplementary Evaluation Compute.
- $500: Project Hosting, Storage & Research Artifacts.
- $500: Independent Release & Reproducibility Review (5 hours @ $100/hour).
Total: $10,000.
Full Funding Goal - $20,000
- $9,000: Research & Engineering (225 hours @ $40/hour).
- $7,000: Independent Accessibility & Technical Review (70 hours @ $100/hour).
- $2,000: Model APIs & Supplementary Evaluation Compute.
- $1,000: Project Hosting, Storage & Research Artifacts.
- $1,000: Independent Release & Reproducibility Review (10 hours @ $100/hour).
Total: $20,000.
My hours will include preparation of cases, development of evaluation software, execution of experiments, analysis of results, modification of the core, and documentation. Independent release review is external work separate from my own hours. Reviewers will be compensated for their work regardless of what they find.
A feasibility study lasting six weeks will be funded with $5,000. The feasibility study will include 20 image/chart cases, two model configurations, and two methods comparing ordinary output with verification and human-review escalation. Three runs per method will produce up to 240 outputs, with 40 first reviews and 10 second reviews. I will release the small test set, working evaluation software, and findings.
An eight-week study including 60 cases, three model configurations, and three methods will be funded with $10,000. Three runs per method will produce up to 1,620 outputs, with 100 first reviews and 25 second reviews.
The full twelve-week study described above will be funded with $20,000.
Additionally, I will receive approximately 19 project hours per week. If funding levels fall between these amounts, I will agree upon a scope that aligns with available funding.
Financial strain on my personal finances is significant. I require assistance to continue this work, and these budgets support future work on this project.
Financial assistance for past hardware costs will be discussed separately with Manifund and clearly articulated in an agreed-upon revised budget and scope.
Who is on your team? What is your history on similar projects?
I am Reginald Crampton, Founder, Project Lead, and Director of Aelira.
In addition to operating my IT consulting business, I have contracted with Outlier and CrowdGen on model-evaluation, red-teaming, and safety projects. My experience in these areas is relevant to analyzing model outputs and identifying areas where they fail.
My co-founder Erik Vuchich directs External Partnerships, Sales & Outreach. Both Erik and I reside in Australia. I will lead research, engineering, and reporting. Erik has indicated that he intends to assist in identifying independent reviewers and disseminating publicly available information through his outreach activities; we will define his role prior to commencing the study. This budget allocates funded research hours to me.
Open source code for Aelira is located at: https://github.com/Aelira-AI/aelira-core. The open source code includes document scanning, supported remediation workflows, and human-review workflows. The core is licensed under AGPL-3.0 and remains in beta validation. The open source code provides a practical foundation for this study, although testing reliability of generated content is still required.
I have a commercial interest in Aelira and will disclose that interest to reviewers. The outputs from this grant will be publicly disclosed, and I will hire reviewers who are unaffiliated with Aelira. Reviewers will be paid for their work irrespective of what they find. Recruitment of reviewers is still pending.
What are the greatest probabilities of failure for this project? What would be the consequences if this project fails?
My ability to complete this project is one of the greatest risks. I have obligations as a caregiver, consultant, and researcher requiring significant time commitments. Funding will allow me to allocate time to this project. Locating qualified independent reviewers may require more time than anticipated.
Additionally, accessing models or costs associated with evaluations may fluctuate.
It is possible that this study will produce limited results. Checks may be too costly, refuse too many tasks that are beneficial, or have negligible impact on reducing errors. Reviewers may disagree, and results from these cases may not translate well to other documents. Regardless, I will publish all limitations/negative results alongside any positive results derived from this study.
I will execute the workload pilot first. Then I will hold model configurations constant for comparative purposes. I will operate within budget constraints established for evaluation expenses. I will tie conclusions drawn from this study directly to evidence obtained.
Deliverables exist at the minimum funding level. Should I fail to complete agreed-upon work, I will notify Manifund and donors promptly and determine how to proceed with reduced scope/unspent funds.
If I cannot obtain grant funding, I will be compelled to allocate additional time to paid consulting work and decrease my commitment to research. Open source code will remain publicly accessible; however, an independently reviewed study will be delayed/halted. After spending nearly one year financing this project myself, external support will provide a significant boost to what I can deliver going forward.
How much money did you raise externally over the last 12 months?
Zero dollars in external funding. I funded all costs associated with this project including hardware purchases, hosting fees, model subscription fees, etc.
I also applied for a BlueDot Rapid Grant for related work. As of today's date, no funding has been awarded. If either application receives funding prior to spending funds requested in this proposal, I will notify Manifund/donors and reduce overlapping requests or establish separate additional milestones prior to spending.