You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Today, every organisation that makes the decision to deploy open-weight model typically has to conduct similar safety and security evaluations independently and repeatedly. The same process is repeated by individual developers and researchers as well. This leads to massive duplication: hundreds and thousands of organisations execute highly similar safety tests, if not same. This consumes compute, energy, researcher time, and engineering resources redundantly.
Duplicates make the work not only ineffective but also amplify the issue of the open-weight safety problem. When every org repeats private evaluations, nobody has a public, comparable, longitudinal picture. As a result, it becomes impossible to state if the open-weight safety has improved or deteriorated over time. And, no regulator or releaser has a shared factual baseline to act on. This matters the most because safety measured at release is not safety in deployment. Anyone can fine-tune safeguards away. All the existing evaluation efforts like METR, Apollo, AI Lab Watch, SaferAI and Midas project have dealt with frontier closed models, where the gap doesn't exist in the same form. No one is systematically tracking safety degradation under fine-tuning across open-weight releases. Requesting for the grant of $13,400 for three-four months to take OSMSI from one model to twenty+.
OSMSI will provide open-weight models with a Safety Card, and a public Safety Index along the six dimensions (Jailbreak resistance, Harmful capability, Prompt injection, Uncertainty calibration, Agentic safety, Interpretability aspects [exploration / v2]. The method needs further work). The goal of OSMSI is to develop a shared library with reproducible safety evaluations which would be useful to developers, organizations and policy-makers as a common baseline. This would enable organizations to then focus their resources on the evaluations that are genuinely specific to their deployment context.
OSMSI treats the evaluation of the safety of models as common infrastructure rather than a task that every organization should independently solve from scratch.
Months 1 - freeze the first version of the framework, publish the Safety Card format, harden the pipelines.
Months 2-3 - evaluate 10–15 open-source models over the five scored dimensions. Establish a held-out private test set alongside the public dataset.
Months 3-4 - publish the leaderboard and report the difference between public/private performance gap as a first-class metric, and run a fine-tuning degradation pilot on a subset of models.
Total requested: $13,400
Timeline: three-four months
No salary. Built alongside full-time employment.
Breakdown:
Compute, GPUs = $8000. Dedicated GPU infrastructure for evaluation runs
Compute, grader = $2000. Model APIs for grading
Storage, hosting, CI/CD = $1500. Artifact and log storage, leaderboard hosting, pipeline CI
Software and subscriptions = $500. Tooling and licences
Contingency = $1400. Model pricing changes, re-runs after methodology revisions
Coffee & Red Bull = $900 ($8 a day)
Total = $13,400
Netan Mangal | AI Research Engineer @ Secure Agentics (UK) | LinkedIn
Research Engineer at Secure Agentics, working on AI security and evaluation for enterprise deployments.
5+ YoE in large scale data-tech, applied AI, ex-crypto exchange.
OSMSI is an open-source project that builds directly on that practical experience evaluating foundation models for safety and deployment.
Contamination : As benchmarks become public, they leak into training data and scores drift upward without safety improving. Outcome: the index degrades into a memorisation test.
$2000 from BlueDot Impact Rapid Grant, July 2026.