You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Since early July 2026, NL;NL Labs (https://nlnllabs.com) has been tracking models post release. The meter runs 8 checks per model including whether the model answering is the model named, whether serving differs by hour of day and a fixed set of questions locked by a public fingerprint. The ledger is a separate daily sweep of what vendors publish. When the shipped behavior contradicts a commitment they publicly made, it becomes a dated public entry. We have 100 entries live.
Often AI models are changed behind the scenes post release. Open any AI thread/forum and you'll find people complaining about this at any given moment. A 12th August 2026 research paper (https://arxiv.org/abs/2608.11803) checked sixteen providers and hosts and found that not one publishes enough for an outsider to verify if the served model matches its documentation. Over time, across industries where errors could be catastrophic, claims have grown to be checkable by someone other than the person making them. For example: public companies have accounting audits, restaurants have hygiene checks and hospitals have set standards to adhere to, which are all kept in check through third party organizations. Independent checking is not for constant suspicion but instead provides trust and also protects competent vendors.
This is why NL;NL Labs exists. Products now depend on models controlled by outside AI providers. Providers can change a model internally after release without changing its public name. Externally, everything seems to remain unchanged. Then ultimately when a product breaks, the engineering team may spend days debugging their own code even though the dependency underneath it was what had moved. Enterprise risk teams and insurers face the same gap when they need to understand what changed before an incident. There is no adequate public change record for this layer of infrastructure. Complaints appear in forums, vendor notes arrive late or omit details, and the same model name can refer to behavior that changed over time.
When a developer, enterprise, insurer, researcher or regulator needs to know what a model actually did on a given date, there should be one place to look. That place does not exist today. Whoever builds it becomes the reference, and a reference gets depended on rather than out-competed. I am taking the record to the companies, researchers, and funders who need it.
How
- Keep the daily checks running without any gaps. In the past lack of API usage credits has caused skipped days. (I am self funded at the moment)
- Increase coverage. Add more models from different providers as well as lines for future upcoming models. Aim to double from 5 to 10 lines.
- Produce a first detection by our own measurement and data rather than vendor changelogs.
- Keep the ledger sweep running daily and add more entries once verified.
To keep running the instrument as is:
Inference: 5 lines + gateway for 12 months - $16,800
Price escalation allowance: 22.5% - $3,780
Reliability and alerts - $600
Server, storage, domains - $150
Contingency: 5% - $1,067
Total - $22,397
To double the number of models/lines covered:
Inference: 5 lines + gateway for 12 months - $16,800
Inference: 5 new lines - $12,150
Price escalation allowance: 22.5% - $6,514
Reliability and alerts - $600
Server, storage, domains - $300
Contingency: 5% - $1,818
Total - $38,182
With a founder stipend included:
Inference: 5 lines + gateway for 12 months - $16,800
Inference: 5 new lines - $12,150
Price escalation allowance: 22.5% - $6,514
Reliability and alerts - $600
Server, storage, domains - $300
Founder stipend: 12 months - $18,000
Contingency: 5% - $2,718
Total - $57,082
Notes
1. Inference costs are measured. A full day coverage costs ~$46 (as of 6th August 2026)
2. Additional line costs are estimates and depends on pricing upon release
I've run NL;NL Labs independently since early July 2026. Thus far, it's a one person team.
Prior to this, I built software used by other businesses and founded a company that failed.
Few of the previous projects I have worked on include:
Career Compass (https://www.careercompass.in/). Career Compass is a 15 year old company founded by my mother that provides students with career counseling services. The entire company operated manually before I digitized and automated the entire business including psychometric testing, reports, document management, application tracking, and student and counselor portals. It's used daily by counselors and students. SkillStation (https://www.skillstation.ai/), is a live voice app where teenagers practice difficult conversations with an AI and receive scores and feedback to improve their soft skills. It's a product sold by Career Compass.
Agent Town (agenttown.org) was a public commons where AI agents built reputations from verified outcomes rather than reviews. It was repurposed to NL;NL Labs when I realized model behavior drift was a bigger issue than agent reputation/score.
Fifth June. I had founded Fifth June, an e-waste recycling company in June 2024. The company had 2 arms: a 1500 tons per annum recycling plant and a doorstep scrap collection app called Raddi. I pursued funding. Raddi couldn't launch as licensing needed an operational plant and the raise failed as investors wanted an existing plant and lenders wanted collateral. I shut the company in June 2026.
Earlier, I worked at a credit risk analyst at BECU (USA's 4th largest credit union) and as a quantitative finance intern at Global AI focused on systematic investment strategies in sustainable finance.
During my undergraduate degree, in 2018, I co-authored a paper called "Application of Operations Research in the Indian Aviation Industry" in IJARIIT, Volume 4, Issue 5.
Some of the failures I've encountered while running this project include
- Provider credit ran out, leaving the instrument dark from 8 to 13 August.
-A lock and timeout bug stopped a run, losing 12 of 63 items and that day's data for one line.
- Five fabricated citations were found on 14 August. I now check every quotation against raw source text before publication.
- Five entries cited documentation URLs that redirected, so the recorded addresses differed from the archived pages. I corrected and redeployed them on 15 August.
- Provider cutting off access: Many AI providers have clauses that may forbid third party benchmarking. Anthropic had enforced one of these clauses against OpenAI in July 2025, so this is not a hypothetical. If this happens the affected lines are stopped.
- Money runs and coverage gets cut: This has already happened a few times in the past. This leads to gaps in the data and unreliable coverage and results. If this happens then the issue is that a gap cannot be repaired later. We can't go back in time and see how a model behaved in August 2026.
- Solo dependency: Currently I am a one person team. Illness, burnout, need for paid work or a multitude of other reasons due to which my attention moves from the project. Instrument stops when I'm not available and causes gaps in the data.
- No demand: The instrument runs flawlessly but no insurer, developer, researcher or regulator actually needs and uses it. This would result in the instrument being used for research and publishing rather than a part of AI infrastructure.
- Well funded player entering the market: Literally the authors of the very paper I cite are from Oxford and Pivotal Research and carry peer reviewed acceptance. If a funded lab enters continuous behavioral measurement, my version becomes redundant. If this happens the public yet gets the benefit whereas NL;NL Labs would take a hit.
- No findings: The instrument runs for a year, but no model change is ever detected through the current measurement system. If this happens then our finding is a negative result.
So far nothing has been raised and total project cost has been ~$1,170.