You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
The open-weight LLMs that people and enterprises deploy are typically quantized, because that allows for faster inference on cheaper hardware. But most evaluations of model safety, ethics, and values use non-quantized base models. Moreover, many open models have political biases or safety guardrails built into them; in practice they are often fine-tuned to remove the biases or guardrails. We don't have good evidence on the final safety, ethics, and values of such fine-tuned models.
This project seeks to understand what happens to the safety, ethics, and values of LLMs when they are quantized or fine-tuned. Are built-in guardrails fragile to quantization? Does fine-tuning to remove guardrails also shift a model's stated values or observed moral decision-making? Preliminary findings at safetyevidence.org/research/abliteration-ugi-values/ suggest that the political values of models shift (in concerning directions) when models are fine-tuned to remove refusals. This is consistent with prior research on how e.g. fine-tuning on insecure code can make a model generally misaligned.
It is possible that open models will be the first truly rogue autonomous agents. We need a lot more evidence on how the quantization and fine-tuning that people do all the time (check out how popular quantized and "abliterated" models are on huggingface!) affect LLM safety, ethics, and values.
Beyond just a research project, we will also provide real-time observability into the safety, ethics, and values of the quantized and finetuned models that people are releasing on huggingface. This will be done on safetyevidence.org and will allow users to make more informed choices about which versions of models to use.
The project will be done by running evals of safety, ethics, and values on quantized and fine-tuned open models. This is fairly straightforward: there are already evals for safety, ethics, and values, it's just that they're not typically run on the plethora of quantized and fine-tuned open models.
(If we find that quantization worsens model safety, ethics, and values, or fine-tuning for refusal removal significantly worsens model ethics and values in other domains, we may also seek to develop quantization and fine-tuning methods that do not worsen the model on these axes and can work on the hardware that hobbyists use.)
Funding goes towards hardware to run LLMs and perform the proposed work.
Funding will go towards hardware for LLM inference and fine-tuning. Exact hardware purchased depends on availability at the time that funding is received. At $5000, a platform such as one of the following will be purchased:
DGX Spark with 128GB unified memory (currently $4300-$5400 including tax)
AI MAX+ 395 / Strix Halo with 128GB unified memory (currently ~$4000 including tax)
Comparable offerings from Apple
At $10000-$12000, we will link two of the above to increase tokens/sec and increase the size of model we can run.
Justification for hardware:
We will have this project, and possibly other related open LLM evaluation projects, run continuously on the hardware, evaluating new models, quantizations, and fine-tunes as they are released. Hardware will have high utilization for AI safety work and will not sit idle.
Having a persistent workspace, as opposed to ephemeral cloud GPUs, will allow for faster iteration on the research and more efficient use of my time on this project.
If some funding remains after hardware purchases, it will be used for API credits and GPU hours to evaluate LLMs that are larger than can be run locally.
No funding will go towards my time and salary, those are already covered.
Anthony Ozerov: PhD student in statistics at UC Berkeley. I have published research in several fields and am moving into AI safety. In AI safety, I maintain safetyevidence.org. An example of the sort of analyses enabled by this project is at safetyevidence.org/research/abliteration-ugi-values/, which I performed using very limited existing evals of "abliterated" (guardrail-removed) models.
Two possible reasons:
There really is no good signal, and we cannot detect if/how fine-tuning and quantization changes safety, ethics, and values. This could happen if existing evals give extremely noisy or contradictory results.
Something becomes obsolete or irrelevant before the project is complete (e.g. the current paradigm of evaluating a static released LLM is no longer meaningful, or we reach AGI and smaller open models are not super relevant in the world, or something else.).
Either of these would be a failure for the direct goals and approach proposed in this project. The former failure could still be helpful as it would be useful information for the community that could spur development of new evals. The latter failure would mean the project has little real-world impact, but could still be academically interesting, and anyway the hardware would still be useful for other AI safety research.
Raised $2000 from BlueDot Impact to fund eval runs for safetyevidence.org.
There are no bids on this project.