@ajskateboarder (the DCT post notes that the output directions are likely not features on their own, however i believe this won't be necessary in order to factor DCTs)
@ajskateboarder
$0 in pending offers
Aditya Joshi
about 13 hours ago@ajskateboarder (the DCT post notes that the output directions are likely not features on their own, however i believe this won't be necessary in order to factor DCTs)
Aditya Joshi
about 13 hours ago
One way I tried going about the problem was investigating the structure of deep causal transcoding, which shows strong generalization abilities for language models (i.e. finding many interesting generalizations given less sampling). The post gives partial evidence for the 'persistent shallow circuits' hypothesis, which if true could considerably alter the difficulty of calculating expectations. It could make the problem simpler since shallow circuits are comparatively simpler structures to locate, and to improve on Pareto frontier for estimating generalizations, all that roughly needs to be done is to enumerate the circuits on the DCT. It could also be true that LLMs and other nets do mechanistically implement generalizations through distributed shallow circuits, and we can't rely on the structure of deep sequential circuits to make the problem easier (and the problem would be harder if this was false). Overall, being able to factor the structure of the DCT seems important for making better estimations.
It seems hard to extract shallow causal structure from DCTs where many features interact to activate one feature, and even harder vice-versa; most of these circuits are representative of the DCT, but not the original LLM/transformer/net. However, it does seem easier to extract circuits that map from one input feature to an output feature in a highly data-efficient manner.
Continue looking into how to best factor DCTs into ensembles of shallow circuits, and how 'surprising' the generalizations these circuits encode are in context of other interp methods (hopefully by relying on some scaling laws/theory/well-reasoned assumptions)
I think the DCT method has broadly not been used very much, and it would be good if people looked more into how they work/generalize (the original authors seem to have made a DCT follow-up in weight-space https://arxiv.org/pdf/2606.29604)
Aditya Joshi
3 months ago(is it ok to consider a broader scope as I outlined in the above doc? this was also the initial motivation for doing this work in particular. I think it would be a better use of the grant and ofc I still plan to work on the ideas I described originally)
| For | Date | Type | Amount |
|---|---|---|---|
| Manifund Bank | about 2 months ago | withdraw | 6800 |
| Manifund Bank | about 2 months ago | withdraw | 900 |
| Manifund Bank | 2 months ago | withdraw | 300 |
| Manifund Bank | 2 months ago | withdraw | 2000 |
| Mechanistic estimation of highly out-of-distribution behaviors | 2 months ago | project donation | +10000 |