Manifund foxManifund
Home
Login
About
People
Categories
Newsletter
HomeAboutPeopleCategoriesLoginCreate
🌳
🌳
Aditya Joshi

@ajskateboarder

$0total balance
$0charity balance
$0cash balance

$0 in pending offers

Projects

Mechanistic estimation of highly out-of-distribution behaviors

Comments

Mechanistic estimation of highly out-of-distribution behaviors
🌳

Aditya Joshi

about 13 hours ago

@ajskateboarder (the DCT post notes that the output directions are likely not features on their own, however i believe this won't be necessary in order to factor DCTs)

Mechanistic estimation of highly out-of-distribution behaviors
🌳

Aditya Joshi

about 13 hours ago
Progress update

What progress have you made since your last update?


One way I tried going about the problem was investigating the structure of deep causal transcoding, which shows strong generalization abilities for language models (i.e. finding many interesting generalizations given less sampling). The post gives partial evidence for the 'persistent shallow circuits' hypothesis, which if true could considerably alter the difficulty of calculating expectations. It could make the problem simpler since shallow circuits are comparatively simpler structures to locate, and to improve on Pareto frontier for estimating generalizations, all that roughly needs to be done is to enumerate the circuits on the DCT. It could also be true that LLMs and other nets do mechanistically implement generalizations through distributed shallow circuits, and we can't rely on the structure of deep sequential circuits to make the problem easier (and the problem would be harder if this was false). Overall, being able to factor the structure of the DCT seems important for making better estimations.

It seems hard to extract shallow causal structure from DCTs where many features interact to activate one feature, and even harder vice-versa; most of these circuits are representative of the DCT, but not the original LLM/transformer/net. However, it does seem easier to extract circuits that map from one input feature to an output feature in a highly data-efficient manner.

https://www.lesswrong.com/posts/8iG7orf6q8hNHT4Xh/ajskateboarder-s-shortform?commentId=XNcbxWHhqYHpMzXXr

What are your next steps?

Continue looking into how to best factor DCTs into ensembles of shallow circuits, and how 'surprising' the generalizations these circuits encode are in context of other interp methods (hopefully by relying on some scaling laws/theory/well-reasoned assumptions)

Is there anything others could help you with?

I think the DCT method has broadly not been used very much, and it would be good if people looked more into how they work/generalize (the original authors seem to have made a DCT follow-up in weight-space https://arxiv.org/pdf/2606.29604)

Mechanistic estimation of highly out-of-distribution behaviors
🌳

Aditya Joshi

3 months ago

(is it ok to consider a broader scope as I outlined in the above doc? this was also the initial motivation for doing this work in particular. I think it would be a better use of the grant and ofc I still plan to work on the ideas I described originally)

Transactions

ForDateTypeAmount
Manifund Bankabout 2 months agowithdraw6800
Manifund Bankabout 2 months agowithdraw900
Manifund Bank2 months agowithdraw300
Manifund Bank2 months agowithdraw2000
Mechanistic estimation of highly out-of-distribution behaviors2 months agoproject donation+10000