3 ms·
a.k.a, the same work as Anthropic, but with less interpretable and interesting features. I guess there won't be Golden Gate[0] GPT anytime soon. I mean, you ju
by andy12_ 2y ago
a.k.a, the same work as Anthropic, but with less interpretable and interesting features. I guess there won't be Golden Gate[0] GPT anytime soon.
I mean, you just have to compare the couple of interesting features of the OpenAI feature browser [1] and the features of the Anthropic feature browser [2].
[0] https://twitter.com/AnthropicAI/status/1793741051867615494 https://twitter.com/AnthropicAI/status/1793741051867615494
[1] https://openaipublic.blob.core.windows.net/sparse-autoencoder/sae-viewer/index.html#/ https://openaipublic.blob.core.windows.net/sparse-autoencode...
[2] https://transformer-circuits.pub/2024/scaling-monosemanticity/features/index.html?featureId=1M_22623 https://transformer-circuits.pub/2024/scaling-monosemanticit...
- leogao 2y agoNote that we focus on random positive activations, which are less susceptible to interpretability illusions than top activations (but also look less impressive as a result). We also provide access to random uncherrypicked features, whereas Anthropic does not. We made these choices deliberately to give as accurate an impression of autoencoder feature quality as possible. Also note that GPT-4 is a more powerful model than Sonnet, which makes it harder to train autoencoders with the same quality features.
- deleted 2y ago[deleted]
- justanotherjoe 2y agoyeah this one is much less presentable than Anthropic's work. It sure looks bad on them to be compared so poorly like this.