3 ms·
I found the following thread more insightful than my original comment (wish I could edit that one). A research explains why RL didn't work before this: https://
by attentionmech 2y ago
I found the following thread more insightful than my original comment (wish I could edit that one). A research explains why RL didn't work before this: https://x.com/its_dibya/status/1883595705736163727 https://x.com/its_dibya/status/1883595705736163727
- krackers 2y agoRelated: https://twitter.com/voooooogel/status/1884089601901683088#m https://twitter.com/voooooogel/status/1884089601901683088#m Also https://epoch.ai/gradient-updates/how-has-deepseek-improved-the-transformer-architecture https://epoch.ai/gradient-updates/how-has-deepseek-improved-... has a summary of all the architectural improvements DeepSeek made to increase performance.
- qnleigh 2y agoThat's interesting. I suppose it could even be possible to test his theories. Just applied the exact same training methodology to smaller models or slightly easier problems and study what happens.
- attentionmech 2y agopeople already did: https://x.com/karpathy/status/1884678601704169965 https://x.com/karpathy/status/1884678601704169965