4 ms·
I don't think the comparison is valid. Releasing code and weights for an architecture that is widely known is a lot different than releasing research about an a
by root_axis 10mo ago
I don't think the comparison is valid. Releasing code and weights for an architecture that is widely known is a lot different than releasing research about an architecture that could mitigate fundamental problems that are common to all LLM products.