4 ms·
There’s honestly so much interesting stuff here, esp. the llm-related things - large concept models (operating on and predicting concepts, not tokens), dynamic
by cube2222 2y ago
There’s honestly so much interesting stuff here, esp. the llm-related things - large concept models (operating on and predicting concepts, not tokens), dynamic byte latent transformers (byte-level alternative to standard tokenization), sparse memory layers (successfully scaling key-value memory layers without an increase in computational requirements).
Here they are presented as separate things, each of which apparently improves quality / efficiency. I wonder what the quality / efficiency increase is of all those methods put together? Maybe that’s what Llama 4 will be?
This looks like a lot of innovation is happening at Meta in those areas, really cool!
- janeway 2y agoSide track, but does anyone have suggestions about how to better present such content. I am struggling with similar docs/demos. As a documentation page, each section is laid out uniformly with section heading, content, link to code and link to paper. However the page itself is a blog post which will be difficult to find again next year. Are there other examples of companies having well presented technical summaries which remain findable from the hime page?
- airstrike 2y agoI'd put a table of contents-like page up front with some exciting short description of each section and use hyperlinks, allowing the user to navigate to the section and back
- ms8 2y agoI hope that Llama 4 or 5 will have a different architecture. All released llamas are +/- same inference with a better training pipeline. The downside is that llamacpp will probably not be able to run new models and maybe it will be too much big rewrite, so we will need new c,cpp,go,rust programs.
- rmbyrro 2y agoit's a bit ironic that Meta ended up becoming the largest "open ai" org. all right, yeah, it's not "open source", but hey, it is open to use and they're publishing their research openly as well.