7 ms·
The paper is hard to read. There is no concrete worked-through example, the prose is over the top, and the equations don't really help. I can't make head or tai
by Scene_Cast2 1y ago
The paper is hard to read. There is no concrete worked-through example, the prose is over the top, and the equations don't really help. I can't make head or tail of this paper.
- joeldg 1y ago[dead]
- lumost 1y agoThis appears to be a position paper written by authors outside of their core field. The presentation of "the wall" is only through analogy to derivatives on the discrete values computer's operate in.
- joe_the_user 1y agoPaper seems to involve a series of analogies and equations. However, I think if the equations accepted, the "wall" is actually derived. The authors are computer scientists and people who work with large scale dynamic system. They aren't people who've actually produced an industry-scale LLM. However, I have to note that despite lots of practical progress in deep learning/transformers/etc systems, all the theory involved just analogies and equations of a similar sort, it's all alchemy and so people really good at producing these models seem to be using a bunch of effective rules of thumb and not any full or established models (despite books claiming to offer a mathematical foundation for enterprise, etc). Which is to say, "outside of core competence" doesn't mean as much as it would for medicine or something.
- ACCount37 1y agoNo, that's all the more reason to distrust major, unverified claims made by someone "outside of core competence". Applied demon summoning is ruled by empiricism and experimentation. The best summoners in the field are the ones who have a lot of practical experience and a sharp, honed intuition for the bizarre dynamics of the summoning process. And even those very summoners, specialists worth their weight in gold, are slaves to the experiment! Their novel ideas and methods and refinements still fail more often than they succeed! One of the first lessons you have to learn in the field is that of humility. That your "novel ideas" and "brilliant insights" are neither novel nor brilliant - and the only path to success lies through things small and testable, most of which do not survive the test. With that, can you trust the demon summoning knowledge of someone who has never drawn a summoning diagram?
- cwmoore 1y agoYour passions may have run away with you. https://news.ycombinator.com/item?id=45114753 https://news.ycombinator.com/item?id=45114753
- jibal 1y agoSomehow the game of telephone took us from "outside of their core field" (which wasn't true) to "outside of core competence" (which is grossly untrue). > One of the first lessons you have to learn in the field is that of humility. I suggest then that you make your statements less confidently.
- ForHackernews 1y agoThe freshly-summoned Gaap-5 was rumored to be the most accursed spirit ever witnessed by mankind, but so far it seems not dramatically more evil than previous demons, despite having been fed vastly more humans souls.
- lazide 1y agoPerhaps we’re reaching peak demon?
- lumost 1y agoI will venture my 2 cents, the equations kinda sorta look like something - but in no way approach a derivation of the wall. Specifically, I would have looked for a derivation which proved for one of/all of 1. Sequence Models relying on a markov chain, with and without summarization to extend beyond fixed length horizons. 2. All forms of attention mechanisms/dense layers. 3. A specific Transformer architecture. That there exists a limit on the representation or prediction powers of the model for tasks of all input/output token lengths or fixed size N input tokens/M output tokens. *Based On* a derived cost growth schedule for model size, data size, compute budgets. Separately, I would have expected a clear literature review of existing mathematical studies on LLM capabilities and limitations - for which there are *many*. Including studies that purport that Transformers can represent any program of finite pre-determined execution length.
- jibal 1y agoIf you look at their other papers, you will see that this is very much within their core field.
- lumost 1y agoTheir other papers are on simulation and applied chemistry. Where does their expertise in Machine Learning, or Large Language Models derive from? While it's not a requirement to have published in a field before publishing in a field. Having a coauthor who is from the target field or a peer review venue in that field as an entry point certainly raises credibility. From my limited claim to be in either Machine Learning or Large Language Models the paper does not appear to demonstrate what it claims. The author's language addresses the field of Machine Learning and LLM development as you would a young student - which does not help make their point.
- stonogo 1y agoIf you can't look at that publication list and see their expertise in macine learning, then it may be that they know more about your field than you know about theirs. Nothing wrong with that! Computational chemists use different terminology than computer scientists but there is significant overlap in the fields.
- JohnKemeny 1y agoHe's a chemist. Lots of chemists and physicists like to talk about computation without having any background in it. I'm not saying anything about the content, merely making a remark.
- chermi 1y agoYou're really not saying anything? Just a random remark with no bearing? Seth Lloyd, Wolpert, Landauer, Bennet, Fredkin, Feynman, Sejnowski, Hopfield, Zechinna, parisi,mezard, and zdebvora, Crutchfeld, Preskill, Deutsch, Manin, Szilard, MacKay.... I wish someone told them to shut up about computing. And I wouldn't dare claim von Neumann as merely a physicist, but that's where he was coming from. Oh and as much as I dislike him, Wolfram.
- deleted 1y ago[deleted]