8 ms·
Catching up on the weird world of LLMs
- d4rkp4ttern 3y agoGreat article. I wonder why this post (and others on this site) have no date.
- grandma_tea 3y agoThere is a date at the very bottom of the article. This one was posted August 3rd
- simonw 3y agoThat's a mobile design flaw: on desktop the date shows up at the top in the right hand sidebar but on mobile you have to scroll all the way to the bottom of the page. I should fix that, it annoys me too. UPDATE: Fixed. https://github.com/simonw/simonwillisonblog/commit/6db7db1740a4ae545164fc81180b42d32a0d8cea https://github.com/simonw/simonwillisonblog/commit/6db7db174...
- earth-adventure 3y agoIt's in the url (fyi)
- tomstuart 3y agoWhat is going on with this comment? It was clearly made a couple of days ago, and Simon’s reply was also posted a couple of days ago, and yet they’re both appearing on an article posted in the last day, and (on that post) showing as being a few hours old. Am I on glue?
- bavell 3y agoGood overview, recommended for anyone wanting a quick lay of the land.
- tayo42 3y agoThat seemed pretty good, assuming none of it was made up by an llm hah I was looking for something like this. Just started playing with some models, it's all pretty overwhelming and technical.
- dTal 3y ago>These things aren’t deterministic so it’s hard to even use things like trial-and-error experiments to figure out what works, which as a computer scientist I find completely infuriating! Nitpick: they can be made "deterministic" in a strict sense by using a deterministic sampling scheme, for instance by turning "temperature" down to 0. This isn't necessarily helpful for prompt experimentation, though - when deployed, prompts are usually combined with other input, so you risk simply substituting the non-determinism of sampling with the unpredictability of whether it will work with all inputs.
- soroushjp 3y agoIn my own experiments with OpenAI's GPT-4 API with temperature set to zero, I was still not getting deterministic outputs, with some small variations between completions. Not sure why, and I haven't had a chance to dig further or talk to their team about why and how this happens.
- mirekrusin 3y agoNon-determinism in GPT-4 is caused by Sparse MoE [0] [0] https://news.ycombinator.com/item?id=37006224 https://news.ycombinator.com/item?id=37006224
- deleted 3y ago[deleted]
- pillefitz 3y agoYou don't have to put the temperature to 0, you can just use a fixed random seed, couldn't you?
- K0balt 3y agoRather than viewing generative AI as a form of artificial intelligence, I posit that it should be seen as an automated tool for tapping into human cultural, linguistic, and empirical knowledge. Data and computation are two sides of the same coin. The 'intelligence' in AI is embedded within the data, with the computational model serving as a tool to access and express this inherent intelligence. I would argue for a change in perspective towards AI, one that recognizes LLMs as powerful tools for accessing the vast wealth of human cultural knowledge rather than viewing them as a separate form of intelligence. We must carefully consider critical ethical considerations about control, access, and trust that will become increasingly relevant as these tools become more integrated into our everyday lives. This paradigm shift carries some implications: LLMs will not achieve superintelligence: Although these models can process information quickly and access a wide range of knowledge, they lack the superior reasoning or inference abilities that would classify them as superintelligence. LLMs as an extension of human thought: These models can automate and amplify human capabilities but do not introduce new abilities beyond what is already present in human thought processes. LLMs as mirrors of human culture and knowledge: These models reflect the recorded artifacts of human language, art, and culture. They can make the inherent intelligence within these artifacts accessible, providing a vast information resource. Implications for the future: Access to this "memetic matrix" of human knowledge will become a fundamental part of being human as these tools become more integrated into our lives, bringing up issues of ownership, access, and the potential for misuse. Thought consolidation and control of inference engines: There's a potential risk that control of inference engines by a small number of companies could lead to a consolidation of thought that threatens democratic governance. I propose a diversity of federated or self-hosted inference tools as solutions to mitigate this risk. The necessity for trust and individuality: As these tools become more influential in our lives, maintaining trust in our individual thoughts and avoiding the uncritical acceptance of synthesized ideas from sources with opaque motives will become increasingly important. Synthetic Inference relies on a vast cultural commons: We cannot allow these commons to be closed off and owned by a few big companies. This resource is the totality of all human knowledge, language, and culture. It belongs to all of humanity. Training data must be open, free, and available for examination.
- hyperliner 3y ago
- sakopov 3y agoCan anyone give a layman's rundown on "it guesses the next token" in situations where it's seemingly able to apply what looks like logic to certain prompt requests?
- tomohelix 3y agoThe way I understand it, our language is an abstraction to our logics. We think using it and we communicate our logic with words. Literally, how do you reason within your head? By forming sentences and argue with yourself using words? At least that is how I do it. I can't make any deep reasoning without first translating it to sentences. So an AI after learning so much text and forming an extremely dense network of connections between different words and phrases, it can mimic something similar to reasoning. At its core, it is still just predicting the next word. But the scale of this "prediction" is so large that it begins to mirror our own "reasoning" process. Because in the end, when we apply logic, we are doing it through our own network of concept connections which is reflected in our language.
- sakopov 3y agoI guess the author did make a complete explanation when they mentioned arrays of floating point numbers representing connections. I just thought there was more to it that was omitted. This seems like a relatively simple solution which simulates an fascinatingly complicated process when performed at scale. Thanks for a great explanation!
- visarga 3y agoIt is applying logic very often, but in an implicit way. How would it solve unseen math and code problems otherwise.
- mcluck 3y agoDoes it solve unseen math? I can consistently get it to trip over even simple math
- stevenhuang 3y ago
- nologic01 3y ago> The fascinating thing is that capabilities of these models emerge at certain sizes and nobody knows why Is there a more technical dive into this statememt? In particular, is it an emergent statistical property of the model and/or the training data size or an emergent illusion of the human observer of model outputs (the way we visually perceive fluid motion once a frame rate exceeds a certain threshold). The nature of this "emergence" is interesting both from a theoretical and a practical point of view. The "large" model size (I put the word in quotes, because large compared to what equivalent task?) might be somehow intrinsic, in which case it will spark a race to make such sized computational resources more common place. But it might also be an imcomplete understanding of the model class.
- NhanH 3y agoThere is a paper saying that around 6-7 billion params, something happened that makes the larger transformer qualitative different than the smaller one. I forgot which paper it was though
- gwern 3y agoWei et al 2022 might be a good starting point: https://gwern.net/doc/ai/scaling/emergence/index#wei-et-al-2022-section https://gwern.net/doc/ai/scaling/emergence/index#wei-et-al-2...
- nologic01 3y agoGood reference. The sparsity of the data points (presumably due to the huge training costs) doesn't reveal too much at present but as the authors suggest the phenomenon might be reproduced in smaller models (which would make it easier to study in greater detail).