5 ms·
Exactly my thoughts, this isn’t really an LLM. I understand it’s just a personal project but it’s probably good to know LLMs are not just lookup tables. Also,
by mpeg 2y ago
Exactly my thoughts, this isn’t really an LLM. I understand it’s just a personal project but it’s probably good to know LLMs are not just lookup tables.
Also, I’m not sure I understand the explanation of why js objects are so great for this, you could do the same thing with a python dict for example or equivalent mapping types in any programming language.
This is also missing an attention mechanism which is very important to understand even for a toy LLM as it’s part of what makes the responses so accurate to the context, rather than it just predicting what word usually comes next to another word (or sequence of words)
- selcuka 2y ago> you could do the same thing with a python dict for example or equivalent mapping types in any programming language. Yes, any language with a hashmap of sorts with O(1) lookups should be equivalent.
- _akhe 2y agoYes the time complexity is how it can all be held in memory without crashing (because it isn't really all held in memory, like it would be with text) but that's only part of it. You could not choose "any programming language" for it because not all of them are good at lookups, and not all languages have callback functionality or the ability to natively enqueue and fork code flow. It would take forever to do this and probably not work in most languages.
- littlestymaar 2y ago> but it’s probably good to know LLMs are not just lookup tables. But there are almost literally lookup tables actually! The thing is they don't perform the lookup on a single word as a key, but on the entire context, which is what makes them special. But besides that it's “just” a lookup table.
- kibibu 2y agoA deep attention network is absolutely not a lookup table
- holoduke 2y agoMultiple lookup tables
- mpeg 2y agoBit of an oversimplification though, no one would look at Postgres and go “it’s just a couple lookup tables put together”
- dragonwriter 2y agoAs an implementation? No. For fundamental understanding of the logical model? That its one big lookup table with a particular form of key nesting is... actually a pretty good model.
- littlestymaar 2y agoThat's exactly what DB indexes are though ;)
- littlestymaar 2y agoAn attention head is quite literally a lookup table!
- _akhe 2y agoExactly right. The others here are confusing LLM with "chat bot", and also they seem to be confusing token prediction with LLM. I have a feeling the mainstream won't really get it until a ChatGPT clone is online and ready to use lol and still it will be "This isn't a true chat bot, this is actually a Markovian language interface!"
- _akhe 2y agoThe library never claims to be an LLM - it's a next token prediction library that you can use to create LLMs. Python - one nice thing is all the "keys" retain their primitive type, where as in JS they all turn into strings. If Python wasn't so slow compared to JS that would matter a lot and I might use it, but the speed comparison isn't even close.
- mpeg 2y agoYou must be joking, the title of this post literally says "build fast LLMs from scratch" but the code is neither an LLM, nor particularly fast. Python dicts are actually different from JS objects, Python uses a hashmap behind the scenes while most fast js vms will apply some heuristics to decide whether to use a hashmap or to make the object as a static struct with a known offset for each key, this explains it better than I can: https://v8.dev/blog/fast-properties https://v8.dev/blog/fast-properties Nevertheless, for your specific use-case I would be really surprised if there was any significant difference between the performance of python and js – hashmaps are fast enough for this – plus with js objects you might be trading insertion performance for access performance in some cases, as there is overhead creating a v8-style "fast property" object
- _akhe 2y agoNot joking. You seem to be confused in thinking LLMs are built with other LLMs, but that's not the case. Otherwise why would you say "it says 'build fast LLMs from scratch' but the code is not an LLM" ?? Why would the code of a library to build LLMs need to also be an LLM? Getting off-topic but Python is incredibly slow at lookups (and most things) compared to JavaScript and it isn't even close, not all dynamic languages are the same. This is pretty widely known and a quick Google search yields plenty of benchmarks and articles! Give it a try. Python is used in AI/ML for its libraries (convenience) not because it's a fast language. There are 3 main reasons: 1) Time complexity of data structures is lower in JS, that's the primary exploit at play here 2) V8 compiles to machine code in less steps than Python and 3) Process forking - the concept that functions can run in parallel in the same thread. Thanks for your comment!