3 ms·
Sorry OP, you seem to have a wrong understanding of what an LLM is, does and how it's trained. And I know that an LLM is not a chatbot, no worries. At most, yo
by gield 2y ago
Sorry OP, you seem to have a wrong understanding of what an LLM is, does and how it's trained. And I know that an LLM is not a chatbot, no worries.
At most, your project is a simple language model, but definitely not a large language model.
After looking at your code, you also seem to have a wrong understanding of embeddings. Other people in this thread have already shared great resources that offer good explanations on these topics.
You seem to be very dismissive of Markov chains (or HMMs) while they've been used in NLP for years and produce really great results. Your project seems to use some concepts of Markov chains but more simplified. Here [1] is a random article that does next-token prediction using Markov chains. It's basically what your project does but in a more scalable and probabilistic way.
>Use cases include: Auto-completion, auto-correct, spell checking, search/lookup, conversation simulation (chatbot), and more.
These use cases are not possible using the concepts used in your project.
[1] https://bespoyasov.me/blog/text-generation-with-markov-chains https://bespoyasov.me/blog/text-generation-with-markov-chain...
- _akhe 2y agoSorry you feel I haven't been grateful enough when people share resources and comments, I really do appreciate it - and I appreciate what you just shared here too. Thanks! I see a lot of parallels to this exercise and what I did - their use of generators is exceptionally slick. I bet theirs is way slower though. Mine has 1 dep (lodash) just for its deep merge function. I didn't even use Wink NLP for text splitting and other tasks because it was 10x or 20x slower than my simpler stuff from Stack Overflow. I actually like your description of "Markov chains but simplified". But really think the devx of `npm i simple-markov-chain` is worse than `npm i next-token-prediction`. I want people to actually use it for fast prediction. AFAIK the only difference between a language model and a "large" one is the amount of data it's trained on. Because that translates into a much larger model. Yeah this concept of embeddings is different than, say OpenAI's. They capture a broad spectrum of semantics, where mine only focuses on frequency and structure. However, I find that very often frequency and structure allows you to bypass semantic analysis and just get to the answer. The frequency method I'm using is I guess a bastardization of tf-IDF vectorization, let's call it "tf-IDF JSON.stringify!" "These use cases are not possible using the concepts used in your project." ^ not sure what you mean by this one, feel free to try it yourself or watch the video demo