5 ms·
I tried the chatbot. jarring to see a large response come back instantly at over 15k tok/sec I'll take one with a frontier model please, for my local coding an
by metabrew 8mo ago
I tried the chatbot. jarring to see a large response come back instantly at over 15k tok/sec
I'll take one with a frontier model please, for my local coding and home ai needs..
- grzracz 8mo agoAbsolute insanity to see a coherent text block that takes at least 2 minutes to read generated in a fraction of a second. Crazy stuff...
- pjc50 8mo agoAccelerating the end of the usable text-based internet one chip at a time.
- kleiba 8mo agoYes, but the quality of the output leaves to be desired. I just asked about some sports history and got a mix of correct information and totally made up nonsense. Not unexpected for an 8k model, but raises the question of what the use case is for such small models.
- djb_hackernews 8mo agoYou have a misunderstanding of what LLMs are good at.
- cap11235 8mo agoPoster wants it to play Jeopardy, not process text.
- paganel 8mo agoNot sure if you're correct, as the market is betting trillions of dollars on these LLMs, hoping that they'll be close to what the OP had expected to happen in this case.
- raincole 8mo agoThe market didn't throw trillions of dollars to develop Llama 3 8B. What GP is expected to happen has happened around late 2024 ~ early 2025 when LLM frontends got web search feature. It's old tech now.
- paganel 8mo agoThe GP’s point was about LLMs generally, no matter the interface. I agree that this particular model is (relatively speaking) ancient in AI the world, but go back 3 or 4 years and this (pretty complex “reasoning” at almost instant speed) would have seemed taken out of a science-fiction book.
- IshKebab 8mo agoI don't think he does. Larger models are definitely better at not hallucinating. Enough that they are good at answering questions on popular topics. Smaller models, not so much.
- kleiba 8mo agoCare to enlighten me?
- vntok 8mo agoDon't ask a small LLM about precise minutiae factual information. Alternatively, ask yourself how plausible it sounds that all the facts in the world could be compressed into 8k parameters while remaining intact and fine-grained. If your answer is that it sounds pretty impossible... well it is.
- kleiba 8mo agoDid you see the part in my original post where it said "Not unexpected for an 8k model"?
- vntok 8mo agoOh I saw it, you still have a fundamentally flawed comprehension of LLMs. The size of the model does not factor as tiny models can use Internet to fetch factual information. But you think they are accurate repositories of knowledge, even though it's physically impossible unless lossless infinite compression algorithms exist (they don't, can't and won't).
- kleiba 8mo agoI think you're overestimating your ability to assess what others think or comprehend.
- kgeist 8mo ago8b models are great at converting unstructured data to a structured format. Say, you want to transcribe all your customer calls and get a list of issues they discussed most often. Currently with the larger models it takes me hours. A chatbot which tells you various fun facts is not the only use case for LLMs. They're language models first and foremost, so they're good at language processing tasks (where they don't "hallucinate" as much). Their ability to memorize various facts (with some "hallucinations") is an interesting side effect which is now abused to make them into "AI agents" and what not but they're just general-purpose language processing machines at their core.
- eternauta3k 8mo agoWould be nice to point this at (pre-LLM) Wikipedia and fill out Wikidata!
- VMG 8mo agoNot at all if you consider the internet pre-LLM. That is the standard expectation when you load a website. The slow word-by-word typing was what we started to get used to with LLMs. If these techniques get widespread, we may grow accustomed to the "old" speed again where content loads ~instantly. Imagine a content forest like Wikipedia instantly generated like a Minecraft word...
- stabbles 8mo agoReminds me of that solution to Fermi's paradox, that we don't detect signals from extraterrestrial civilizations because they run on a different clock speed.
- dintech 8mo agoIain M Banks’ The Algebraist does a great job of covering that territory. If an organism had a lifespan of millions of years, they might perceive time and communication differently to say a house fly or us.
- xyzsparetimexyz 8mo ago:eyeroll:
- pennomi 8mo agoYeah, feeding that speed into a reasoning loop or a coding harness is going to revolutionize AI.
- deleted 8mo ago[deleted]