6 ms·
This really is the fastest growing technology of all time. Do you feel the curve? I remember Mixtral8x7b dominating for months; I expected data bricks to do th
by harryp_peng 2y ago
This really is the fastest growing technology of all time. Do you feel the curve?
I remember Mixtral8x7b dominating for months; I expected data bricks to do the same! but it was washed out of existence in days, with 8x22b, llama3, gemini1.5...
WOW.
- krainboltgreene 2y agoI must be missing something because the output from two years ago feels exactly the same as the output now. Any comment saying the output is significantly better can be equally pared with a comment saying the output is terrible/censored/"nerfed". How do you see "fastest growing technology of all time" and I don't? I know that I keep very up to date with this stuff, so it's not that I'm unaware of things.
- HeatrayEnjoyer 2y agoThe best we had two years ago was GPT-3, which was not even instruction tuned and hallucinated wildly.
- cma 2y agoAre you trying the paid gpt or just free 3.5 chatgpt?
- krainboltgreene 2y ago100% of the time when I post a critique someone replies with this. I tell them I've used literally every LLM under the sun quite a bit to find any use I can think of and then it's immediately crickets.
- kolinko 2y agoUsually people who post such claims haven’t used anything beyond gpt3. That’s why you get questions. Also, the difference is so big and so plainly visible that I guess people don’t know how to even answer someone saying they don’t see it. That’s why you get crickets.
- imtringued 2y agoRT-2 is a vision language model fine tuned on the current vision input and actuator positions as the output. Google uses a bunch of TPUs to produce a full response at a cycle rate of 3 Hz and the VLM has learned the kinematics of the robot and knows how to pick up objects according to given instructions. Given the current rate of progress, we will have robots that can learn simple manual labor from human demonstrations (e.g. Youtube as a dataset, no I do not mean bimanual teleoperation) by the end of the decade.
- cma 2y agoYou see no difference between non-RLHFed GPT3 from early 2022 and GPT-4 in 2024? It's a very broad consensus that there is a huge difference so that's why I wanted to clarify and make sure you were comparing the right things. What type of usages are you testing? For general knowledge it hallucinates way less often, and for reasoning and coding and modifying its past code based on English instructions it is way, way better than GPT-3 in my experience.
- harryp_peng 2y agoI always use GPT4 to write boiler plate code etc. It probably automates 50% of my tasks, pretty good.
- Workaccount2 2y agoUsually when I encounter sentiment like this it is because they only have used 3.5 (evidently not the case here) or that their prompting is terrible/misguided. When I show a lot of people GPT4 or Claude, some percentage of them jump right to "What year did Nixon get elected?" or "How tall is Barack Obama?" and then kind of shrug with a "Yeah, Siri could do that ten years ago" take. Beyond that you have people who prompt things like "Make a stock market program that has tabs for stocks, and shows prices" or "How do you make web cookies". Prompts that even a human would struggle greatly with. For the record, I use GPT4 and Claude, and both have dramatically boosted my output at work. They are powerful tools, you just have to get used to massaging good output from them.
- parineum 2y ago> or that their prompting is terrible/misguided. This is the "You're not using it right" defense. It's an LLM, it's supposed to understand human language queries. I shouldn't have to speak LLM to speak to an LLM.
- Filligree 2y agoThat is not the reality today. If you want good results from an LLM, then you do need to speak LLM. Just because they appear to speak English doesn't mean they act like a human would.
- jiggawatts 2y agoPeople don’t even know how to use traditional web search properly. Here’s a real scenario: A Citrix virtual desktop crashed because a recent critical security fix forced an upgrade of a shared DLL. The output is a really specific set of errors in a stack trace. I watched with my own two eyes an IT professional typed the following phrase into Google: “Why did my PC crash?” Then he sat there and started reading through each result… including blog posts by random kids complaining about Windows XP. I wish I could say this kind of thing is an isolated incident.
- Aeolun 2y agoI mean, you need to speak German to talk to a German. It’s not really much different for LLM, just because the language they speak has a root in English doesn’t mean it actually is English. And even if it was, there’s plenty of people completely unintelligible in English too…
- Eisenstein 2y agoIt's fine, you don't have a use for it so you don't care. I personally don't spend any effort getting to know things that I don't care about and have no use for; but I also don't tell people who use tools for their job or hobby that I don't need how much those tools are useless and how their experience using them is distorted or wrong.
- hehdhdjehehegwv 2y agoI do massive amounts of zero shot document classification tasks, the performance keeps getting better. It’s also a domain where there is less of a hallucination issue as it’s not open ended requests.
- krainboltgreene 2y agoI didn't ask what you do with LLMs, I asked how you see "fastest growing technology of all time".
- hehdhdjehehegwv 2y agoI didn’t say that?
- steve_adams_86 2y agoIt strikes me as unprecedented that a technology which takes arbitrary language-based commands can actually surface and synthesize useful information, and it gets better at doing it (even according to extensive impartial benchmarking) at a fairly rapid pace. It’s technology we haven’t really seen before recently, improving quite quickly. It’s also being adopted very rapidly. I’m not saying it’s certainly the fastest growth of all time, but I think there’s a decent case for it being a contender. If we see this growth proceeding at a similar rate for years, it seems like it would be a clear winner.
- harryp_peng 2y agoBut humans aren't 'original' ourselves. How do you do 3*9? You memorized it. It's striking how humans could reason at all.
- oarsinsync 2y ago> How do you do 3*9? You memorized it I put my hands out, count to the third finger from the left, and put that finger down. I then count the fingers to the left (2) and count the fingers to the right (2 + hand aka 5) and conclude 27. I have memorised the technique, but I definitely never memorised my nine times table. If you’d said ‘6’, then the answer would be different, as I’d actually have to sing a song to get to the answer.
- hehdhdjehehegwv 2y agoFunny thing is I’m still in love with Mistral 7B as it absolutely shreds on a nice GPU. For simple tasks it’s totally sufficient.
- qeternity 2y agoLlama3 8B is for all intents and purposes just as fast.
- minosu 2y agoMistral 7b inferences about 18% faster for me as a 4bit quantized version on an A100. Thats definitely relevant when running anything but chatbots.
- tmostak 2y agoAre you measuring tokens/sec or words per second? The difference matters as generally in my experience, Llama 3, by virtue of its giant vocabulary, generally tokenizes text with 20-25% less tokens than something like Mistral. So even if its 18% slower in terms of tokens/second, it may, depending on the text content, actually output a given body of text faster.