14 ms·
It does seem hacky, but then again the whole concept of conversational LLMs is. You're just asking it to add an extra word to a given conversation and after a b
by nottheengineer 3y ago
It does seem hacky, but then again the whole concept of conversational LLMs is. You're just asking it to add an extra word to a given conversation and after a bit, it spits out an end token that tells your application to hand control back to the user.
I think latent space and text space aren't as far apart as you think. LLMs are pretty stupid, but very good at speech. They are good at writing code because that's very similar, but fall apart in things that need some actual abstract thinking, like math.
Those text space hacks do tend to work and stuff like "think step by step" has become common because of that.
LoRAs are closer to what you mean and they're great at packing a lot of understanding into very little data. But adjusting weights for a single conversation just isn't feasible yet, so we're exploring text space for that purpose. Maybe someome will transfer the methods we discover in text space to embedding space to make them more efficient, but that's for the future.
- famouswaffles 3y ago>They are good at writing code because that's very similar, but fall apart in things that need some actual abstract thinking, like math. Pretty odd assertion. LLMs are not "good at speech, bad at abstract thinking". What do these have to do with speech ? https://general-pattern-machines.github.io/ https://general-pattern-machines.github.io/ https://arxiv.org/abs/2212.09196 https://arxiv.org/abs/2212.09196 It doesn't even hold with your example because GPT-4 is pretty good at Math, nowhere near "falling apart".
- riku_iki 3y ago> GPT-4 is pretty good at Math, nowhere near "falling apart". its good at tasks which were included into training dataset in some variations.
- PaulHoule 3y agoKinda able to do some math tasks some of the time whereas you can use techniques from the arithmetic textbook to get the right answer all of the time with millions of times less CPU even including the overhead of round-tripping to ASCII numerals which is shockingly large compared to what a multiply costs. Kinda "the problem" with LLMS is that they successfully seduce people by seeming to get the right answer to anything 80% of the time.
- famouswaffles 3y agoMath is a lot more than just Arithmetic.
- PaulHoule 3y agoYeah but if you can only do arithmetic right X% of the time you aren't going to get other answers right as often as would really be useful. That said, LLMs have a magic ability to "short circuit" and get the right answer despite not being able to get the steps right. I remember scoping out designs for NLP systems about 5 years ago and frequently conclude that "that won't work" because information was lost at an early stage but in retrospect by short circuiting a system like that can outperform its parts but it still faces a ceiling on how accurate the answers are because the reasoning is not sound.
- btilly 3y agoHuman reasoning is amazingly not sound. When you add in various patterns, double-checks, and memorized previous results, what human reasoning can do is astounding. But it is very, vary far from sound.
- riku_iki 3y agoall currently available reasoning approaches are limited. I guess the topic is how far GPT in reasoning is from human. We can take some simple tests: - can GPT play chess as well as humans, as benchmark of reasoning games? - did GPT prove some nontrivial math theorems or solve some math problems where humans couldn't find solution yet?
- PaulHoule 3y agoOne thing I thought was amusing was that there was a burst of articles about Cyc that got mentioned when Doug Lenat died including this arXiv paper https://arxiv.org/abs/2308.04445 https://arxiv.org/abs/2308.04445 and that one said that Cyc had over 1,100 special purpose reasoning engines. The general purpose resolution solver was nowhere near fast enough to be really useful. Early one there was https://en.wikipedia.org/wiki/General_Problem_Solver https://en.wikipedia.org/wiki/General_Problem_Solver which would be capable in principle of finding a winning move in a chess position but because it worked by exhaustive search it would practically take too long. The thing is that a good chess playing program is not generally intelligent just as a chess grandmaster isn't necessarily good at anything other than chess, it just has special purpose heuristics (as opposed to algorithms) that find good chess move. ChatGPT-like systems will be greatly improved by coupling them to other systems such as "write a Python/SQL script then run it", "run a query against bing and summarize the results", and "go find the chess engine and ask it what move to make", that is, like Cyc, it will get a swiss army knife of tools that help it do things it's not good at but it doesn't create general intelligence any more than Cyc did. Robert Penrose in the Emperor's New Mind suggests that there must be some quantum magic in the human mind because the human mind is able to solve any math problem whereas any machine is limited by Gödel's theorem. It's silly, however, because we don't humans are capable of proving any theorem: look at how we struggled with Fermat for nearly 360 years or how https://en.wikipedia.org/wiki/Collatz_conjecture https://en.wikipedia.org/wiki/Collatz_conjecture seems not even tantalizingly out of reach. The difference might be that humans feel bad when they get the wrong answer whereas ChatGPT certainly doesn't. (as much as its empty apology can be satisifying to people) This isn't just an attribute of humans, working with other animals such as horses I'm convinced that they feel bad when they screw up too.
- Nevermark 3y agoI have played around with GPT-4 and some fairly simple but completely new math ideas. It was fabulous at identifying special cases I overlooked, that disproved conjectures.
- sudokuist 3y agoExample?
- Nevermark 3y agoI was playing around with prime numbers, and simple made up relationships between them, such as between the square of a prime N vs. the set of primes smaller than N, etc. It caught me out with specific examples that violated my conjectures. In one case the conjecture held for all but one case, another conjecture was generally true but not for 2 and 3. In one case it thought a conjecture I made was wrong, and I had to push it to think through why it thought it was wrong until it realized the conjecture was right. As soon as it had its epiphany, it corrected all its logic around that concept. It was very simple stuff, but an interesting exercise. The part I enjoyed the most was seeing GPT-4's understanding move and change as we pushed back on each other's views. You miss out on that impressive aspect of GPT-4 in simpler sessions.
- sudokuist 3y agoHave you tried formalizing your ideas with Isabelle? It has a constraint solver and will often find counterexamples to false arithmetical propositions[1]. 1: https://isabelle.in.tum.de/overview.html https://isabelle.in.tum.de/overview.html
- famouswaffles 3y agoNo it's just pretty good in general lol.
- cerved 3y agomy experience is that it's pretty subpar
- astrange 3y agoAre you using code interpreter? It's better. The mobile app doesn't offer it though, and also has a system prompt that causes some strange behavior - sometimes it will put emojis in the text and then apologize for using emojis.
- bigyikes 3y agoCare to share a GPT conversation you’ve had? I’m interested in what sorts of prompts lead you to this opinion. My experience is the opposite.
- cerved 3y agoA bit too much of a hassle. But if you're willing to share some of your good experiences, I'm curious
- sudokuist 3y agoCan LLMs solve sudoku yet?
- JieJie 3y agoThese folks think so. https://github.com/jieyilong/tree-of-thought-puzzle-solver https://github.com/jieyilong/tree-of-thought-puzzle-solver
- sudokuist 3y agoThey don't have 9x9 puzzles. Any guesses as to why they only tried 3x3, 4x4, and 5x5 but not 9x9? This work is interesting. I wouldn't have guessed 3x3 puzzles would be solvable by a large Markov chain. It would be interesting to know how large of a context is necessary to solve 9x9 puzzles. No existing model can currently solve 9x9 puzzles even though the recursive backtracking algorithm can solve any given puzzle in less than a second.
- JieJie 3y agoWell, you just said sudoku. As others have pointed out, maybe intelligence derived from language just isn't very good at math? It's not like linear algebra comes naturally to humans, we have to be specially trained. I've been taking Khan Academy classes and believe me, math sure doesn't come naturally to me. I realize tempers are high on this subject, but I literally just wanted to point it out, in case you hadn't seen it. I wasn't trying to dunk on you or anything.
- chaxor 3y agoWhy are people so intent on incorrectly asserting these models are Markov chains? It makes sense to use the analogy as an educational tool for exposition, but it more often seems that many use it as a way to minimize the notion that these models could ever possibly be useful for anyone. Is this just simply to make it more intuitive for others that it's a sequence model? Because it seems about as helpful as 'email is just bits' when everyone and their grandma knows about the relation between transformers, GAT, and circulant matrices.
- dongping 3y agoI'm not sure about its math, but GPT-4 fails miserably at simple arithmetic questions like 897*394=? The GPT-3.5 turbo is fined-tuned for arithmetic according to ClosedAI (noted in one of the change logs), so it is sometimes slightly better, but nevertheless always fails equations like 4897*394=?
- Filligree 3y ago> I'm not sure about its math, but GPT-4 fails miserably at simple arithmetic questions like 897*394=? That's, um, about 300,000? ... 353,418 actually. But I'm not going to blame the AI too much for failing at something I can't do either.
- dongping 3y agoOne can resort to traditional vertical multiplication (which requires patience), or do 897*394 = (900-3) * (400-6) = 900*400 - 6*900 - 400*3 + 3*6 = 360,000 - (5,400 + 1,200) + 18 = 360,018 - 6,600 = 353,418
- 6510 3y ago8*3=24 and 800*300 =240000 8*9=72 and 800* 90 = 72000 8*4=32 and 800* 4 = 3200 9*3=27 and 90*300 = 27000 9*9=81 and 90* 90 = 8100 9*4=36 and 90* 4 = 360 7*3=21 and 7*300 = 2100 7*9=63 and 7* 90 = 630 7*4=28 and 7* 4 = 28 -------------------------- 353418
- dash2 3y agoBut you are smart enough to use a computer or calculator. And AI is a computer. So the naive expectation would be that it would be capable of doing as well as a computer. Also, you probably could do long multiplication with paper and pencil if you needed to. So a reasoning AI (which has read many many descriptions of how to do long multiplication) should be able to also.
- Filligree 3y ago
- nottheengineer 3y agoPattern reproduction is very close to speech in my opinion. Formal grammars even have it in the name and approaches like https://news.ycombinator.com/item?id=37125118 https://news.ycombinator.com/item?id=37125118 show that LLMs are indeed very fit for that purpose. I think I have to walk that claim about math back and try to phrase what I meant differently: LLMs have a hard time with problems that don't translate well into the text space, i.e. abstract problems. Math used to be one of those because early tokenizers were designed just with text in mind and LLMs weren't good enough to overcome those limitations. OpenAI put in a lot of effort into their tokenizers to make GPT3.5 and GPT4 better at math specifically. The second paper you linked is very interesting and I think it supports my original assertion of text space and latent space being close. The first graph shows GPT3.5 doing much better at pattern reproduction and language tasks while humans still hold an advantage in the more abstract tasks like story analogies. Higher order relations being a key thing that's measured maybe makes this task a bit too perfect for arguing my case, but it does show that humans have an advantage in more abstract situations. I think any problem that can be viewed as being mostly a form of translation is a good one for LLMs and if you can express a problem as that, you can get better results. To get back to the main point: latent space and text space, or feature space in general, being close is what I believe causes all of this. Happy to hear counterexamples.
- westurner 3y agoFWIU recently there's?: - Increase the input prompt token limit (2023-09: 32K tokens in: OpenAI GPT-4 Enterprise, Giraffe (LLama 2)) - Fine tune [a "LoRA" atop a foundation model] - TODO: ~Checkpoint w/ Copy-on-Write
- orbital-decay 3y ago> LLMs are pretty stupid, but very good at speech. They are good at writing code because that's very similar, but fall apart in things that need some actual abstract thinking, like math. Isn't that more of a training method issue? Try teaching a caveman to count by making him memorize the sticks and words pairs like LLM does, and you will get similar results, as he won't know about the stateful counting algorithm. Humans get the powerful reasoning ability through gradual learning of new abstractions, and won't be able to extract anything useful from a textbook on quantum physics until they learned the basics first.