5 ms·
It is as though its mathematical abilities are incomplete in their training, and wildly, incomprehensibly convoluted: I tried many base64 strings and they all
by PostOnce 4y ago
It is as though its mathematical abilities are incomplete in their training, and wildly, incomprehensibly convoluted:
I tried many base64 strings and they all decoded correctly until:
It "decoded" the base64 string for "which actress is the best?" except that it replaced "actress" with "address"... there is no off-by-one error that brings you to that.
You may try 100 base64 strings and they all decode correctly... only to find, in fact, that it DOES NOT know how to decode base64 reliably.
This tool could be a 50x accelerator for an expert, but absolutely ruinous to a non-expert in any given field.
I also got it to draw an icosahedron whose points were correct but whose triangles were draw incorrectly, so if I create a convex hull over it, it's correct.
The kinds of mistakes it makes are so close but so far at the same time. It sometimes writes complete working programs that are off by a single variable assignment, or sometimes they're just perfect, other times, they're nonsensical and call magic pseudocode functions or misunderstand the appropriate algorithm for a context (e.g. audio vs text compression).
It can provide citations for legal opinions -- but decades old citations that don't reflect current precedent.
God help us all if they plug it into some robot arms or give it the ability to run arbitrary code it outputs on a network interface.
Let's say they dump another 10 billion dollars into it and dectuple the size of the network, will it suddenly become legitimately capable, and not just "wow that's close" but actually startlingly competent in many more fields?
I could see this thing causing a war by all manner of means, whether its putting many out of work, making beguiling suggestions, outputting dangerous code, or, I'm sure, a million things that don't spring immediately to my small mind.
- xg15 4y ago> I tried many base64 strings and they all decoded correctly until: It "decoded" the base64 string for "which actress is the best?" except that it replaced "actress" with "address"... there is no off-by-one error that brings you to that. I'm still baffled by weird failure modes like this. Out of couriosity, did you also give it base64 that just contained random letters, so it can't jump to any word associations?
- nl 4y agoThink of it as overly aggressive error correction at the language level. It has a context of some Base64 code. Given that is almost always seen associated with computer code, is "address" or "actress" more likely. It "knows" the algorithm for decoding base64, and can follow those steps. But it can't overcome it's built-in biases for optimizing the most likely output given the context. (This problem is solvable, but I think that thinking about it like this helps understand why it behaves like it does)
- xg15 4y ago> Given that is almost always seen associated with computer code, is "address" or "actress" more likely. Sorry, but I don't buy it. I don't think "address" is a particularly likely word to appear in code, especially the kind of code that uses base64 (usually high-level). It appears even less often inside base64 encoded content.
- visarga 4y agoThis is all irrelevant. A language model should not run code itself, instead it should have a code execution environment, where it can read the error messages and iterate. It's terribly inefficient and error prone to run code directly. People also code on computers, not on paper.
- nl 4y agoThe point is to develop good intuitions for how large language models behave. If someone can develop a differentiable script runner that would be great! But the intuition about how the language model is behaving is useful for more than this specific problem.
- nl 4y agohttps://github.com/search?q=base64+address https://github.com/search?q=base64+address gives 9M+ results The original use for base64 was to send binary content to an email address.
- visarga 4y ago> I tried many base64 strings and they all decoded correctly until: You're holding it wrong. Let's not kill flies with cannons. How many million times less efficient is to do that than run the code on CPU? And still makes errors, as you said. Because it's a probabilistic model, not a deterministic computer. It's like a car bad at flying.
- ClumsyPilot 4y agoThis indicates it equally unreliable at a broad range of tasks. Applying it in self driving, life insurance, etc. will produce terrible outcomes.
- verdenti 4y ago
- ravi-delia 4y agoEh, the squishier the task the better it performs. It's ability to decode base 64 couldn't be less related
- ClumsyPilot 4y agoEhat is the basis for this claim? Same algorythm and dataset are used, the only difference is that we suck at detecting errors in squishy tasks. Squishy tasks is where people purposefully hide corruption, fraud and discrimination
- fragmede 4y ago> It can provide citations for legal opinions -- but decades old citations that don't reflect current precedent. From what I've seen of "citations" in other areas (eg asking it to generate stack overflow answers), I'm surprised the citations are even real. It seemed to be wholly making up citations, complete with real-enough looking URLs!
- btbuildem 4y agoSounds like we've tried a lot of the same things! I was asking it to generate an image and encode it as Base64 -- failed miserably. Then it turned out whatever image I had it cook up, the Base64 version would be the same malformed string. For "legal advice" it was super helpful in finding sections of the legal code relevant to my query. It also happily returned cases where rulings where the accused was found guilty and not guilty -- but searching for these cases in the archives of the courts it claimed they came from found no results. The way I understand the reasons behind these.. anomalies? hallucinations? is that this is precisely how this model works - it constructs sequences. For natural language, it works well enough, but if you need to deal with facts, then it can easily stray into a dream world.
- pcthrowaway 4y agoI asked it to give me song lyrics, and I got complete fiction. What's funny is that the fiction sound like it could be the accurate song lyrics given the song title and band, and it was poetic too. If you're curious, I asked it for "Bukowski" by Modest Mouse, because I wanted to see what its interpretation of the song would be. When I fed it the correct lyrics, it claimed to recognize them, and apologized for the inaccuracy earlier, then had some intelligible things to say about it, until it reverted to analyzing the made-up song lyrics.
- HTTP418 4y agoSounds pretty human to me.
- RosanaAnaDana 4y agoThat's wild. I had a similar issue asking to to translate to and from binary representations of characters. It would occasionally do fine. Most of the time it would add in a reference to "Charles". It's bizzare.
- usgroup 4y agoBare in mind that humans have a fundamental sense of meaning. You write and speak to express meaning : that’s syntax . Your valid sentences map onto valid meanings. You can speak nonsense but you can also tell it’s nonsense. GPT sentences map into meaning only incidentally due to training — that is, it has no sense of semantics. One can systematically thus generate an infinity of defeating cases.