5 ms·
I may be missing something obvious here, and I'm not a Python expert, but isn't this "get rid of all the emojis, except for one I like" code snippet completely
by keithwinstein 4y ago
I may be missing something obvious here, and I'm not a Python expert, but isn't this "get rid of all the emojis, except for one I like" code snippet completely wrong (despite the article's assertions that "it works fine," "the LLM used a neat hack," etc.)?
It seems to always just erase the file, because it declares `keep_emojis` to be a list of one element (`keep_emojis = ['¯\_(ツ)_/¯']`) and then goes USV-by-USV through the file to check if the encoded/decoded coded char is "in" keep_emojis. And there's basically no way that any single char can be. So the output is always empty.
If we wrote it like `keep_emojis = '¯\_(ツ)_/¯'` instead, then it would be a little closer, but in that case the program ends up preserving all non-ASCII characters (including all emojis), because these characters become the empty string after emoji.encode('ascii', 'ignore') and are therefore always "in" the keep_emojis string, so they never get replaced.
The explanation also doesn't make any sense, where it writes: "In the above code, the assumption is made that any ASCII character that cannot be encoded using UTF-8 is an emoji."
That's silly, because every ASCII character can be encoded in UTF-8. It also doesn't correspond to what the code is doing -- what it's really doing (I think) is testing for USVs that can't be encoded in ASCII, but then throwing the char away no matter what.
And, of course, the vast majority of non-ASCII coded characters in Unicode are not emoji, so this is not really a good way to satisfy the user's request. The roundtrip through ASCII seems like a bad idea -- if the user wants you to preserve some emojis while getting rid of most emojis, converting to ASCII (where there are no emojis) doesn't seem like it will be helpful.
Of course it's amazing that ChatGPT is already good enough to produce something close enough to fool venture capitalists into trusting it without trying the code, and these techniques will doubtless continue to improve.
- thundergolfer 4y agoI don’t think you’re missing something. The code looks busted, wouldn’t work at all.
- posnet 4y agoThis actually gets at the heart of the problem with the current batch of LLMs. They have no concept of truth, and so will gladly generate tokens that are "likely", but are at best clearly wrong, or worse look correct to a non-expert, but are subtly wrong. With software at least you can run it and inspect the behavior for flaws. Or have the results be validated by an actual programmer for correctness. And hopefully the developer using the LLM understands the limitations, and understands how it can hallucinate incorrect but convincing results. However even software engineers I know, who are aware of all of the above issues, will ask ChatGPT questions about fields they aren't familiar with and view the results as authoritative. It's the same effect as when you read a news article about a topic you're familiar with and can immediately see flaws, shortcuts and exaggerations; but then you turn the page to a topic you aren't an expert it and say to yourself "what an interesting article". The example from the article is even worse, since the author asked for an explanation of the code, and it hallucinated a clear and convincing but completely wrong explanation. Again to a programmer it is obvious, but to a layperson it's actively deceptive. I could easily imagine a scenario where some asks ChatGPT, what is the best combination of chemicals to clean mold from a ceramic surface, and it responding with "Bleach and Vinegar are the perfect way to clean a ceramic surface." and following up with an explanation like "Vinegar de-greases the surface allowing the the disinfectant power of the bleach to penetrate into the mold". Which all sounds reasonable, unless you know beforehand that those chemicals mixed produces toxic chlorine gas. Now that was definitely a contrived example, but unless you know and constantly question the output of an LLM you could easily be misled. Especially if you are asking about topics outside of it's training data or with minimal training examples. Maybe with enough training data or some combination of transformers with reinforcement learning that has a "truth" metric, hallucinating completely incorrect information can be reduced to an acceptable level. But at this point it seems intractable.
- polishdude20 4y agoNext thing in the future is to ask the llm to generate code. The llm generates it, runs it, writes tests, confirms they work etc. Now will it write correct tests? Maybe?
- thephyber 4y agoYou seem to be talking about giving the LLM the responsibility to do both strategy and tactics. In my experience, it can be useful in tactics, but usually fails to understand the concepts of strategy. My personal feeling is that an LLM is insufficient for strategy and is not the only technology suitable for implementing the tactics. I think it makes more sense to treat each as a module in a system, and build different modules to compete against each other.
- mikewarot 4y ago>They have no concept of truth 1 - They aren't optimized for truth, they are optimized for best appearing answers. 2 - What things do you know that you are willing to state on HN without being fear of being contradicted?
- posnet 4y agoI know they aren't optimized for truth, that was the point of my comment. But based on my own experience of showing people ChatGPT, without explaining the flaws they tend to treat the output as if it was. I don't understand the second question.
- mikewarot 4y ago2 was about my feeling that pretty much anything that is a "fact" seems to dissolve here on HN. Restated thus: The quickest way to falsify something is to state it as a fact here on HN.
- mistermann 4y agoEven here there are epistemic issues: >> This actually gets at the heart of the problem with the current batch of LLMs. Here you are implying that the fault lies [solely] with LLMs. >> They have no concept of truth, and so will gladly generate tokens that are "likely", but are at best clearly wrong, or worse look correct to a non-expert, but are subtly wrong. Hear you are asserting that you have the means yourself to reach truth. This is extremely easy to do if you're speaking abstractly, but try executing that at the concrete level (as above) and it's pretty difficult to avoid imperfection. Regarding the difficulty of truth, is the problem here entirely with LLMs, or is the problem with reality itself? When ChatGPT (or anyone/anything) says something "is" true (and some people agree) someone else says it "is" not, how are we to decide which is correct? Where does this "is" that people "are [only] perceiving" come from? Where, and what, "is" "reality"? Is the universe equal/identical to reality? That's how a lot of people talk...but is it true? INB4 "we could all be brains in jars", which is one of the most common responses to arise "purely by coincidence" when this topic is raised.
- mikewarot 4y ago>ChatGPT is already good enough to produce something close enough to fool venture capitalists VC funding is centered around making high risk/high payout bets in quantity. If they took the time it would take a functioning prudent investor to carry out due diligence, the startups would have already used up their runway and died. VCs trade risk for reward, and make it up on volume. Fooling a VC isn't a high bar in the eyes of this Midwesterner.
- CGamesPlay 4y agoThis actually gets at the heart of the problem with the current batch of hype around LLMs. The people writing this kind of blog either don't have the ability to understand, or didn't bother to check, that they just fired their software development team and replaced with with the software equivalent of a cargo cult. Just because the code highlights in the IDE does not mean the revenue will come. The crazy thing is that this is super common in many posts about LLMs! Here's another example that was recently on HN: https://news.ycombinator.com/item?id=34473783 https://news.ycombinator.com/item?id=34473783 - as called out here https://buttondown.email/hillelwayne/archive/programming-ais-worry-me/ https://buttondown.email/hillelwayne/archive/programming-ais...
- spudlyo 4y agoThis is so confused, it's making my head hurt. What I think the author wanted was to remove ASCII emoticons, because there is no such thing as ASCII emoji. Another confusing thing is that good 'ole shruggie (which in this form I'd classify as an emoticon) is not pure ASCII, and relies on Unicode codepoint 0x30c4 for the smiley face, and a couple more besides. The author did such a poor job of describing what they wanted, it's no surprise the generated program is nonsense. If anything, I'm now less convinced that non-technical people will be able to use conversational AI tools to usurp my hard-earned societal role as a technologist.
- citrin_ru 4y agoHumans are generally not very good at describing what they want. Transferring messy descriptions into the code asking questions in the process is what software developers do.
- kolinko 4y agoIt’s iterative. I pasted the same prompt/code to gpt4, and said the following: „it removed polish characters from the text!” It correctly figured out the problem and fixed it - removing only unicode emojis, not all the unicode characters. Then I told him: „but it still didn’t remove things like :) :/ xD” (not specyfying that I mean non-ascii emoticons, not emojis really). It replied by saying I meant emoticons and writing regexp for those and other ascii emoticons.