6 ms·
Could you give an example of something we recently solved that was considered an unsolvable problem six months beforehand? I don’t have any specific examples, b
by chefandy 2y ago
Could you give an example of something we recently solved that was considered an unsolvable problem six months beforehand? I don’t have any specific examples, but it seems like most of the huge breakthrough discoveries I’ve seen announced end up being overstated and for practical usage, our choice of LLM-driven tools is only marginally better than they were a couple of years ago. It seems like the preponderance of practical advancement in recent times has come from the tooling/interface improvements rather than generating miracles from the models themselves. But it could be that I just don’t have the right use cases.
- intelkishan 2y agoPerformance of OpenAI o3 in the ARC-AGI challenge fits the bill, however the model is not released publicly.
- gallerdude 2y agoCompletely disagree… there are a crazy amount of cases that didn’t work, until the models scaled to a point they magically did. Best example I can think of is the ARC AGI benchmark. It was seen to measure human-like intelligence through special symmetries and abstract patterns. From GPT-2 to GPT-4 there was basically had no progress, then o1 got about 20%. Now o3 has basically solved the benchmark.
- chefandy 2y agoI guess what I'm probably not seeing from my vantage point is that translating into a better experience with the tools available. I just cancelled a ChatGPT plus subscription because it just didn't seem useful enough to justify the price. I absolutely understand that there are people for whom it is, but nearly everyone I see that talks a lot about the value of AI either has use cases that I don't care about such as automated "content" generation or high-volume lowish-skill code generation, or they see achieving a progressively more difficult set of benchmarks as a useful end in itself. I like copilot autocomplete when I'm coding, but the quality of that hasn't dramatically changed. I don't give a damn about benchmarks-- I only care what I get from it practically. I have absolutely no interest in using ChatGPT as a therapist or companion because I value human connection and have access to it. So far I simply don't see significant changes in what comes out vs what gets typed in for practical usage. I wouldn't give ChatGPT logic problems to solve except maybe for generating code because I know code well enough to quickly evaluate its output. If the caveat is "hey FYI this thing might hide some frustratingly plausible looking bullshit in the answer so double-check its work," then what good is it really for hard problems if you just have to re-do them anyway? The same thing is true with image generation. Sure, it's better in ways that are sort-of meaningful for low-value professional or hobby usage, but it's barely budged the barriers to becoming good enough for high-end media production. I totally believe that this technology is improving and when you're looking at it in isolation, those improvements seem meaningful. But I just don't see that yet translating into things most of the general public can sink their teeth into. With things like the (still) shitty google search "enhancements", and users being forced into AI-driven chat workflows or having big loud not-really-useful UI elements dedicated to AI features, in some ways they've made people's experience using computers meaningfully worse. Just like with Mastodon, I see a huge disconnect with the tech crowd's excitement with what's happening with the technology, and how that ends up working for users that need to actually solve their problems with that technology.
- munchler 2y agoTake a look at the ARC Prize, which is a test for achieving "AGI" created in 2019 by François Chollet. Scroll down halfway on the home page and ponder the steep yellow line on the graph. That's what OpenAI o3 recently achieved. [0] https://arcprize.org/ https://arcprize.org/ [1] https://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough
- mrshadowgoose 2y agoReviewing the actual problems is highly recommended: https://kts.github.io/arc-viewer/ https://kts.github.io/arc-viewer/ They're not particularly difficult, but clearly require reasoning to solve.
- UnlockedSecrets 2y agounless you train directly against solving those problems... in which case how could you theoretically design a test that could stand against training directly against the answer sheet?
- munchler 2y agoThat's why they keep the evaluation set private: "Submit a solution which scores 85% on the ARC-AGI private evaluation set and win $600K." [0] https://arcprize.org/guide https://arcprize.org/guide
- EdwardDiego 2y agoSo we're only 12% from AGI? I'm dubious tbh. Given we still can't simulate a nematode.
- munchler 2y agoThat's why I put "AGI" in quotes. The point is that six months ago, no one expected an LLM to score this well.
- liamwire 2y agoNot quite what you asked for, but it seems tangentially related and you might find it interesting: https://r0bk.github.io/killedbyllm/ https://r0bk.github.io/killedbyllm/
- janalsncm 2y agoWould be interesting to have a list of startups killed by ChatGPT as well.