5 ms·
Worth mentioning, though, that people have already tried running all of them through LLMs at this point. So this is proof of the models actually getting strong
by CSMastermind 5mo ago
Worth mentioning, though, that people have already tried running all of them through LLMs at this point.
So this is proof of the models actually getting stronger (previous generations of LLMs were unable to solve this one).
- Tarq0n 5mo agoNot definitively. LLMs are stochastic with respect to input, temperature and the exact prompt. It's possible that the model was already capable of it but never received the exact right conditions to produce this output.
- teiferer 5mo agoEvery model is able to solve each problem, given the right prompt. (Worst case, the prompt contains the solution.)
- pontifier 5mo agoInteresting... Exhaustive brute force prompting might expose previously unknown capabilities in existing models. Seems like a whole can of worms.
- Calazon 5mo agoExhaustive brute force prompting is completely unfeasible. The number of potential prompts is impossibly large.
- teiferer 5mo agoIt "exhaustive brute forcing" approach does not need an LLM in the loop. Just brute force the possible outputs instead. They will contain all the most beautiful novels you can imagine!
- deleted 5mo ago[deleted]
- imiric 5mo ago> So this is proof of the models actually getting stronger (previous generations of LLMs were unable to solve this one). No, it's not. While I don't dispute that new models may perform better at certain tasks, the fact that someone was able to use them to solve a novel problem is not proof of this. LLM output is nondeterministic. Given the same prompt, the same LLM will generate different output, especially when it involves a large number of output tokens, as in this case. One of those attempts might produce a correct output, but this is not certain, and is difficult if not impossible for a human not expert in the domain to determine this, as shown in this thread.
- notahacker 5mo agoAs others have pointed out, a key part of the prompt used here may have been "don't search the internet" as it would most likely have defaulted to starting off with existing approaches to that problem...
- famouswaffles 5mo agoThis is one of a number of such results achieved only in the last few months with only the last crop of models. They have undoubtedly gotten better in this domain. Saying anything else is just denial. You can run these same problems on GPT-4 or 5 all you want, you'll get nowhere. In fact people did, and you're hearing about it now because it's these crop of models that are getting meaningful results.
- _ccwi 5mo agoMinor aside, these models do not return the same answer every time you prompt it. Makes it harder to reason over their effectiveness.
- rjh29 5mo agoYou don't need to say "Minor aside" either. Thankfully language is a creative endeavour not a scientific one.
- deleted 5mo ago[deleted]
- rjh29 5mo agoContext: parent originally said "you should not say 'worth mentioning', if it's worth mentioning you can just say it". That sentence has now been edited out so my comment looks weird.
- _ccwi 5mo agoYour reply was so rude it convinced me to edit. Your second reply is a distortion of my original message too.