4 ms·
These are always cherry-picked, though. They tell you about the 1/10 that went really impressively, ignoring the other 9 shots at the task where the clanker sta
by Dilettante_ 2mo ago
These are always cherry-picked, though. They tell you about the 1/10 that went really impressively, ignoring the other 9 shots at the task where the clanker started to try selling tungsten cubes (in person, wearing a blue shirt).
- jstanley 2mo agoThis is exactly the kind of take that the comment you replied to is talking about.
- makaking 2mo agoFair. But that 1/10 continues to get more and more impressive. Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times. Anthropic reports "Claude figured out how to tie general relativity with quantum mechanics." Would you hand-wave it away saying that it's cherry-picked?
- haldujai 2mo ago> Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times. I would not hand wave that away - but Claude 5 feels closer to Claude 1 than the hypothetical Claude 7 you propose. I do not believe one can extrapolate LLMs that far ahead despite the very substantial progress so far.
- fourseventy 2mo agoLike two days ago Claude solved a century old math conjecture
- haldujai 2mo agoIf you are referring to the Jacobian conjecture Claude only provided a counterexample, not a disproof, neither of which is necessarily a “new brilliant scientific idea”.
- throw-the-towel 2mo agoThe whole "AI is a parrot" argument feels like moving the goalposts so quickly, you could actually hear the whooshing sound they make as they move.
- haldujai 2mo agoNo, that’s not my argument. Rather it is that a counterexample to a mathematical proof whether produced by human or AI may be a scientific discovery but it is not inherently new and brilliant by definition. The case for AI is also weakened when the model is steered by an expert.
- throw-the-towel 2mo agoSo what is your definition of "new and brilliant"?
- fourseventy 2mo agoA counterexample proves that the conjecture is false, not sure what you are talking about. And we can play word games all day about what counts as a "new brilliant scientific idea" but the fact is that no human had been able to solve it.
- mistercheph 2mo agoIt wouldn't surprise me that a company with an effectively unlimited budget could fund enough researchers to solve breakthrough problems, all while using LLM's (which are excellent research and exploration tools!) and then claim that the LLM found the breakthrough. The amount of money Anthropic has to play with is >10,000x what entire fields of research have, I think people don't appreciate how few resources we spend to support people working on hard problems that don't have clear commercial applications. Very similar story with security research, LLM's are a super useful tool while hunting vulnerabilities, but it turns out when the entire software industry starts throwing tens of billions of dollars at vuln. research, a lot of stuff gets unearthed, something security people have been insisting on for years and complaining that their work is underfunded and under-resourced.
- cardboard9926 2mo ago> Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times Sadly, general public (us) is never seeing that model
- nolok 2mo agoWhen I talk about my kid to friends I talk them about he did that awesome thing, I don't specifically insist on the 99 times before where he miserably failed. They're not hidden, and we all know they exists and on occasion laugh about a few particularly funny ones, but overall the idea is that they don't matter much in terms of development, what matters is that if he succeeded once from now on his percentage of success will keep improving. I don't believe in all the LLm is AI is AGI dream, it's too easy to trigger failure case that show a lack of basic thinking no matter how good they do on these tests. But I also can recognize the insane things that are made possible by them. PS: I believe llm true power comes from hive/ant behavior, that's why we're so amazed by goal and agentic and sub agent PS2: it's rather easy to figure out when we're there : when they can /goal it into improving itself until it does strictly better than itself at those benchmark, they've essentially reached mini singularity.
- camgunz 2mo agoYeah but you're also not like "my genius kid will put you all out of work".