3 ms·
How can these models do anything close to RSI when they can’t even self check their output? Gemini for example is so self confidently wrong about 30% of the ti
by smackeyacky 13d ago
How can these models do anything close to RSI when they can’t even self check their output? Gemini for example is so self confidently wrong about 30% of the time for me on certain tasks. I tell it that its answer is wrong and it issues a mea culpa but goes back to being wrong in short order. I feel like the AI industry is still massively overstating their projections.
- kakugawa 13d agoThey can only do it in the (narrow) domains that are verifiable.
- StevenWaterman 13d agoAs someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like
- deleted 13d ago[deleted]
- pinkmuffinere 13d agoI think your reply has a somewhat familiar structure -- "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". You might be completely correct! But these sorts of claims push the onus back onto the other person, without accepting any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?
- StevenWaterman 13d agoThe frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems. Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.
- peterashford 13d agoI agree with you somewhat but I also just this morning read an article from a Blender educator who tried to replicate the Blender demos and couldnt get the same quality of results nor get results without errors that werent evident in Anthropic's demos
- x-complexity 12d ago> Is there any data you can provide to support your claim, or any result you can contribute here? By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations. You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.
- croon 12d agoWouldn't this also mean that all previous generations that were proclaimed as intelligent and working were in fact... not? It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven. To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.
- pinkmuffinere 12d agoTo preface -- I try not to be dogmatic/politicized on AI, so I will genuinely consider your arguments! Please try to convince me. (indeed, I am the grandparent commenter) I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?
- 12d ago
- mancerayder 13d agoMaybe Gemini is loosey goosey on purpose so we angrily correct it - then feed something on the back end that trains a different model? It's so abysmally bad on Google search... and it's free. Isn't Google the great pioneer of the product is us?
- emodendroket 13d agoThe kind of results you get from the one on Google Search and a dedicated "Gemini Pro" response are totally different. I'm assuming that's cost savings.
- drodgers 13d ago> Gemini That's definitely part of your problem. In my recent experience, error rates for astra/fable are at or below human level. Just like when directing humans, it pays to ask probing questions ('Are you sure about X?', 'Did you check for Y?', 'Please run Z just to double check.') if you really care about the result being correct.
- glhaynes 13d agoAnd if you're building something of any importance, you need to have verification steps at checkpoints. It's honestly just engineering. Weak models tasked with review can catch a decent amount of the mistakes that weak models make and help them be much better, especially if you have them verify against authoritative sources. Strong models make far fewer mistakes to begin with. And, for the mistakes they do make, a swarm of reviewers (same model or somewhat weaker, reviewed by the stronger model) can really help reduce the error rate further.
- PantaloonFlames 13d agoGauging state of the art against what is available for free or for very cheap per token cost is like gauging the maximum theoretical transport potential by riding a bicycle. You’re using something that is very energy efficient; you cannot extrapolate that experience to conclude that SOTA models are not doing something much different.
- deleted 13d ago[deleted]
- singingfish 13d agoI can't see a pathway for these things to be able to learn from experience in any meaningful way - the energy budget seems to prohibit it. Artificial neural networks are already many orders of magnitude more energy intensive than natural neural networks. Despite this, large nervous systems are very energy intensive as well. For example in humans 20% of the energy budget goes to the brain which is 5% of the body weight. So I refrer to LLMs as language extrusion confabulation machines. Language extrusion was a term I heard the linguist Emily Bender use. Confabulation because my observation is that talking to an LLM is very much similar to my experience of interacting with Korsakov syndrome patients some years ago. I look forward to the hype settling down to see what we end up with.
- rcxdude 12d ago> Artificial neural networks are already many orders of magnitude more energy intensive than natural neural networks. What metric are you using to compare? By most counts, the energy budget of an instance of an LLM in a datacenter is lower than the energy a person uses. Of course, if you count energy per neuron connections vs weights then you'll likely get a quite different number. But then again LLMs do a lot of things with far fewer weights than the brain does neuron connections, even if you only count neurons in some parts of the brain. And of course you can point to capabilities that the brain has but LLMs lack, but on the whole it feels like it's pretty difficult to make a useful like-for-like comparison here.
- singingfish 12d agoI can't make sense of your comment. Firstly because of the obvious massive over-build of GPU infrastructure the AI companies are engaging in. Secondly because the instance of the LLM in the data centre that users interact with is only a small part of the story. The training phase is clearly prohibitively expensive, thus the fact that these things have no way to learn from experience except by smoke and mirrors. Also your comment feels like the classic climate denial discourse - say something a bit complicated and a bit difficult to follow that looks at a very small out of context part of the story to cast doubt.
- 12d ago
- RataNova 12d ago[flagged]