3 ms·
As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like
by StevenWaterman 13d ago
As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like
- deleted 13d ago[deleted]
- pinkmuffinere 13d agoI think your reply has a somewhat familiar structure -- "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". You might be completely correct! But these sorts of claims push the onus back onto the other person, without accepting any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?
- StevenWaterman 13d agoThe frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems. Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.
- peterashford 12d agoI agree with you somewhat but I also just this morning read an article from a Blender educator who tried to replicate the Blender demos and couldnt get the same quality of results nor get results without errors that werent evident in Anthropic's demos
- x-complexity 12d ago> Is there any data you can provide to support your claim, or any result you can contribute here? By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations. You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.
- croon 12d agoWouldn't this also mean that all previous generations that were proclaimed as intelligent and working were in fact... not? It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven. To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.
- pinkmuffinere 12d agoTo preface -- I try not to be dogmatic/politicized on AI, so I will genuinely consider your arguments! Please try to convince me. (indeed, I am the grandparent commenter) I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?
- croon 12d agoI'll respond/comment on the parts I hope are relevant to you, in no particular order: Yes, a proof is a proof regardless how you get there. We however don't hear about when they fail, and I doubt their 10000 agents (from their own statement) would necessarily reach another solution/proof (this by leaning towards using user data after finding out others were close). They could as well have attacked another Millenium problem, but they didn't. In whichever case, we will have to wait and see if they (either company) can reach novel solutions/proofs without significant amount of human provided data for the LLM to bridge the gaps. Further, and this is more of a policy opinion/prediction: If the numerable obtainable (albeit very hard) problems are solved, assuming training data is needed, will it push out future researchers from entering the field due to lack of reachable goals, thus cutting off future training data? LLMs have been great at replacing gateway jobs. But those jobs are what leads to frontier training data (be it maths, physics, chemistry, economics, graphics, prose, etc).
- mancerayder 12d agoMaybe Gemini is loosey goosey on purpose so we angrily correct it - then feed something on the back end that trains a different model? It's so abysmally bad on Google search... and it's free. Isn't Google the great pioneer of the product is us?
- emodendroket 12d agoThe kind of results you get from the one on Google Search and a dedicated "Gemini Pro" response are totally different. I'm assuming that's cost savings.