Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
felix089
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
felix089
7mo ago
This app cracked the GEO code
32.
▲
by
felix089
7mo ago
Thanks! Yea I think the best ones are when science is actually quite clear but politics get in the way so you see their bias
33.
▲
by
felix089
7mo ago
thanks happy to hear. Yes for debate mode the max number of models is actually only 6. More than that didn't really add anything in my preliminary test. Only for direct comparison in the poll mode you can choose up to 50, then it'
34.
▲
by
felix089
7mo ago
Whoever just asked this, very funny: https://opper.ai/ai-roundtable/questions/does-mr-krabs-evade...
35.
▲
by
felix089
7mo ago
I actually asked this question before posting, just to be sure... edit: their reply is quite funny actually "In a display of absolute consensus, the AI Roundtable unanimously validated its own existence,"
36.
▲
Show HN: AI Roundtable – Let 200 models debate your question
(opper.ai)
118 points
by
felix089
7mo ago
|
98 comments
37.
▲
by
felix089
8mo ago
This looks interesting, what's your goal with this project?
38.
▲
by
felix089
8mo ago
Yea, I thought the same before the test and was pretty surprised. But RE the data, it's actually not a gig platform where people get paid. Rapidata answered this in another comment below. They integrate micro-surveys into mobile apps (
39.
▲
by
felix089
8mo ago
Because it's interesting to me, it doens't mean they have to share them publicly btw
40.
▲
by
felix089
8mo ago
What were you trying to test here?
41.
▲
by
felix089
8mo ago
agreed
42.
▲
by
felix089
8mo ago
Sounds interesting, would be nice to see the questions if you're open to sharing?
43.
▲
by
felix089
8mo ago
This is amazing thanks for sharing!
44.
▲
by
felix089
8mo ago
Good question, I used the API defaults across the board since it felt like the most reasonable baseline to compare. Flash lite getting 10/10 was definitely very surprising
45.
▲
by
felix089
8mo ago
Yea funny coincidence, but this is not at all how the human answers were collected. Rapidata answered this in another comment below. They integrate micro-surveys into mobile apps (like Duolingo, games, etc) as an optional opt-in instead of
46.
▲
by
felix089
8mo ago
Flash lite succeeded in every test, smth got lost in editing, just updated it. thx!
47.
▲
by
felix089
8mo ago
Good catch, something got lost in editing just updated, flash lite succeeded in every test, which is pretty surprising!
48.
▲
by
felix089
8mo ago
They answered it in another comment somewhere below, there's no incentive for a correct answer
49.
▲
by
felix089
8mo ago
my take as well, reliablity is the biggest concern, with more context available during inference or orchestration like yours it definitely gets better
50.
▲
by
felix089
8mo ago
Interesting find!
51.
▲
by
felix089
8mo ago
Great idea honestly! I wonder how long it'll take until they are able to solve these reliably though
52.
▲
by
felix089
8mo ago
Agreed, it makes me wonder what other logic tests / evals can be built from this to expand this type of evaluation.
53.
▲
by
felix089
8mo ago
They are amazing, super fast turnaround for the data also
54.
▲
by
felix089
8mo ago
Sonnet 4.6 wasn't part of the test in my case but would be interesting to see the baseline responses. It might be that it gets it right regardless, but will have to test it.
55.
▲
by
felix089
8mo ago
Good point on the noise, that might be it
56.
▲
by
felix089
8mo ago
Exactly, same pattern across almost every failure, but sonar models, which just go wild
57.
▲
by
felix089
8mo ago
Interesting, which Gemini model? And how did you ask for symbolic reasoning, just added it to the prompt?
58.
▲
by
felix089
8mo ago
It's similar for me, it generates so much content without me asking. if I just ask for feedback or proofreading smth it just tends to regenerate it in another style. Anything is barely good to go, there's always something it wants
59.
▲
by
felix089
8mo ago
I don't think it qualifies as a stupid question either, it does make sense
60.
▲
by
felix089
8mo ago
If you were forced to answer either or, which one would you pick? I think that's where the interesting dynamic comes from. Most humans would pick drive, also seen in the human control, even if it is lower that I thought it'd be
More ›