Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
5555watch
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
5555watch
2mo ago
In my understanding the first Deep Think / Pro models were already very good as they were doing some kind of parallel repeated reasoning, thus were slow and expensive. So if chatjimmy speeds enables a fast deep think level performance,
32.
▲
by
5555watch
2mo ago
I mean in their code the simulation is incorrect. Someone found the source: https://github.com/pem725/Dunning-Kruger They're sampling slope and bias for the line from U[0,1] and U[0,100], the expected values of wh
33.
▲
by
5555watch
2mo ago
It's an interesting read. Curiously, it doesn't really debunk anything. The fact that given X and Y random and independent, that Y-X is correlated with X doesn't disprove the Dunning Kruger. It in fact proves that Y = 1 X is
34.
▲
by
5555watch
2mo ago
The extrapolation can also be a learned skill, especially in math. How many papers took result X, extended it to Y using known building blocks, and applied to Z. By the way, convex hull permits extrapolating past the training data. LLM won&
35.
▲
by
5555watch
2mo ago
They're simulating randomness incorrectly: relationship between true and perceived will average 0.5, not 0; and bias will average 50%, not 0%. That's why their "random data" is sloped. Add negative relationship and negat
36.
▲
by
5555watch
2mo ago
Thanks for finding the code! Now it's much more clear. The simulated data tries generating the true relationship between actual and perceived scores from 0.0 to 1.0, and bias in self-reporting from 0% to 100%. So the output graph shoul
37.
▲
by
5555watch
2mo ago
So what does that change? If there's no error bars on the graphs you can't discuss significant differences easily. And if they do differ, plotting differences will yield the U shape curve.
38.
▲
by
5555watch
2mo ago
Very hard to understand the meat behind all the fluff of the article, especially as the simulation code is not available, and as the presented simulated and original graphs are effectively the same (I don't see a disagreement). It'
39.
▲
by
5555watch
2mo ago
I like the illustration that the models are working on a convex hull of known information. Filling gaps with linear combinations of known facts and results. They can't exit the hull until the "intuition" starts spawning point
40.
▲
by
5555watch
2mo ago
For reference, Safari is doing around 53, and Chrome is around 60 for M5, of which they were bragging about [0]. So curious to see whether there's any gains from your custom browser. [0]: https://blog.google/chromium&#x
41.
▲
by
5555watch
2mo ago
Aha, I'm getting 97/100 on Safari
42.
▲
by
5555watch
2mo ago
What's the Speedometer 3.1 [0] rating for that browser paired with an M5? [0]: https://browserbench.org/Speedometer3.1/
43.
▲
by
5555watch
2mo ago
But it's not proactive then? There's just a hidden loop or some cron-like signal that feeds data to a reactive loop.. I'd imagine proactive as something like, hmm, no signal from X, I wonder how they're doing...
44.
▲
by
5555watch
2mo ago
I wonder into what weird research niche would all such AI schools converge to after enough time.
45.
▲
by
5555watch
2mo ago
I'm curious, how hard/expensive it is to burn a really large model into silicon, and why aren't we doing this already? Or, when we will start doing this, who's going to be able to do that in scale? I'm seeing the TA
46.
▲
by
5555watch
2mo ago
It's a division by zero. We can change infinity to 'undefined'
47.
▲
by
5555watch
2mo ago
I used to chat with "the paper" until both me and the AI were roughly on the same page at what's happening. It's important, as a "full read" of the "full paper" is not always used efficiently if not d
48.
▲
by
5555watch
2mo ago
I'm probably a weird outlier. Coming from academia, it's ranging from 1x to infinity-x (as in, certain tasks wouldn't even be touched if not for AI). For stuff that I'm used to (R) I can write nice and compact spaghetti
49.
▲
by
5555watch
2mo ago
I love that they let you switch to a more common q'k notation!
50.
▲
by
5555watch
3mo ago
A silver lining to think about.. Currently there's a big issue (at least in some countries) that mathematicians must publish from 5 to 7-8 papers within a 5-year window to keep the tenure. So it's a minimum, much more is expected
51.
▲
by
5555watch
3mo ago
So with this release do they kill the 5.5-Pro model with super long thinking and reasoning? 5.6-Sol-Ultra is not the equivalent, right?
52.
▲
by
5555watch
4mo ago
Will it also have hardcoded self-lobotomy if asked about cutting edge ML or LLM solutions? (Looking at Fable here)
53.
▲
by
5555watch
4mo ago
I think it's also important and heavily overlooked to develop and maintain open source "pro" level models. Those that are able to think for 80 minutes and yield heavy solutions. I'm not an expert in LLMs so it's har
54.
▲
by
5555watch
4mo ago
> model except to limit its effectiveness in developing frontier LLMs Does this imply that they're actively using it for their frontier development and that it's very effective?
55.
▲
by
5555watch
4mo ago
How are you grading the student submissions? Also, do you catch students who fully use AI and don't follow the Honor code? If so, how?
56.
▲
by
5555watch
5mo ago
My guess is that with the new UI they've either mistakenly or deliberately reduced the thinking effort or the budget, even if they write "extended thinking" in the app. And it seems that they're (mistakenly or deliberate
57.
▲
by
5555watch
5mo ago
I would guess it's because ChatGPT Pro allows for 80min "think". I've never had even remotely similar think times with Gemini Deep Think. It's generally around 10-15min for math problems, and get increasingly shorte
58.
▲
by
5555watch
6mo ago
Why no (high) variants in the comparison models?
59.
▲
by
5555watch
6mo ago
It's not "13 parameters to reason", they just rotated the full 8B parameter space in 13 dimensions and found a rotation that was still able to reason. Depending on the latent structure, it's possible a nice rotation that