Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
n_bhavikatti
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
n_bhavikatti
6mo ago
They actually did test GPT-5: https://www.science.org/doi/10.1126/science.aec8352 (see the figure under Conclusion). Its rate of endorsement of user action, 52%, was the same as GPT-4o. So based on their setup it
2.
▲
by
n_bhavikatti
6mo ago
In STEM/objective matters (math, science, coding), answers are more clearly defined as either right or wrong. This is where hallucination is more difficult/unlikely. But in personal matters, everything is subjective. AI tends to d
3.
▲
by
n_bhavikatti
7mo ago
The temperature clamp fix and "Optuna++" actions by the agents (the cause of basically all improvement to eCLIP) indicate they are good at finding bugs and hyper-parameter tuning. But when it comes to anything beyond that, such as