Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
qt31415926
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
How good are frontier models at physics?
(arxiv.org)
95 points
by
qt31415926
17d ago
|
50 comments
2.
▲
by
qt31415926
17d ago
Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks" John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks ar
3.
▲
by
qt31415926
25d ago
You're mistakening Tristan Buckmaster for Sebastien Bubeck. Seb is the one where there's at least 2 (unless the personal friend is Dheeraj) allegations, not Tristan
4.
▲
by
qt31415926
4mo ago
Isn't that what's solved by this method? Your SSO provider (e.g. Okta) is now what gates each employee's resource access for different MCP resources.
5.
▲
by
qt31415926
1y ago
the commenter never said they came up with nothing, they said o3 came up with something better.
6.
▲
by
qt31415926
2y ago
On our apps we consistently see a p50 3-4x speed difference between iOS and Android (though there are more lower end android devices). Hard to fathom if it's all due to variability in android devices vs RN being less performant on Andr
7.
▲
by
qt31415926
2y ago
Which parts of reasoning do you think is missing? I do feel like it covers a lot of 'reasoning' ground despite its on the surface simplicity
8.
▲
by
qt31415926
2y ago
> React doesn't make you a better developer, it makes you a better React developer. React's pure component functional style translates really well to nearly every other type of software development.
9.
▲
by
qt31415926
2y ago
808 ELO was for GPT-4o. I would suggest re-reading more carefully
10.
▲
by
qt31415926
3y ago
A lot of the internet would break if YouTube removed/tweaked their embedded video player so I doubt he has to worry.
11.
▲
by
qt31415926
3y ago
He's pointing out a motte/bailey meaning he's against the motte. If to him e/acc is the motte, then he's likely against
12.
▲
by
qt31415926
3y ago
It's actually impressive that Weird Al has made it as far as he did now that I think about it
13.
▲
by
qt31415926
3y ago
FYI, it was sarcasm (the emoji at the end is the giveaway)
14.
▲
by
qt31415926
3y ago
He groups them together because ultimately the result is that the science can't be trusted. He doesn't go so far to claim that one was intentionally faked vs gross incompetence.
15.
▲
by
qt31415926
3y ago
I don't think the statement reads that the 44% and the 26% should be additive. Especially given the zombie graphic where it looks like they overlap the 26 on top of the 44, where the orange bar is the 26% and the remaining yellow bar i
16.
▲
by
qt31415926
3y ago
Thanks, didn't know. This article seems to have been written by a reader though so not the same authors of the sketchy lithium work
17.
▲
by
qt31415926
3y ago
I'm impressed you were able to not get tangled up by the youtube/captcha/color hex/roman numeral mess! The youtube one is what screwed me over and over out of all my attempts
18.
▲
by
qt31415926
3y ago
I don't think it's easy. Verification is much easier than generating correct solutions for this. Looking at the JS, these rules use RNG such that you can have an inconsistent or impossible password. E.g. if the only youtube video
19.
▲
by
qt31415926
3y ago
Hmm I looked into it, and looked at papers/pdfs in google scholar's advanced search with her as an author that mentioned LLMs or GPT in the past 3 years. Every single one was a criticism about how they couldn't actually under
20.
▲
by
qt31415926
3y ago
In her field doesn't mean that's what she researches, LLMs are loosely in her field but the methods are completely different. Computational linguistics != deep learning. Deep learning does not directly use concepts from linguisti
21.
▲
by
qt31415926
3y ago
Her field has also taken the largest hit from the success of LLMs and her research topics and her department are probably no longer prioritized by research grants. Given how many articles she's written that have criticized LLMs it'
22.
▲
by
qt31415926
3y ago
Current topic aside, I feel like that stochastic parrots paper aged really poorly in its criticisms of LLMs, and reading it felt like political propaganda with its exaggerated rhetoric and its anemic amount of scientific substance e.g. >
23.
▲
by
qt31415926
4y ago
oh my bad, totally misread that
24.
▲
by
qt31415926
4y ago
It's trained on pre-2021 data. Looks like they tested on the most recent tests (i.e. 2022-2023) or practice exams. But yeah standardized tests are heavily weighed towards pattern matching, which is what GPT-4 is good at, as shown by it
25.
▲
by
qt31415926
4y ago
Curious since it does well on the LSAT, SAT, GRE Verbal.
26.
▲
by
qt31415926
4y ago
For Vietnamese, ChatGPT has been much more effective for me. You can tell it very specific things to modify as well as give it additional context which you can't really do with Google Translate. Especially with pronouns, google will ju
27.
▲
by
qt31415926
9y ago
makes sense, so they just looped that segment. thanks!
28.
▲
I'm not a conspiracy theorist, but what is happening in the SpaceX livestream?
2 points
by
qt31415926
9y ago
|
2 comments
29.
▲
by
qt31415926
9y ago
This is such an amazing post! So resourceful, thank you so much.
30.
▲
by
qt31415926
13y ago
Given his upbringing, I would say he never had the potential in the first place. From how he talks about his time at NASA, it seems like he was used as an equation-solving tool 10 years. “At that time, I led my life like a machine ― I woke
More ›