Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Bjorkbat
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
Bjorkbat
1y ago
Ooh, fun, a FizzBuzzFeed quiz!
32.
▲
by
Bjorkbat
1y ago
Sure, but if you're trying to get there by training a model on video games then you're likely going to wind up inadvertently creating a video game simulator rather than a physics simulator. I don't doubt they're trying t
33.
▲
by
Bjorkbat
1y ago
Also, to drive my point further home, in one of the demos they were operating a jetski during a festival. If the jetski bumps into a small Chinese lantern, it will move the lantern. Impressive. However, when the jetski bumped into some s
34.
▲
by
Bjorkbat
1y ago
Genuinely technically impressive, but I have a weird issue with calling these world simulator models. To me, they're video game simulator models. I've only ever seen demos of these models where things happen from a first-person o
35.
▲
by
Bjorkbat
1y ago
Somewhat related, but I’ve been feeling as of late what can best be described as “benchmark fatigue”. The latest models can score something like 70% on SWE-bench verified and yet it’s difficult to say what tangible impact this has on actual
36.
▲
by
Bjorkbat
1y ago
Alright, it does look pretty charming, and I especially like that it's open-source since pretty much anyone buying a domestic robot is likely to be a tinkerer of some sort, but at the same time it reminds me of the Jibo ( https:/&
37.
▲
by
Bjorkbat
1y ago
Yeah, this is something I find kind of tricky. I definitely believe that AI companies should get permission from rightsholders to train on their works, but actually compensating them for their works seems pointless. To make the royalties
38.
▲
by
Bjorkbat
1y ago
I recall it as less an evolution and more a complete tonal shift the moment o3 was evaluated on ARC-AGI. I remember on Twitter Sam made some dumb post suggesting they had beaten the benchmark internally and Francois calling him out on his
39.
▲
by
Bjorkbat
1y ago
Personally I think a more effective analogy would be if someone used a textbook and created an online course / curriculum effective enough that colleges stop recommending the purchase of said textbook. It's honestly pretty diffic
40.
▲
by
Bjorkbat
1y ago
Something missed in arguments such as these is that in measuring fair use there's a consideration of impact on the potential market for a rightsholder's present and future works. In other words, can it be proven that what you are
41.
▲
by
Bjorkbat
1y ago
Honestly feels like the whole Soham Parekh thing on Twitter is one giant joke with the one sincere / honest remark being the original from @Suhail. Like, I can't wrap my head around this many people having some kind of experience
42.
▲
by
Bjorkbat
1y ago
Yeah, I was about to say, it sounds a lot like this guy is just riding an intense high from getting Claude to build some side-project he's been putting off, which I feel is like 90% of all cases where someone writes a post like this. B
43.
▲
by
Bjorkbat
1y ago
> These are just good chess moves. The "super-intelligence" bit is just hype/spin for the journalists and layperson investors. Which is kind of what I figured, but I was curious if anyone disagreed.
44.
▲
by
Bjorkbat
1y ago
It's a smart purchase, it's just that I don't see how these datasets factor into super-intelligence. I don't think you can create a super-intelligent AI with more human data, even if it's high-quality data from pai
45.
▲
by
Bjorkbat
1y ago
Fair enough. If you aren't willing to give your friend $14 billion to join your company so you can hang out more, then are you two really friends?
46.
▲
by
Bjorkbat
1y ago
I don't actually think this is the case, but nonetheless I think it would be kind of funny if LLMs somehow "discovered" linguistic relativity ( https://en.wikipedia.org/wiki/Linguistic_relativity ).
47.
▲
by
Bjorkbat
1y ago
Point remains though, they crushed the benchmark using a specialized model that you’ll probably never have access to, whether personally or through a company. They inflated expectations and then released to the public a model that underperf
48.
▲
by
Bjorkbat
1y ago
Related, when o3 finally came out ARC-AGI updated their graph because it didn’t perform nearly as well as the version of o3 that “beat” the benchmark. https://arcprize.org/blog/analyzing-o3-with-arc-agi
49.
▲
by
Bjorkbat
1y ago
He's still a great designer, the problem though is that without the right kind of editorializing force he'll make mistakes, usually in the form of compromising practicality and functionality for the sake of aesthetics. I should p
50.
▲
by
Bjorkbat
1y ago
I was about to comment on a past remark from Anthropic that the whole reason for the convoluted naming scheme was because they wanted to wait until they had a model worth of the "Claude 4" title. But because of all the incremental
51.
▲
by
Bjorkbat
1y ago
A while back someone here on Hacker News made a pretty insightful comment that as great of a designer as Jony Ive is, a large part of his success is owed to the fact that he had an "editor" in the form of Steve Jobs. Once Jobs pa
52.
▲
by
Bjorkbat
1y ago
My weird theory: the alpha releases are more expensive than people realize, and OpenAI can afford to launch money-losing alphas on Pro because few users have Pro accounts and the lose less money per Pro user. If anyone could use Codex, you’
53.
▲
by
Bjorkbat
1y ago
I kind of had the feeling LLMs would be better at Python vs other languages, but wow, the difference on Multi SWE is pretty crazy.
54.
▲
by
Bjorkbat
1y ago
Fair point, we use metaphor to explain and understand a variety of topics, and a lot of those metaphors are best understood through pop culture analogies. A reasonable compromise then is that you can train an AI on Wikipedia, more-or-less.
55.
▲
by
Bjorkbat
1y ago
I broadly agree in that sure, unfettered access to copyrighted material will AI more capable, but more capable of what exactly? For national security reasons I'm perfectly fine with giving LLMs unfettered access to various academic pub
56.
▲
by
Bjorkbat
1y ago
I mean, I don't disagree with you when you say that something that would take an hour or more to implement would only take 10 minutes or so with AI. That kind of aligns with my personal experience. If something takes an hour, it'
57.
▲
by
Bjorkbat
1y ago
I'm always a little bit skeptical whenever people say that AI has resulted in anything more than a personal 50% increase in productivity. Like, just stop and think about it for a second. You're saying that AI has doubled your pro
58.
▲
by
Bjorkbat
1y ago
Makes me think of the horse browser ( https://gethorse.com ), namely the fact that unlike pretty much all other browsers it's a paid, subscription product. You actually have to pay $60 a year in order to use it. Sounds absol
59.
▲
by
Bjorkbat
1y ago
/vg/ also had a pretty cool amateur game dev general thread (/agdg/). No one was making any hidden gems there, but it wasn't trash either. At any rate, I liked it.
60.
▲
by
Bjorkbat
2y ago
Minor pet-peeve of mine, I really don't like the term "superforecaster". First time I encountered it was in association with some guy who was making predictions a year or two out. Which to be fair it actually is kind of impr
More ›