Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Bjorkbat
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
91.
▲
by
Bjorkbat
2y ago
I think it's still an interesting way to measure general intellience, it's just that o3 has demonstrated that you can actually achieve human performance on it by training it on the public training set and giving it ridiculous amou
92.
▲
by
Bjorkbat
2y ago
It's easy to miss, but if you look closely at the first sentence of the announcement they mention that they used a version of o3 trained on a public dataset of ARC-AGI, so technically it doesn't belong on this list.
93.
▲
by
Bjorkbat
2y ago
To me it's very frustrating because such little caveats make benchmarks less reliable. Implicitly, benchmarks are no different from tests in that someone/something who scores high on a benchmark/test should be able to gene
94.
▲
by
Bjorkbat
2y ago
> Also, 1 odd thing I noticed is that the graph in their blog post shows the top 2 scores as “tuned” Something I missed until I scrolled back to the top and reread the page was this > OpenAI's new o3 system - trained on the ARC-A
95.
▲
by
Bjorkbat
2y ago
It's not so much the cost as much the fact that they got a slightly better result by throwing 172x more compute per/task. The fact that it may have cost somewhere north of $1 million simply helps to give a better idea of how absu
96.
▲
by
Bjorkbat
2y ago
I was impressed until I read the caveat about the high-compute version using 172x more compute. Assuming for a moment that the cost per task has a linear relationship with compute, then it costs a little more than $1 million to get that sco
97.
▲
by
Bjorkbat
2y ago
That’s pretty common and not actually limited to H1-Bs I know universities will do this with certain open positions where they already have a candidate in mind but are required to advertise an opening, can’t remember the specifics why thoug
98.
▲
by
Bjorkbat
2y ago
I always assumed firing one of these things was a suicide mission, but some NATO guy on Quora has suggested otherwise, and it seems pretty convincing. https://www.quora.com/Would-the-Davy-Crockett-tactical-nucle... Basicall
99.
▲
by
Bjorkbat
2y ago
Like I said, 3D printers have gotten incrementally better. Buddy of mine makes high-quality prints on one that’s way cheaper than what I owned back in the day. And yet nothing has really changed because he’s still using it to print dumb tc
100.
▲
by
Bjorkbat
2y ago
> I am constantly surprised how prevalent this attitude is. ChatGPT was only just released in 2022. Is there some expectation that these things won't improve? I mean, in a way, yeah. Last 10 years were basically one hype-cycle after
101.
▲
by
Bjorkbat
2y ago
I'm weirdly not too surprised due to this belief I have that software developers would make effective criminals. A lot of this boils down to a belief I have that not getting caught in the first place is easy. Murders have something l
102.
▲
by
Bjorkbat
2y ago
A tangent perhaps, but I've felt this with AI, albeit with nuanced differences. In a nutshell, there's this weird tension between being kind of dismissive of it due to failed personal expectations, but also a desire to be as obje
103.
▲
by
Bjorkbat
2y ago
I'm really curious to see how this pricing plays out. I constantly hear on Twitter how certain influencers would be more than willing to pay more than $20/month for unlimited access to the best models from OpenAI / Anthropic
104.
▲
by
Bjorkbat
2y ago
As someone who's struggled to really get into podcasts, I'm convinced that most people who enjoy podcasts don't really actively listen to them, they just like that extra bit of noise in the background while they do something
105.
▲
by
Bjorkbat
2y ago
Not meeting expectations != not better than the previous models. The Information reporting was a bit more clear on this. Orion is better than GPT-4, it's just that they were expecting a leap in capabilities comparable to what we saw g
106.
▲
by
Bjorkbat
2y ago
It's kind of, I don't know, "weird", observing how there's all these news outlets reporting on how essentially every up-and-coming model has not performed as expected, while all the employees at these labs haven
107.
▲
by
Bjorkbat
2y ago
I agree that existing benchmarks are no longer useful now that there's basically nothing left in them that seems to stump LLMs. But when I hear that models are failing to meet expectations, I imagine what they're saying is that th
108.
▲
by
Bjorkbat
2y ago
Kind of reminds me of an idle thought I have every now and then. Between the sheer difficulty of establishing any kind of foothold on Mars, and the vast amount of uninhabited land, it’s curious that more thought hasn’t been given into the
109.
▲
by
Bjorkbat
2y ago
Tried my standard go-to for testing, asked it to generate a voronoi diagram using p5js. For the sake of job security I'm relieved to see it still can't do a relatively simple task with ample representation in the Google search re
110.
▲
by
Bjorkbat
2y ago
As someone who otherwise hates genAI, I must admit, this is actually a very cool demo and a very sensible application of AI.
111.
▲
by
Bjorkbat
2y ago
Kind of thrown off by the Phil Fish comments. Like, are we talking about the guy who made Fez? That Phil Fish? Is it 2012?
112.
▲
by
Bjorkbat
2y ago
As much of a curmudgeon as I am on AI I do sincerely believe that one area that it's effective at is going from absolute beginner to decent understanding on something completely new and foreign. Recently I've been somewhat curious
113.
▲
by
Bjorkbat
2y ago
A complete tangent, but I think a big reason why I'm kind of dismissive of AI is because people who speculate on what it would enable make it honestly sound kind of unimaginative. > AI models will soon serve as autonomous personal a
114.
▲
by
Bjorkbat
2y ago
This is going to be the real test for Waymo. From anecdotal experience, Austin has more inclement weather, and its road infrastructure is lacking in some parts of town.
115.
▲
by
Bjorkbat
2y ago
Yeah, now that you mention it I also see that. It was clearly meant to spawn after 3 seconds. Seems on successive attempts it also doesn't quite wait 3 seconds. I'm kind of curious if they did a little bit of editing on that one
116.
▲
by
Bjorkbat
2y ago
The idea of Github having a unique "taste" advantage resonates with me a lot. I don't like the fact that Github is using my code to feed Microsoft's AI ambitions, but I dislike Bitbucket and Gitlab more simply on the gr
117.
▲
by
Bjorkbat
2y ago
This is far, far more rigorous than the experiment Microsoft behind Microsoft's claim that Copilot made devs 55% faster. The experiment in question was to split 95 devs into two groups and see how long it took each group to setup a web
118.
▲
by
Bjorkbat
2y ago
Worth mentioning that LlaMa 70b already had pretty high benchmark scores to begin with https://ai.meta.com/blog/meta-llama-3-1/ Still impressive that it can beat top models with fine-tuning, but now I’m mostly imp
119.
▲
by
Bjorkbat
2y ago
I never developed with Flash, but I grew up on Flash media, in particular old Newgrounds Flash games / movies. I remember some of the pains of the day. The fact that Flash media usually came with a loading bar is seems quaint compare
120.
▲
by
Bjorkbat
2y ago
Related https://thehistoryofweb.design/ The book is an interesting read. The death of Flash feels like the start of a dark age for web design when you compare what came before with what came now. Granted, this may be less
More ›