Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Bjorkbat
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
Bjorkbat
2y ago
Can't believe I'm actually rooting for the copyright cartels in this fight. But that does make me think, that in a sane society with a functional legislature I wouldn't have to pick a dog in this fight. I'd have have en
62.
▲
by
Bjorkbat
2y ago
I've honestly thought of hacker spirit as embodying a kind of homesteader ethos in a way. There's this homesteading book I bought a long time ago when I was in college, rich with illustrations on how to do everything from raise a
63.
▲
by
Bjorkbat
2y ago
I'm amused by all the flags this article has. It reinforces this belief that "vibe-coding" isn't something that evolved organically, but was forced. I wouldn't go as far as to call it "astroturfed", I be
64.
▲
by
Bjorkbat
2y ago
Tangential, but this reminds me of something someone said on Twitter that has resonated with me ever since. Startups targeting developers / building developer tooling are arguably one of the worst startups to build, because no matter
65.
▲
by
Bjorkbat
2y ago
My first thought seeing this and looking at benchmarks was that if it wasn’t for reasoning, then either pundits would be saying we’ve hit a plateau, or at the very least OpenAI is clearly in 2nd place to Anthropic in model performance. Of c
66.
▲
by
Bjorkbat
2y ago
I'm going to add a bit of a left-field contribution since his work is less generative coding more mathematics and geometry in general, but it has inspired me when it comes to generative coding. I'm assuming the works he's do
67.
▲
by
Bjorkbat
2y ago
Related, if you like startup simulators, then you'll love The Founder ( http://thefounder.biz/ ). Incidentally Francis Tseng also worked on Half-Earth Socialism, which is also quite fun, and both games happen to be open-
68.
▲
by
Bjorkbat
2y ago
Yeah, I have to admit it looks kind of "uncanny", but at the same time it kind of has a sense of personality to it that other humanoid robots lack, in part because it has face with two distinct eyes. The knit "suit" is a
69.
▲
by
Bjorkbat
2y ago
Yeah, I saw the pass@7 figure as well, and I'm not sure what to make of it. On the one hand, solving nearly half of all tasks is impressive. On the other hand, a machine that might do something correctly if you give it 7 attempts i
70.
▲
by
Bjorkbat
2y ago
>But sonnet solved over 25% of them and made 60 grand. Technically it didn’t since all these tasks were done some time ago. On that note, I feel like putting a dollar amount on the tasks it was able to complete is misleading. In the rea
71.
▲
by
Bjorkbat
2y ago
I think that's a premature conclusion to take from this benchmark. Something to keep in mind is that Expensify is kind of an anomaly in that it hires freelancers by creating a well-articulated Github issue and telling them to go solve
72.
▲
by
Bjorkbat
2y ago
To be fair I'm pretty sure Gary Marcus did in fact bring that up at one point, and if I recall correctly the incident was embarrassing enough for Google to briefly pause the generation of AI images until it could correct the awkward bi
73.
▲
by
Bjorkbat
2y ago
My belief is that software engineering benchmarks are still a poor proxy for performance on real world software engineers tasks, and that there's a decent chance a new model might saturate a benchmark while being kind of underwhelming.
74.
▲
by
Bjorkbat
2y ago
This kind of reminds me of when there was a lot of hype around messenger apps and this idea that we'd just do everything through a chat interface / chat bot. It never panned out, arguably because the technology wasn't quite t
75.
▲
by
Bjorkbat
2y ago
Before people lose their minds on AlphaGeometry, I thought I'd share this gem the r/math subreddit lending some insight into how the original AlphaGeometry appears to work from the perspective of someone far more literate in math
76.
▲
by
Bjorkbat
2y ago
Most of my observations have been that people are using it to make personal software. That is to say, software with an intended user base of just yourself and maybe friends and family. For software meant to be consumed by the masses it
77.
▲
by
Bjorkbat
2y ago
This kind of reminds me of back when the hype cycle was focused on Messenger apps and the idea of most online behavior being replaced with a chatbot. God I hated the smug certainty of (some, definitely not all!) UX designers at the time pr
78.
▲
by
Bjorkbat
2y ago
So, I still think this is a cool tool for search reasons, but otherwise the tendency to hallucinate makes it questionable as a researcher. Hypothetically speaking, if the time you saved is now spent verifying the statements of your AI resea
79.
▲
by
Bjorkbat
2y ago
Actually sounds pretty cool, but the graph on expert level tasks is confusing my expectations. Saying it has a pass rate of less than 20% sounds a lot like saying this thing is wrong most of the time. Granted, these strike me as difficult t
80.
▲
by
Bjorkbat
2y ago
So am I to understand that they used their internal tooling scaffold on the o3(tools) results only? Because if so, I really don't like that. While it's nonetheless impressive that they scored 61% on SWE-bench with o3-mini combine
81.
▲
by
Bjorkbat
2y ago
I have to admit I'm kind of surprised by the SWE-bench results. At the highest level of performance o3-mini's CodeForces score is, well, high. I've honestly never really sat down to understand how elo works, all I know is t
82.
▲
by
Bjorkbat
2y ago
I'm not going to say that this isn't creative coding, since that sounds elitist. What I will say is that this is a very shallow interpretation of creative coding. With very little learning and modest effort you can make very fu
83.
▲
by
Bjorkbat
2y ago
I'm personally kind of anti-NFT myself, but I have to admit I do like the generative art NFT scene, especially back when Hic Et Nunc was around and people were selling fun little things they made for the equivalent of $5 of TEZ. To me
84.
▲
by
Bjorkbat
2y ago
This is in line with my perspective that writing as a form of self-reflection and thinking out loud is underrated. Likewise, programming as a form of thinking out loud is also underrated. For complex problems, I find that it doesn't m
85.
▲
by
Bjorkbat
2y ago
I'm still skeptical on the notion that we can remove the human bottleneck on code because code has verifiable solutions. It's true only to the extent that there's sufficient test coverage to prevent any unwanted side effects.
86.
▲
by
Bjorkbat
2y ago
Glad someone brought this up. I'm personally fine with o3 being tuned on the train set as a way to teach models "the rules of the game", what annoys me is that this wasn't also done with the o1 models or r1. It's a
87.
▲
by
Bjorkbat
2y ago
I have to admit it's really hard for me to take the noise about agents seriously because they keep bringing up using it for things that don't feel like chores to me, but are actually kind of delightful. I like the anticipation an
88.
▲
by
Bjorkbat
2y ago
Can't take this seriously knowing that this is the same Mustafa Suleyman who... - Was basically acqui-hired by Microsoft from Pi AI (seems a little biased to recommend a book from one of your own) - Left DeepMind due to allegations of
89.
▲
by
Bjorkbat
2y ago
Tangentially related, I get strong Metaverse/NFT vibes around predictions on agents. Namely, a lot of predictions were made around NFTs that just didn’t make sense or were kind of dumb. My pet favorite was this notion that in the futu
90.
▲
by
Bjorkbat
2y ago
If I recall correctly the authors of the benchmark did mention on Twitter that for certain issues models will submit an answer that technically passes the test but is kind of questionable, so yeah, good point.
More ›