Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
blixt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
by
blixt
1y ago
Maybe for support but it’s a real world problem unrelated to language models that they do help me with. And ordering food at a restaurant is an age old problem, I just don’t enjoy making the call personally so I got value out of using a voi
62.
▲
by
blixt
1y ago
This is very interesting, and nice learnings in there too, thank you for sharing! It seems the author monitored the LLM, stopped it from going off-track a few times, fixed some unit test code manually, etc. Plus this is strictly re-implemen
63.
▲
by
blixt
1y ago
I think OpenAI is the first 100% AI-focused company to throw this many engineers (over 1,000 at this point?) at every part of the agentic workflow. I think it's a tremendous amount of discovery work. My theory would be that once we see
64.
▲
by
blixt
1y ago
We now have some very interesting elements that can become a workhorse worth paying hundreds of dollars for: - Reasoning models that can remember everything it spoke to the user about in the past few weeks* and think about a problem for 20
65.
▲
by
blixt
1y ago
I keep switching away from and back to Cursor (mainly due to frontier models explicitly fighting their apply model, the first few times it’s funny to see the LLM itself write “this is frustrating” but I digress). And every time I find it ha
66.
▲
by
blixt
1y ago
Let’s not forget that every round trip with the LLM costs latency (and extra input tokens). We now have parallel tool calls which sometimes works in some models[1]. But it’s great because now a model can say “write these 3 files then read t
67.
▲
by
blixt
1y ago
My list of uses of AI includes: - Turning a lot of data into a small amount of data, such as extracting facts from a text, translating and querying a PDF, cleaning up a data dump such as getting a clean Markdown table from a copy/paste
68.
▲
by
blixt
1y ago
I think it's in human nature to force any topic to be all "good" or "bad". I agree with most criticisms this author has about the performance of AI -- it _is_ very bad at writing essays, and dare I say most things (
69.
▲
by
blixt
1y ago
I actually thought the video they posted[1] would have information, but that was just 9 minutes of two guys congratulating each other. [1] https://twitter.com/OpenAI/status/1925235156157440438
70.
▲
by
blixt
1y ago
They mentioned "microVM" in the live stream. Notably there's no browser or internet access. It makes sense, running specialized Firecracker/Unikraft/etc microkernels is way faster and cheaper so you can scale it up.
71.
▲
by
blixt
1y ago
They certainly can automate their own SWE but I wonder if that’s as good as getting full computer use logs (terminal, web browsing, code acceptance/rejection, etc etc — as claimed in the linked article) from millions of individuals and
72.
▲
by
blixt
1y ago
> Enabled from the insight from our heavily-used Windsurf Editor, we got to work building a completely new data model (the shared timeline) and a training recipe that encapsulates incomplete states, long-running tasks, and multiple surfa
73.
▲
by
blixt
1y ago
I'm not sure I see the behavior in the Gemini 2.0 Flash model's image output as a strength. It seems to me it has multiple output modes, one indeed being masked edits. But it also seems to have convolutional matrix edits (e.g. &qu
74.
▲
by
blixt
1y ago
I remember as a young kid living in Norway, there'd be the "russefeiring"[0] around May where students finishing their final semester will don a brightly colored overall and cause mayhem in the town. I remember getting shot a
75.
▲
by
blixt
2y ago
I think interactions between many types and functions would be harder to migrate, especially if you go so granular as to target individual types independently. Maybe if it was a migration from some checkpoint (timestamp / version /
76.
▲
by
blixt
2y ago
It's easy to hate on the ways web development goes wrong, but I think it's really cool that our ecosystem has so many options for people to get started. New libraries, and thus new ways of doing things, push old established practi
77.
▲
by
blixt
2y ago
Seconding this. Also curious if this is done with microkernels (I put Unikraft high on the list of tech I'd use for this kind of problem, or possibly the still-in-beta CodeSandbox SDK – and maybe E2B or Fly but didn't have as good
78.
▲
by
blixt
2y ago
Needs a (2023) tag. But definitely the release of ARC2 and image outputs from 4o got me thinking about the JEPA family too. I don't know if it's right (and I'm sure JEPA has lots of performance issues) but seems good to have
79.
▲
by
blixt
2y ago
Yeah if we get an open model that one could apply a LoRA (or similarly cheap finetuning) to, then even problems like reproducing identity would (most likely) be solved, as they were for diffusion models. The coherence not just to the prompt
80.
▲
by
blixt
2y ago
Every component of a deep neural network is understood by many people, it's the interaction between the numbers trained that we don't always understand. Likewise, I would say that we understand the components on a CPU, and the ins
81.
▲
by
blixt
2y ago
This is super cool! I think new kinds of experiences can be built with infinite generative UIs. Obviously there will need to be good memory capabilities, maybe through tool use. If you end up taking this further and self hosting a model you
82.
▲
by
blixt
2y ago
Yeah I wasn’t very imaginative in my examples, with 4o you can also perform transformations like “rotate the camera 10 degrees to the left” which would be hard without a specialized model. Basically you can run arbitrary functions on the ex
83.
▲
by
blixt
2y ago
Still sounds a bit like we've seen it all already – dynamic linking introduced a lot of ways for software that wasn't buggy today to become buggy tomorrow. And Chrome uses an absurd amount of computing power (its bare minimum is m
84.
▲
by
blixt
2y ago
This argument could be made for every level of abstraction we've added to software so far... yet here we are commenting about it from our buggy apps!
85.
▲
by
blixt
2y ago
Yeah, it seems like somewhere in the semantic space (which then gets turned into a high resolution image using a specialized model probably) there is not enough space to hold all this kind of information. It becomes really obvious when you
86.
▲
by
blixt
2y ago
Yeah Gemini has had this for a few weeks, but much lower resolution. Not saying 4o is perfect, but my first few images with it are much more impressive than my first few images with Gemini.
87.
▲
by
blixt
2y ago
What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, t
88.
▲
by
blixt
2y ago
I don't really think the tone or the examples of this presentation are all that useful, and while the advice slide[1] seems well-intended, I just think most of it is only really known post-hoc because reality is not so clean cut. When
89.
▲
by
blixt
2y ago
What I found building multiplayer editors at scale is that it's very easy to very quickly overcomplicate this. For example, once you get into pub/sub territory, you have a very complex infrastructure to manage, and if you're
90.
▲
by
blixt
2y ago
I usually explained ECS to people unfamiliar with the concept more like a relational database. It's not a perfect 1:1, but I feel like it resonates well if you think of components as highly normalized tables, and entities being a table
More ›