Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
libraryofbabel
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
libraryofbabel
5mo ago
> She says explicitly it's not an empirical hypothesis. It's just a label for how they function. Then… what’s the point of the label, if it’s not making any empirically-meaningful claims about LLMs at all? I know that LLMs invo
32.
▲
by
libraryofbabel
5mo ago
It would have been nice to see some version of “I am very surprised by how far LLMs have come since I wrote the stochastic parrots paper, here is how I have revised my thinking.” But there is nothing like that and the author is just doublin
33.
▲
by
libraryofbabel
5mo ago
That’s correct, and yes - not less compute total on the main model (actually slightly more, since checking failed draft tokens costs you compute), but faster because inference is memory-bandwidth bound. And like you I also think of it as li
34.
▲
by
libraryofbabel
5mo ago
Speculative decoding is an amazingly clever invention, almost seems-too-good-to-be-true (faster interference with zero degradation from the quality of the main model). The core idea is: if you can find a way to generate a small run of dra
35.
▲
by
libraryofbabel
5mo ago
> Do you think you could remember most of Z80 ASM? I find when you learn things at 15 they tend to stick around. (Stuff I learned last week, not so much!) Even just looking at your example, I remembered that HL is a 16 bit register and y
36.
▲
by
libraryofbabel
5mo ago
Much nostalgia. The TI-83 Z80 was how I learned assembly as a teenager, so I could write better calculator games than was possible with TI Basic. Many others here had a similar experience, I’m sure. It’s been a couple decades, but I’m sure
37.
▲
by
libraryofbabel
5mo ago
So, to reiterate my example: you'd have been fine with people claiming in 2019 that we would eventually scale LLMs to the capabilities of Opus 4.7 + Claude Code? Because I would have said then that was a fantasy, because "LLMs are
38.
▲
by
libraryofbabel
5mo ago
I think this is a case of that mildly apocryphal Richard Feynman quote: "if you think you understand quantum mechanics, you don't understand quantum mechanics." I understand LLM architecture internals just fine. I can write y
39.
▲
by
libraryofbabel
5mo ago
> that specific version we're aligning toward is just the only one that makes some kind of rational sense, among a trillion of other meaningless gibberish-producing ones. Oh, the space of possibilities is unimaginably vaster than th
40.
▲
by
libraryofbabel
5mo ago
LLM’s aren’t software (except in an uninteresting obvious sense); they are “grown, not made” as the saying is. And sure, they can find which weights activate when goblins come up (that’s basic mechanistic interpretability stuff), but it’s n
41.
▲
by
libraryofbabel
5mo ago
It's interesting that some people are responding to your comment as if this proves that AI is a sham or a joke. But I don't think that's what you're saying at all with your reference to Terence McKenna: this is a serio
42.
▲
by
libraryofbabel
5mo ago
The author did give him credit for the whole you-can-make-the-fries-super-salty-to-increase-demand-for-drinks thing in Theme Park, which I remember vividly. (I, too, dropped many hours on Theme Park as a kid.) Although I imagine there’s abo
43.
▲
by
libraryofbabel
5mo ago
Oh wow, you blow my mind with your linguistic erudition; I had no idea it was possible to use male-gendered terms in a generic way! Well, all is forgiven, then. Seriously, just... don't? This isn’t some woke political thing and I disli
44.
▲
by
libraryofbabel
5mo ago
Not a “bro” (there are women on this site you know), and perhaps you’re missing the British understatement in my “maybe a little too uncritical of its subject” line. Obviously the book is totally biased in favor of Hassabis and Deepmind
45.
▲
by
libraryofbabel
5mo ago
Is anyone else reading Sebastian Mallaby’s new book about Demis and Deepmind: The Infinity Machine: Demis Hassabis, DeepMind, and the Quest for Superintelligence ? It’s pretty good, and goes a lot into his background before Deepmind (chess
46.
▲
by
libraryofbabel
5mo ago
Their CEO is on record as saying this. You may think he's lying, but that's just your opinion; given the pricing and how it stacks relative to the pricing of inference providers of comparable open source models (who are certainly
47.
▲
by
libraryofbabel
5mo ago
This is already happening. For new Anthropic enterprise accounts you are billed at api token prices (maybe with a small volume discount). Anthropic makes a profit on those tokens. (Sure, that profit does not cover the model training costs,
48.
▲
by
libraryofbabel
5mo ago
These are flaws from 6-12 months ago. You might want to spend some time talking to Opus 4.7 or GPT 5.5. I can assure you that they can count letters just fine. You’re right that AI isn't perfect, but it’s pretty good. Especially since
49.
▲
by
libraryofbabel
5mo ago
I’m a Brit who lives in the US. Can confirm 99% of British people have never heard of the war of 1812, and even if they are a military history nerd and they have heard of it, they will consider it a minor sideshow to the main event of the e
50.
▲
by
libraryofbabel
5mo ago
> Storing on GPU would be the absolute dumbest thing they could do No. It’s not dumb. There will be multiple cache tiers in use, with the fastest and most expensive being on-GPU VRAM with cache-aware routing to specific GPUs and then pro
51.
▲
by
libraryofbabel
6mo ago
Are you sure about that? They charge $6.25 / MTok for 5m TTL cache writes and $10 / MTok for 1hr TTL writes for Opus. Unless you believe Anthropic is dramatically inflating the price of the 1hr TTL, that implies that there is some
52.
▲
by
libraryofbabel
6mo ago
To extend your point: it's not really the storage costs of the size of the cache that's the issue (server-side SSD storage of a few GB isn't expensive), it's the fact that all that data must be moved quickly onto a GPU i
53.
▲
by
libraryofbabel
6mo ago
That is because LLM KV caching is not like caches you are used to (see my other comments, but it's 10s of GB per request and involves internal LLM state that must live on or be moved onto a GPU and much of the cost is in moving all tha
54.
▲
by
libraryofbabel
6mo ago
They are caching internal LLM state, which is in the 10s of GB for each session. It's called a KV cache (because the internal state that is cached are the K and V matrices) and it is fundamental to how LLM inference works; it's no
55.
▲
by
libraryofbabel
6mo ago
> There are a lot of companies who would gladly drop half a million on a GPU to have private inference that Anthropic or OpenAI can’t use to steal their data. Obviously, and certainly companies do run their own models because they place
56.
▲
by
libraryofbabel
6mo ago
Nit: It doesn’t have to live in GPU memory. The system will use multiple levels of caching and will evict older cached data to CPU RAM or to disk if a request hasn’t recently come in that used that prefix. The problem is, the KV caches are
57.
▲
by
libraryofbabel
6mo ago
> you can download it, run it on your systems In theory, sure, but as other have pointed out you need to spend half a million on GPUs just to get enough VRAM to fit a single instance of the model. And you’d better make sure your use case
58.
▲
by
libraryofbabel
6mo ago
I wonder if the difficulties LLMs have with “seeing” complex detail in images is muddying the problem here. What if you hand it the cube state in text form? (You could try ascii art if you want a middle ground.) If you want to isolate the i
59.
▲
by
libraryofbabel
6mo ago
Stephenson’s piece is a classic, but it was written in 1996, when things were very different in the tech industry and geopolitically. Much more up to date (and with an explicit debt to Stephenson) is Samanth Subramanian, The Web Beneath Th
60.
▲
by
libraryofbabel
6mo ago
That's interesting research, but I think a more important reason that you don't have access to them (not even via the bare Anthropic api) is to prevent distillation of the model by competitors (using the output of Anthropic's
More ›