Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aesthesia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
151.
▲
by
aesthesia
5mo ago
That's also a CPU that came out four years later than the A100. The contemporaneous B200 is not optimized for FP32 and does 74.45 TFLOP/s. For FP16 it's at ~2 PFLOP/s.
152.
▲
by
aesthesia
5mo ago
See the later post testing a newer Mythos checkpoint, though: https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber...
153.
▲
by
aesthesia
5mo ago
Yeah, this is a big part of it. Labs have been hill climbing on Python for years, plus AI devs are usually most familiar with Python anyways.
154.
▲
by
aesthesia
5mo ago
The Nemotron model has attention layers interspersed with the Mamba layers, and I didn't see any attention layers in the model. It looks like the attention layers are present but show up as blocks with an RMSNorm followed by two sequen
155.
▲
by
aesthesia
5mo ago
SAEs are useful, and the Qwen release is great, but this is a different thing entirely.
156.
▲
by
aesthesia
5mo ago
Looking at your experiment code, it seems like the retrieval experiments are done with the reconstructed vectors of dimension D rather than the compressed vectors of dimension d, which doesn't have any direct performance improvements.
157.
▲
by
aesthesia
5mo ago
This is a neat idea. When I'm looking up models I usually want to see something about the architecture, but also some of the hyperparameters for the specific model---residual dimension, total number of layers, tokenizer configs. There&
158.
▲
by
aesthesia
5mo ago
Calling the AISLE experiment a "benchmark" is generous. They tested three code snippets on each model.
159.
▲
by
aesthesia
5mo ago
It maps (-inf, inf) to (0, inf) in about as nice a way as you could expect (addition turns into multiplication). When you want to constrain a value to be positive, parameterizing it with exp is usually a good option.
160.
▲
by
aesthesia
5mo ago
The "model card" concept actually comes from a pre-LLM Google paper ( https://arxiv.org/abs/1810.03993 ), where the example cards did fit on a single page. The concept quickly became a standard component of AI
161.
▲
Don't forget: The plural of anecdote is data
(blog.danwin.com)
1 points
by
aesthesia
5mo ago
|
0 comments
162.
▲
by
aesthesia
5mo ago
There's some validity to these criticisms, but it would be a lot more credible to cite someone whose job isn't "loudly promote any claim that sounds negative for AI, regardless of how well-founded it is."
163.
▲
by
aesthesia
5mo ago
It would be twice that, since nVidia always lists "with sparsity" FLOPS as the headline number. But I bet they got a bunch of research credits to do this.
164.
▲
by
aesthesia
5mo ago
There's a similar but unreleased project here: https://github.com/DGoettlich/history-llms I've been waiting for them to publish the 4B model for a while so I'm glad to have something similar to play with
165.
▲
by
aesthesia
5mo ago
But of course the monarch was a queen for the majority of the 19th century. While there's definitely post-1930 information that made it into the training data, I suspect the reason this happened is that the model is not very sure what
166.
▲
by
aesthesia
6mo ago
The same This American Life episode raised serious doubts about Dr. Steel's claims, which is mentioned in the article you link: > When reporters tried to corroborate Dr. Steel’s claims, however, holes started appearing, according to
167.
▲
by
aesthesia
6mo ago
I'm not totally convinced by this: > It might appear that this is an argument against scale, and the Bitter Lesson. That is not the case. I see this as a move that lets scale do its work on the right object. As with chess, where enc
168.
▲
by
aesthesia
6mo ago
I notice the experiments are all run with Gaussian token embeddings and weight matrices, which is a very different scenario than you would get in a real model. It shouldn't be much more difficult to try this with an actual model and da
169.
▲
by
aesthesia
6mo ago
A top-k approximation still requires k forward passes; that's k times as expensive as just computing the exact value. Unless you're doing a prefix-unconditional prediction, in which case you still likely need quite a large token -
170.
▲
by
aesthesia
6mo ago
They mention fine tuning an abliterated (post-trained) Qwen3.5 on Karoline Leavitt transcripts, but they don't mention doing this for the base models they test, and I suspect they didn't. For their use case (generating plausible t
171.
▲
by
aesthesia
6mo ago
This could be interesting work---it's definitely possible that pre-training corpus filtering has a hard-to-erase effect on post-trained model behavior. But it's hard to take this article seriously with the slop AI research report
172.
▲
by
aesthesia
6mo ago
What model is doing this prediction? The only way a transformer predicts the "next KV vector" is by sampling the next token and then running a forward pass with that token.
173.
▲
by
aesthesia
6mo ago
> The second layer, predictive delta coding, stores only the residual of each new KV vector from the model's own prediction of it I don't understand this. The key and value vectors for any given layer + token are created by the
174.
▲
by
aesthesia
6mo ago
Leaked/extracted system prompts for other chat models, particularly ChatGPT, are often around this size. Here's GPT-5.4: https://github.com/asgeirtj/system_prompts_leaks/blob/main/O...
175.
▲
by
aesthesia
6mo ago
I think this would come off a lot better if the recommendations weren't so absolute. I like the effect of a multicolored slab of highlights calling out every LLM cliche in a passage. Yes, the slop style is not just the sum of these ind
176.
▲
by
aesthesia
6mo ago
At least in some fields, advanced courses are the most likely to have lower cost textbooks. Real analysis textbooks are usually cheaper than calculus textbooks. It's the introductory courses that tend to have $200 behemoths attached to
177.
▲
by
aesthesia
6mo ago
The new tokenizer is interesting, but it definitely is possible to adapt a base model to a new tokenizer without too much additional training, especially if you're distilling from a model that uses the new tokenizer. (see, e.g., https
178.
▲
by
aesthesia
6mo ago
> Anthropic has also admitted that the bugs found by Mythos had not been found by using a prompt like "find the bugs", but by running many times Mythos on each file with increasingly more specific prompts, until the final run t
179.
▲
by
aesthesia
4y ago
This reminds me of the “no hello” proposal for workplace chat messages (e.g. https://nohello.net/ ). It’s much less of a big deal for personal communication, but I can understand wanting someone to just say what they want to
180.
▲
by
aesthesia
4y ago
Interestingly, the traditional algorithmic solution is inherently asymmetric: it gives better outcomes to one gender than the other.
More ›