Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andy12_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
121.
▲
Gauss, an Agent for Autoformalization
(math.inc)
6 points
by
andy12_
1y ago
|
0 comments
122.
▲
by
andy12_
1y ago
This is nice and all for tourists and people that live in the city center. But then if you don't live in the actual city and need to drive for half an hour to reach the city, spend half an hour searching for parking, then take a half a
123.
▲
by
andy12_
1y ago
> They just choose the most probabilistic next token That does not imply that a model should hallucinate. A trivial counterexample is a small LLM trained up to 100% accuracy to output x mod 100 for any input x in the range 0-1000000 and
124.
▲
by
andy12_
1y ago
The same can be said about any recurrent network. To predict the token n+1 you could recalculate the hidden state up to token n, or reuse the hidden state of token n from the previous forward pass. The only difference is the amount of memor
125.
▲
by
andy12_
1y ago
What? No. The intermediate hidden states are preserved from one token to another. A token that is 100k tokens into the future will be able to look into the information of the present token's hidden state through the attention mechanism
126.
▲
by
andy12_
1y ago
Is there a way of making wayland actually usable with Nvidia GPUs? I never manage to make it work, and it makes the whole system feel slow and sluggish compared to X11
127.
▲
by
andy12_
1y ago
What's up with that search bar for bangs? Every time I write a new letter it appears to reload the webpage, polluting the browser history, and it's extremely slow, frequently missing letters if I type fast enough. That is by far t
128.
▲
by
andy12_
1y ago
How is this possible? I mean, I thought that sometimes you had no choice but to separate computation into several kernels. But here they literally allow cuda threads to dinamically perform tasks assigned by scheduler threads? I only have a
129.
▲
by
andy12_
1y ago
It really is an obvious hint. Anyone can check that since the blackout there is never less than 4000 or 5000 MW of nuclear at all times (compared to 3000 MW of nuclear in the days before the blackout). Also combined cycle. On the day of the
130.
▲
by
andy12_
1y ago
Not to be that guy but... clearly Ad Hominem.
131.
▲
by
andy12_
1y ago
Honestly, I don't think that it would be too hard. With the grid function you can do a lot of things, specially because you can use grid.cell(rowspan:3, colswap:4) to make cells that span multiple rows or columns, use fractional sizes
132.
▲
by
andy12_
1y ago
Is there any reason why you can't use Typst for any of the stuff you mentioned? I can't see why you couldn't (except for interactive forms, which is already being worked on [1]. The pdf-writer low-level backend seems to have
133.
▲
by
andy12_
1y ago
I also I'm very interested. I had played around a lot with Differentiable Logic Networks a couple of months ago and how to make the learned wiring scale to bigger number of gates. I had a couple of ideas that seemed to worked in a smal
134.
▲
by
andy12_
1y ago
I also worked a long time ago in recreating the original Deep Differentiable Logic Network paper [1], so I have a couple of additions to make. > I wanted to see if I could learn the wires in addition to the gates. I still think it’s poss
135.
▲
Spurious Rewards: Rethinking Training Signals in RLVR
(rethink-rlvr.notion.site)
1 points
by
andy12_
1y ago
|
0 comments
136.
▲
by
andy12_
1y ago
I thought the same, but Per-Layer Embeddings as a name doesn't make sense in any context, and MatFormer does exactly what the blogpost says PLE does. I just think it's more probable that the blogpost was written by several authors
137.
▲
by
andy12_
1y ago
I think that it's a poorly named reference to this paper [1] that they mention in the blogpost. If I had to give it another more descriptive name, I would probably name it "Per-Layer Embedding Dimensionality" [1] https:/
138.
▲
by
andy12_
1y ago
I can't say that I don't love Gemini. I use it a lot, and the huge context window does help. But I can also say that I much prefer how Claude writes code.
139.
▲
by
andy12_
1y ago
I haven't tried many variations yet because a basic prompt seems to work well, thought it is important to remind Gemini of not using ";" inside the text of the cards if you use it as a separator. I imagine that with better pr
140.
▲
by
andy12_
1y ago
You are so right. I just began doing this today for my exams, and it really feels like cheating with how easy it is. That, and also creating a podcast with NotebookLM with all the pdfs as source and using as a prompt "Make this a long
141.
▲
by
andy12_
1y ago
I have just tried this afternoon to create with Gemini 2.5 Pro Anki cards to study for my exams. I've been doing it raw: I just paste the whole material (like 100k worth of tokens) into aistudio and generate the flashcards in txt forma
142.
▲
by
andy12_
1y ago
Online Word (or Microsoft 365, or whatever it is called) regularly took me 2 minutes to load a 120 page document. I'm being very literal here. You could see it load in real time approximately 1 page a second. And it wasn't a netwo
143.
▲
VR-CLI: Learning to Reason for Long-Form Story Generation
(arxiv.org)
2 points
by
andy12_
1y ago
|
0 comments
144.
▲
by
andy12_
1y ago
Interestingly, when compering benchmarks of Experimental 03-25 [1] and Experimental 05-06 [2] it seems the new version scores slightly lower in everything except on LiveCodeBench. [1] https://storage.googleapis.com/model-car
145.
▲
by
andy12_
2y ago
I mean, it could very well be that it generates image patches autoregressively, but in a pyramidal way (first a very low resolution version, the "canvas", and then each individual patch). This is very similar to VAR [1] We can
146.
▲
by
andy12_
2y ago
The main progress value is that Test-Time Training appears to work very well in practice. I think that as labs begin to test it as scale in LLMs, it will become commonplace in next-generation models.
147.
▲
by
andy12_
2y ago
> defining two different return types and using the wrong one could happen in any language This specifically is the kind of bug that is avoided with strong typing. The compiler screams at you when using the wrong return type. For example
148.
▲
by
andy12_
2y ago
It seems they kind of do this already. [1] "We send out payments to different partners each month to plant and protect trees in biodiversity hotspots across the globe." [1] https://blog.ecosia.org/ecosia-financial-
149.
▲
by
andy12_
2y ago
Because ad revenue can be used for the non-profit? Ecosia's whole thing is that when you use it, you indirectly help plant trees - "the search engine that plants trees".
150.
▲
by
andy12_
2y ago
Actually, the 5 million figure is for the compute cost for the base 600B parameter model. Training R1 was just 8000 steps of reinforcement learning, so I expect that the vast, vast majority of the training cost is already included in the pr
More ›