Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andy12_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
andy12_
9mo ago
Is it that weird that AI agents (and arguably also humans) are faster and more efficient to use if standardized APIs/UIs are available?
92.
▲
by
andy12_
9mo ago
This is probably related to this [1] if anyone is wondering. https://news.ycombinator.com/item?id=46527950
93.
▲
by
andy12_
9mo ago
> then you have an extremely biased sample to gauge the overall mood of their populace. I think that if a good chunk of the people that don't agree with their government are basically forced to emigrate you don't get to turn ar
94.
▲
by
andy12_
9mo ago
Because the general idea here is that image and video models, when scaled way up, can generalize like text models did[1], and eventually be treated as "world models"[2]; models that can accurately model real world processes. These
95.
▲
by
andy12_
9mo ago
This is extremely similar to Karpathy's idea of a "cognitive core" [1]; an extremely small model with near-0 encyclopedic knowledge and basic reasoning and tool-use capabilities. [1] https://x.com/karpathy
96.
▲
by
andy12_
9mo ago
In which sense is it regulated? Are they regulated in any way that matters for this discussion? Have their ecological consequences been avoided by regulation? The oil and gas industries continue to be the biggest culprits of climate change,
97.
▲
by
andy12_
10mo ago
If you were capable of time travel and you could go to the past and convince world government of the evil oil and gas industries, and that their expansion should be prevented, would you have done it? Would you have prevented the technologic
98.
▲
by
andy12_
10mo ago
I don't care about the supposed ecological consequences of AI. If we need more water, we build more desalination plants. If we need more electricity, we build more nuclear reactors. This is purely a technological problem and not a mor
99.
▲
by
andy12_
10mo ago
You just have to extrapolate the improvements in consistency in image model from the last couple of years and apply it to these kinds of video models. When in a couple of years they can generate consistent videos of many physical phenomena
100.
▲
by
andy12_
10mo ago
I remember doing this kind of test in a vanilla transformer trained on my laptop on a small text dataset. I basically added N^3 attention where each layer could pay attention to previous layers. It didn't improve anything and was much
101.
▲
by
andy12_
10mo ago
It seems like this could be solved by partial structured output, where the structure of the JSON itself is constrained, but the values of the JSON entries is not (so even if "quantity" here is set to int, the model can output &quo
102.
▲
by
andy12_
10mo ago
I'm confused about the "Accuracy vs Cost" section. Why is Gemini 3 Pro so cheap? It's basically the cheapest model in the graph (sans Llama 4 and Mistral Large 3) by a wide margin, even compared to Gemini 3 Flash. Is tha
103.
▲
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
(arxiv.org)
2 points
by
andy12_
10mo ago
|
1 comments
104.
▲
by
andy12_
10mo ago
Transformers show remarkable versatility across domains, suggesting the existence of inductive biases beneficial across modalities. In this work, we explore a new way to instil such generic biases in vision transformers (ViTs) by pretrainin
105.
▲
by
andy12_
10mo ago
Kind-of. You could theoretically use LoRA for this, in fact, but it probably wouldn't have enough capacity to make it a proper substitute of the attention mechanism. Instead a full MLP is trained as input chunks get processed.
106.
▲
by
andy12_
10mo ago
This is an oversimplification of what Titans does. The model performs nested learned, where the model learns during inference, and during training the model weights learn _how and what_ to learn during inference. If the input contains junk
107.
▲
by
andy12_
10mo ago
That's weird, from my own tests Nano banana pro has no problem generating complex infographics with legible text.
108.
▲
by
andy12_
10mo ago
No, the "large _language_ model" name is a misnomer nowadays. Some time ago it was indeed common to get a pure-text model and inject embeddings from a separately trained image-encoder (which generated "meh" results), but
109.
▲
by
andy12_
10mo ago
Do note that that is a different model. The one we are talking about here, DeepSeekMath-V2, is indeed overcooked with math RL. It's so eager to solve math problems, that it even comes up with random ones if you prompt it with "Hel
110.
▲
by
andy12_
11mo ago
You might not believe this, but there are a lot of people (me included) that were extremely excited about the Gemini 3 release and are pleased to see the SOTA benchmark results, and this is reflected in the comments.
111.
▲
by
andy12_
11mo ago
For my part, I don't know why, but Zig's syntax feels wrong to me. I don't even know why. I really want to like its syntax, as Zig seems really promising to me, but I just don't, which makes it not very enjoyable for m
112.
▲
by
andy12_
11mo ago
Actually, no! Look at this in the paper > In extending from studying per-example to bulk memorization, we propose a novel inversion of the previous interpretation of loss curvature: while individual memorized points are associated with h
113.
▲
by
andy12_
11mo ago
Very concise summary of the procedure described in this paper: 1. Run the model once across a dataset to estimate loss curvature per MLP weight matrix via K-FAC (activation/gradient covariances). 2. Decompose each weight matrix into cu
114.
▲
From Memorization to Reasoning in the Spectrum of Loss Curvature
(arxiv.org)
65 points
by
andy12_
11mo ago
|
14 comments
115.
▲
by
andy12_
1y ago
I have never noticed any major difference in performance of ChatGPT between English and Spanish. The truth is that as long as the amount of training data of a given language is above some threshold, knowledge transfers between languages.
116.
▲
by
andy12_
1y ago
Does anyone with more knowledge on the subject know if this technology has any gotchas? It seems really promising, and when making some estimates with GPT5-thinking, it seems that it could be as cheap as conventional batteries in a kWh-basi
117.
▲
Concrete "battery" developed at MIT now packs 10 times the power
(news.mit.edu)
3 points
by
andy12_
1y ago
|
2 comments
118.
▲
by
andy12_
1y ago
You don't? Now I use Gemini to code and optimize CUDA kernels. When I first used GPT3 in the OpenAI playground I was extremely impressed when I managed to get it to output a hello world program in C.
119.
▲
by
andy12_
1y ago
When I say 10 times cheaper, I mean when comparing models of the same capabilities. The kind of performance you get now for a 200$ subscription, a year ago probably would have costed 2000$.
120.
▲
by
andy12_
1y ago
Models drop in price x10 each year. Us, common folk, getting access to these kinds of models is just a matter of time.
More ›