Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
highfrequency
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
highfrequency
10mo ago
Per the author’s links, he warned that deep learning was hitting a wall in both 2018 and 2022. Now would be a reasonable time to look back and say “whoops, I was wrong about that.” Instead he seems to be doubling down.
92.
▲
The Great Filter
(mason.gmu.edu)
2 points
by
highfrequency
10mo ago
|
0 comments
93.
▲
ArXiv Monthly Submissions Chart
(arxiv.org)
1 points
by
highfrequency
10mo ago
|
0 comments
94.
▲
by
highfrequency
10mo ago
Great respect for Ilya, but I don’t see an explicit argument why scaling RL in tons of domains wouldn’t work.
95.
▲
Distributed Representations (Geoff Hinton) [pdf]
(cs.toronto.edu)
2 points
by
highfrequency
10mo ago
|
0 comments
96.
▲
The Smart Squeeze
(hypersoren.xyz)
1 points
by
highfrequency
11mo ago
|
0 comments
97.
▲
by
highfrequency
11mo ago
> But for API use, the models are easily substituted, so market share is fleeting. The LLM interface being unstructured plain text makes it simpler to upgrade to a smarter model than than it used to be to swap a library or upgrade to a n
98.
▲
by
highfrequency
11mo ago
Sure, every business owner has incentives that point to delivering a worse product (eg cheaper pizza ingredients increase margins). For most businesses there is a strong counteracting incentive to do a great job so the customer returns ne
99.
▲
by
highfrequency
11mo ago
The universal theme with general purpose technologies is 1) they start out lagging behind current practices in every context 2) they improve rapidly, but 3) they break through and surpass current practices in different contexts at different
100.
▲
by
highfrequency
11mo ago
Is GPT-5.1-Codex better or worse than GPT-5.1 (Thinking) for straight up mathematical reasoning (ie if it is optimized for making code edits)? Said another way: what is the set of tasks where you expect GPT 5.1 to be better suited than GPT-
101.
▲
by
highfrequency
11mo ago
> 1) Prioritize your ease of being over any other consideration: parties are like babies, if you’re stressed while holding them they’ll get stressed too. Every other decision is downstream of your serenity: e.g. it's better to have
102.
▲
by
highfrequency
11mo ago
If AI can solve all of your interview questions trivially, maybe you should figure out how to use AI to do the job itself.
103.
▲
Agent Labs Are Eating the Software World
(nibzard.com)
1 points
by
highfrequency
11mo ago
|
1 comments
104.
▲
by
highfrequency
11mo ago
> When your throughput increases by an order of magnitude, you're not just writing more code - you're making more decisions. > These aren't just implementation details - they're architectural choices that ripple th
105.
▲
Deep Learning 33 Years Ago (Karpathy) (2022)
(karpathy.github.io)
3 points
by
highfrequency
1y ago
|
0 comments
106.
▲
by
highfrequency
1y ago
Fortunately, we can have LLMs write code and keep all the benefits of normal software (determinism, reproducibility, permanent bug fixes etc.) I don’t think anyone is advocating for web apps to take the form of an LLM prompt with the app ge
107.
▲
by
highfrequency
1y ago
> posed a fascinating question: How can a relatively modest embedding space of 12,288 dimensions (GPT-3) accommodate millions of distinct real-world concepts? Because there is a large number of combinations of those 12k dimensions? You
108.
▲
Basic Guide to Einsum
(ajcr.net)
4 points
by
highfrequency
1y ago
|
1 comments
109.
▲
by
highfrequency
1y ago
> EVM-compatible, built on Reth Anyone know what this actually means? Both literally (what is Reth?) and what it means qualitatively: are Stripe’s crypto efforts competing with Ethereum or strengthening it?
110.
▲
by
highfrequency
1y ago
> We manage to dominate the world mainly by using brute force to simplify our environment and then maintaining and building systems on top of that simplified environment. If we didn't have the proper tools to selectively ablate our
111.
▲
by
highfrequency
1y ago
> Sometimes we dislike things simply because we have a concept of ourselves as not liking them. Astute!
112.
▲
by
highfrequency
1y ago
To crack NLP we needed a large dataset of labeled language examples. Prior to next-word prediction, the dominant benchmarks and datasets were things like translation of English to German sentences. These datasets were on the order of millio
113.
▲
by
highfrequency
1y ago
> without language, your thoughts are just emotions. Is that true though? Seems like you can easily have some cognitive process that visualizes things like cause and effect, simple algorithms or at least sequences of events.
114.
▲
by
highfrequency
1y ago
Do you have a link handy for where he says this explicitly?
115.
▲
by
highfrequency
1y ago
> The “establishment” mistakenly assumes that a shameless person wants the approval of their community, when it turns out that, much like any cult or counterculture, that person’s goal was to attract a following, regardless of who the me
116.
▲
by
highfrequency
1y ago
Enjoyed the article. To play devil’s advocate, an entirely different explanation for why huge models work: the primary insight was framing the problem as next-word prediction. This immediately creates an internet-scale dataset with trillion
117.
▲
Sol Price (Farnam Street)
(fs.blog)
1 points
by
highfrequency
1y ago
|
0 comments
118.
▲
by
highfrequency
1y ago
Interesting that for these small models, it is optimal for the embedding parameters to be a huge fraction of the total (170e6/250e6) = 68%!
119.
▲
by
highfrequency
1y ago
This is awesome - thanks for sharing. Appreciate the small-scale but comprehensive studies testing out different architectures, model sizes and datasets. Would be curious to see a version of your model size comparison chart but letting the
120.
▲
by
highfrequency
1y ago
> We can literally define an airplane parametrically in a configuration file and press a button. In a matter of minutes we have a complete quick-and-dirty analysis of how the whole aircraft performs—as mkBoom flies the aircraft through a
More ›