Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cgdl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Erdős Problem #1026
(terrytao.wordpress.com)
6 points
by
cgdl
10mo ago
|
0 comments
2.
▲
Feynman vs. Computer
(entropicthoughts.com)
90 points
by
cgdl
10mo ago
|
25 comments
3.
▲
by
cgdl
1y ago
Very cool. For the INT4 QAT model, what is the recommended precision for the activations and for the key and values stored in KV cache?
4.
▲
by
cgdl
1y ago
Which model does the demo use?
5.
▲
by
cgdl
1y ago
I'd say llm inference requires both memory capacity and bandwidth. Cerebras provides bandwidth with on-chip SRAM, but not capacity (an entire wafer has only 44GB SRAM).
6.
▲
by
cgdl
1y ago
Indeed, and even if the cost per wafer was 300K, since about say 20-50 wafers are needed, its still 6MM to 15MM for the system. So likely it would appear this is VC subsidized.
7.
▲
by
cgdl
1y ago
Do you distinguish betwen "chips" and the wafer-scale system? Is the wafer-scale system significantly less than 3MM? EDIT: online it seems TSMC prices are about 25K-30K per wafer. So even 10Xing that a wafer-scale system should be
8.
▲
by
cgdl
1y ago
Exactly what I was thinking. What sort of latency do you think one would get with 8x B200 Blackwell chips? Do you think 1500 tokens/sec would be achievable in that setup?
9.
▲
by
cgdl
1y ago
Do we know how far this event was from earth? Wouldn't that distance be the determiner of what the relative contraction observed on earth would be?
10.
▲
by
cgdl
2y ago
I recently came across a critique of the Turing test that seems relevant here. Given the test's limited duration (five minutes in this study) and the constrained rate of human communication, it’s theoretically possible to anticipate ev
11.
▲
by
cgdl
2y ago
Thank you for the great discussion. You've put your finger on the right thing I think. We can now dispense with the old VC-type thinking (i.e., that it's because the hypothesis space is not complex enough that we get generalizatio
12.
▲
by
cgdl
2y ago
Thank you, this makes sense. I am thinking of this as an abstraction/refinement process where an abstract notion of the longer completion is refined into a cogent whole that satisfies the notion of a good completion. I look forward to
13.
▲
by
cgdl
2y ago
Thank you. In my mind, "planning" doesn’t necessarily imply higher-order reasoning but rather some form of search, ideally with backtracking. Of course, architecturally, we know that can’t happen during inference. Your example of
14.
▲
by
cgdl
2y ago
Yes, and that's the problem. What Zhang et al [2] showed convincingly in the Rethinking paper is that just focusing on the hypothesis space cannot be enough since the same hypothesis space fits real and random data so it's already
15.
▲
by
cgdl
2y ago
Agreed, but PAC-Bayes or other descendants of VC theory is probably not the best explanation. The notion of algorithmic stability provides a (much) more compelling explanation. See [1] (particularly Sections 11 and 12) [1] https:/
16.
▲
by
cgdl
2y ago
Very interesting. A related paper from a couple of years ago proposed a similar idea to understand generalization in deep learning: https://arxiv.org/abs/2203.10036