Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
neilmovva
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
26 ms
·
1.
▲
by
neilmovva
20d ago
hi, I'm one of the founders of Sail. I'm very sorry that you had a bad experience with us! We are serious about serving models correctly, and always publish a link to the exact HF checkpoint we're using for each model in our
2.
▲
by
neilmovva
10mo ago
I think Hopper's native matmul tile is 64x64, and Blackwell is 128x128. see this blog for a reference on Blackwell: https://hazyresearch.stanford.edu/blog/2025-03-15-tk-blackwe...
3.
▲
by
neilmovva
1y ago
The multilingual example in the launch graphic has Qwen3 producing the text: > "Bonjour, pourriez-vous me dire comment se rendreà la place Tian'anmen?" translation: "Hello, could you tell me how to get to Tiananmen Sq
4.
▲
by
neilmovva
1y ago
Today, training in "low precision" probably means computing FP8 x FP8 -> FP32. The FP32 accumulation is still important, but otherwise yes this works, especially if we're talking about MXFP8 as supported on Blackwell [0].
5.
▲
by
neilmovva
1y ago
Not really: 5090: 210 TF / $2k == 105 TF/$k B200: 2250 TF / $40k == 56 TF/$k Getting only 2x the FLOPs per dollar probably isn't worth the hassle of having to rack 10x as many GPUs, while having no NVLink.
6.
▲
by
neilmovva
1y ago
I was surprised to see 5090's theoretical BF16 TFLOPs at just 209.5. That's not even 10% of the server Blackwell (B200 is 2250, and GB200 is 2500). B200 costs around $30-40k per GPU, so they are pretty close in performance per dol
7.
▲
by
neilmovva
2y ago
A bit surprised that they're using HBM2e, which is what Nvidia A100 (80GB) used back in 2020. But Intel is using 8 stacks here, so Gaudi 3 achieves comparable total bandwidth (3.7TB/s) to H100 (3.4TB/s) which uses 5 stacks of
8.
▲
by
neilmovva
3y ago
I agree that synchronization causes overhead, so 2x GPUs won't achieve the ideal 0.5x total runtime. But here, taking your Alpaca benchmark as an example, we are seeing 2x GPUs get 3.6x runtime with Huggingface, or 1.15x with Unsloth M
9.
▲
by
neilmovva
3y ago
promising results, excited to try it out! question on the perf benchmarks: why do all the results with 2 GPUs & DDP take longer than the single GPU case? Both benchmarks do the same amount of work, one training epoch, so this negative
10.
▲
by
neilmovva
3y ago
I can't universally agree with the headline statement. The article focuses on the pros of SRAM, which are real -- peak bandwidth (e.g. 5 TB/s out of the H100’s L2) and lower energy per bit transferred (the rule of thumb I remember
11.
▲
by
neilmovva
3y ago
I like this review: https://www.lighterra.com/papers/modernmicroprocessors/ A bit dated, but the major ideas used in current CPUs are all covered!
12.
▲
by
neilmovva
4y ago
> moves to Austin because it is less “vulnerable to climate change” > commutes by plane hmmm
13.
▲
by
neilmovva
4y ago
A bit underwhelming - H100 was announced at GTC 2022, and represented a huge stride over A100. But a year later, H100 is still not generally available at any public cloud I can find, and I haven't yet seen ML researchers reporting any
14.
▲
Nvidia Announces H100 NVL – Max Memory Server Card for Large Language Models
(anandtech.com)
122 points
by
neilmovva
4y ago
|
107 comments
15.
▲
by
neilmovva
4y ago
Yes, we all place a lot of trust in cloud vendors today. FHE is a way to move the trust boundary back to the client - let the server be as malicious or insecure as it wants. Raw compute could even become much cheaper, since any machine anyw
16.
▲
by
neilmovva
4y ago
Thanks for the feedback, I understand your hesitation. We don't just want to advertise guarantees - we want you to never trust third-party servers again. Fully homomorphic encryption makes this possible by never letting sensitive data
17.
▲
by
neilmovva
4y ago
Thanks for checking it out! Responses inline: > That sounds like loading the entire database every time Yup, we do perform computation over the entire database for every read - there is zero correlation between the server's work and
18.
▲
by
neilmovva
4y ago
Thanks! Yup, it's not always practical to make a huge number of queries when you expect many of them to come back empty. Instead, we first perform private lookups against a Bloom filter, to find out which keys actually hold data (e.g.
19.
▲
by
neilmovva
4y ago
Thanks! Yup, private retrieval is interesting as a product because it's a fundamentally new capability; there aren't really competitors we can show incremental improvements against. If you're still interested in the space, we
20.
▲
by
neilmovva
4y ago
Our FHE scheme uses lots of Number Theoretic Transforms (NTTs), which are pretty computationally expensive. NTT is a good candidate for acceleration, and there is quite a bit of interest from the zk community in doing so ( https://
21.
▲
by
neilmovva
4y ago
While OPT-175B is great to have publicly available, it needs a lot more training to achieve good results. Meta trained OPT on 180B tokens, compared to 300B that GPT-3 saw. And the Chinchilla scaling laws suggest that almost 4T tokens would
22.
▲
Show HN: Send private valentines, using homomorphic encryption
(valentine.blyss.dev)
4 points
by
neilmovva
4y ago
|
0 comments
23.
▲
by
neilmovva
4y ago
I was curious about exactly how much burst heat could be absorbed, so I asked WolframAlpha [0]. In a 15" workstation laptop, I think the CPU could quite reasonably pull +100 watts over steady-state TDP for 30 seconds. (100-gram aluminu
24.
▲
by
neilmovva
6y ago
Latest in a trend of silicon industry consolidation. A few other major moves in the embedded market over the last five years: NXP + Freescale in 2015 Microchip + Atmel in 2016 ON Semi + Fairchild in 2016 Infineon + Cypress in 2020
25.
▲
by
neilmovva
7y ago
The author's comments on cache sizes are a bit reductive. Not all "L3" is created equal, and designers always make tradeoffs between capacity and latency. In particular, the EPYC processors achieve such high cache capacities
26.
▲
by
neilmovva
7y ago
Actually, the L3 cache is also sharded across chiplets, so there's a small (~8MB) local portion of L3 that is fast, while remote slices will have to go over AMD's interdie connection fabric and incur a serious latency penalty. On
27.
▲
by
neilmovva
8y ago
That still means multiple silicon dies, which we have known how to do for a while (see: Intel Core 2 Quad from 2006, and more recently AMD Epyc). Having more dies lets you dissipate more heat, but then it's kinda hard to build low-late
28.
▲
by
neilmovva
9y ago
It matters when defining parallel work distribution. Unless memory bandwidth is homogeneous across the whole board (i.e. each TPU on a board gets 600 GB/s to its peers), we can't do model parallelism across ASICs efficiently, and
29.
▲
by
neilmovva
9y ago
Intel's "10nm" has roughly twice the transistor density vs. Samsung/TSMC "10nm", so I wouldn't compare based on the advertised process names.
30.
▲
by
neilmovva
9y ago
The fact that "2TB is addressable" is irrelevant. Putting NAND on the board doesn't improve latency/bandwidth nearly enough to function like vram. Nvidia has also supported unified virtual memory since Pascal, meaning yo
More ›