Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fooblaster
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
fooblaster
8mo ago
CUDA was profitable very early because of oil and gas code, like reverse time migration and the like. There was no act of incredible foresight from jensen. In fact, I recall him threatening to kill the program if large projects that made it
62.
▲
by
fooblaster
8mo ago
It was definitely luck, greg. And Nvidia didn't invent deep learning, deep learning found nvidias investment in CUDA.
63.
▲
by
fooblaster
8mo ago
I was really happy to see that blue spec was fully open sourced in recent years. Does anyone have experience with a non trivial project with it? Does it have any traction anymore in real silicon development.
64.
▲
by
fooblaster
9mo ago
calling neural engine the best is pretty silly. the best perhaps of what is uniformly a failed class of ip blocks - mobile inference NPU hardware. edge inference on apple is dominated by cpus and metal, which don't use their NPU.
65.
▲
by
fooblaster
9mo ago
Looks like they have made some progress on a native model in recent months: https://github.com/amd/IRON/tree/devel
66.
▲
by
fooblaster
9mo ago
sms is the Nvidia definition of processor, and cuda device properties returns it, not anything else. If you want a marketing number, use cuda cores, it doesn't consistently match to anything in the hardware design.
67.
▲
by
fooblaster
9mo ago
b200 is 148 sms, so no
68.
▲
by
fooblaster
9mo ago
The versal stuff isn't really an FPGA anymore. The chips have PL on them, but many don't. The consumer NPUs from AMD are the same versal aie cores with no PL. They just aren't configurable blocks in fabric anymore and don
69.
▲
by
fooblaster
9mo ago
10 MB mail app. 690 MB local llm to write snarky emails for you.
70.
▲
by
fooblaster
9mo ago
wow, say more..
71.
▲
by
fooblaster
9mo ago
I'd like to know more. I expect these systems are 8xvh1782. Is that true? What's the theoretical math throughput - my expectation is that it isn't very high per chip. How is performance in the prefill stage when inference is
72.
▲
by
fooblaster
9mo ago
This architecture is likely going to be a dead end for AMD. It has been in the wild for several years, yet still has no open programming model, multiple compiler stacks with poor software support. I find it likely that AMD drops this archit
73.
▲
by
fooblaster
9mo ago
The amd npu and versal ML tiles (same underlying architecture) have been an complete failure. Dynamic programming models like cu tile do not work on them at all, be cause they require an entirely static graph to function. AMD is going to wa
74.
▲
by
fooblaster
9mo ago
I wouldn't trust any benchmarks on the vendors site. Microsoft went down this path for years with FPGAs and wrote off the entire effort.
75.
▲
by
fooblaster
9mo ago
Yep, but they are still 50x faster than any fpga.
76.
▲
by
fooblaster
9mo ago
Show me a single FPGA that can outperform a B200 at matrix multiplication (or even come close) at any usable precision. B200 can do 10 peta ops at fp8, theoretically. I do agree memory bandwidth is also a problem for most FPGA setups, but x
77.
▲
by
fooblaster
9mo ago
Yeah, I wouldn't have guessed it would be helping me write systemverilog.
78.
▲
by
fooblaster
9mo ago
FPGAs will never rival gpus or TPUs for inference. The main reason is that GPUs aren't really gpus anymore. 50% of the die area or more is for fixed function matrix multiplication units and associated dedicated storage. This just isn&#
79.
▲
by
fooblaster
9mo ago
Great! How do you program it?
80.
▲
by
fooblaster
9mo ago
Well, anthropic just purchased a million TPUs from Google because even with a healthy margin from Google, it's far more cost effective because of Nvidia's insane markup. That speaks for itself. Nvidia will not drop their margin be
81.
▲
by
fooblaster
9mo ago
If you think they are going to catch up with Google's software and hardware ecosystem on their first chip, you may be underestimating how hard this is. Google is on TPU v7. meta has already tried with MTIA v1 and v2. those haven't
82.
▲
by
fooblaster
9mo ago
And yes, all their competitors are making custom chips. Google is on TPU v7. absolutely nobody is going to get this right on the first try among their competitors - Google didn't.
83.
▲
by
fooblaster
9mo ago
There is a pretty big moat for Google: extreme amounts of video data on their existing services and absolutely no dependence on Nvidia and it's 90% margin.
84.
▲
by
fooblaster
9mo ago
I find this all very cool, but why is this useful for game engine scripting. anyone know?
85.
▲
by
fooblaster
10mo ago
Anyway, perhaps we can chat in the executorch discord.
86.
▲
by
fooblaster
10mo ago
And you wouldn't happen to know about a torchscript replacement that is currently in-flight that is not based on export?
87.
▲
by
fooblaster
10mo ago
Its clear from listening to podcasts/interviews, he does not want to say anything to get on elons bad side. Interviewers appear to also not be eager to broach the subject.
88.
▲
by
fooblaster
10mo ago
So what are your users doing to get around this? Hoisting all control flow out?
89.
▲
by
fooblaster
10mo ago
That is true, but that doesn't mean Nvidia is not engaging in engineering to intentionally kneecap competition. Triton and other languages like that are a huge threat and CUtile is a means to combat that threat and prevent a hardware a
90.
▲
by
fooblaster
10mo ago
Let's see if developers sleepwalk into another trap to keep us locked into nvidia's hardware for the next decade.
More ›