Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dlewis1788
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
dlewis1788
1y ago
Looks like more than KV is having an issue. Just tried to load dash.cloudflare.com and no bueno.
2.
▲
by
dlewis1788
2y ago
100% - probably why vLLM is now the default back-end in Dynamo.
3.
▲
by
dlewis1788
2y ago
100% valid - Nvidia is trying to address that now with cuTile and the new Python front-end for CUTLASS.
4.
▲
by
dlewis1788
2y ago
CUDA is an entire ecosystem - not a single programming language extension (C++) or a single library, but a collection of libraries & tools for specific use cases and optimizations (cuDNN, CUTLASS, cuBLAS, NCCL, etc.). There is also tool
5.
▲
by
dlewis1788
2y ago
For training, yes, but no indications on inference workloads. Apple has said they would use their own silicon for inference in the cloud.
6.
▲
by
dlewis1788
3y ago
I didn't even know about Apple's AMX instructions until I clicked on your link. Very interesting - thanks!
7.
▲
by
dlewis1788
3y ago
My understanding is for certain types of networks BF16 will train better than FP16, given the additional protection against exploding gradients and loss functions with the extended range of BF16 - at the loss of precision.
8.
▲
by
dlewis1788
3y ago
Confirmed Apple M1 lacks bfloat16 support completely - M1: hw.optional.arm.FEAT_BF16: 0 vs M2: hw.optional.arm.FEAT_BF16: 1