Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
awnihannun
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
2 ms
·
1.
▲
by
awnihannun
10mo ago
Right, my comment was mostly about decoding speed. For prefill you can get a speed up but there you are less latency bound. In our benchmarks with MLX / mlx-lm it's as much as 3.5x for token generation (decoding) at batch size 1 o
2.
▲
by
awnihannun
10mo ago
For a bit more context, those posts are using pipeline parallelism. For N machines put the first L/N layers on machine 1, next L/N layers on machine 2, etc. With pipeline parallelism you don't get a speedup over one machine -
3.
▲
WWDC25: Explore large language models on Apple Silicon with MLX [video]
(youtube.com)
7 points
by
awnihannun
1y ago
|
1 comments
4.
▲
by
awnihannun
1y ago
Everything you want to know about running LLMs with MLX on Apple silicon: - Introduction - MLX LM Introduction - Text generation - Quantization - Fine-tuning - LLMs in MLXSwift
5.
▲
by
awnihannun
9y ago
I agree with your point. It can be hard for a US native English speaker to recognize a Scottish accent. But, other Scottish people certainly don't have trouble with understanding a Scottish accent. So I view that as a certificate that
6.
▲
by
awnihannun
9y ago
There are a few categories that I think TensorFlow is notably strong in. Namely: 1. Deployment. 2. Coverage of the library / built-in functionality. 3. Device management. For more details, I wrote a comparison of PyTorch and TensorFlow