Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zhihaojia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
zhihaojia
1y ago
1. In MPK, each task is mapped to an individual SM. The amount of work handled by a task is similar to that of a thread block in the traditional kernel-per-operator approach. 2. TL;DR: MPK automatically analyzes inter-task dependencies by t
2.
▲
by
zhihaojia
1y ago
Thanks for reproducing our results!
3.
▲
by
zhihaojia
1y ago
You are right that CUDA graph can help reduce launch overhead but does not support overlapping computation/communication across layers, since data dependencies are described at the kernel level.
4.
▲
by
zhihaojia
1y ago
Ooops, missed one sentence in my previous response. Stanford's MegaKernel project tackles a similar challenge but focuses on manual CUDA implementation. While MPK takes a compiler-driven approach—users express their LLMs at the PyTorch
5.
▲
by
zhihaojia
1y ago
Thanks for the great feedback! Stanford's MegaKernel project tackles a similar challenge but focuses on manual CUDA implementation. While MPK takes a compiler-driven approach—users express their LLMs at the PyTorch level, and MPK autom
6.
▲
by
zhihaojia
1y ago
Yes, it would be a lot of fun if MPK can enable torch.compile to generate megakernels. Torch-generated kernels are currently too slow for latency-sensitive workloads.
7.
▲
by
zhihaojia
1y ago
Thanks for the feedback! Yes, we believe the approach is general and applicable to other ML workloads.
8.
▲
by
zhihaojia
1y ago
The task implementations used by MPK are currently optimized for A100. While the Mirage compiler can generate task implementations for other architectures such as Hopper and Blackwell, but we haven't integrated things together yet. Thi
9.
▲
by
zhihaojia
1y ago
JAX's operator fusion ( https://apxml.com/courses/advanced-jax/chapter-2-optimizing-... ) can fuse a few local operators (e.g., matmul and elementwise computation) into a single kernel. But JAX's approach
10.
▲
by
zhihaojia
1y ago
The github repo includes a tutorial for using MPK: https://github.com/mirage-project/mirage/tree/mpk
11.
▲
by
zhihaojia
1y ago
Thanks a lot for your positive feedback! We believe that MPK can enhance existing LLM serving systems, especially for low-latency LLM serving. We are very excited about the opportunity to collaborate with others on direction.
12.
▲
by
zhihaojia
1y ago
Thanks for reading the post and github README. Supporting training is definitely feasible but the benefit may not be as significant as low-latency inference since training generally involves much larger kernels, making kernel launch overhea
13.
▲
by
zhihaojia
1y ago
This is the writer of the blog post. You are right that Stanford's work is a parallel effort. The main difference is that our focus is on compilation: making it easier to generate megakernels automatically.
14.
▲
Accelerating LLM Serving with Speculative Inference and Token Tree Verification
(github.com)
3 points
by
zhihaojia
3y ago
|
1 comments
15.
▲
by
zhihaojia
3y ago
SpecInfer is a system that accelerates generative LLM serving with speculative inference and token tree verification. The key idea is to use an LLM as a token tree verifier instead of an incremental decoder. We show that this reduces LLM in