Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
junrushao1994
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
junrushao1994
3y ago
TVM Unity has a CUDA backend, and TensorCore MMA instructions are supported, so it wouldn't be hard to turn this option on. It's on our plan, but we haven't looked to enable them by default in the first place, mainly because
32.
▲
by
junrushao1994
3y ago
Thanks for sharing! It's definitely a bit painstaking to get a real-world LLM running at all on an iPhone due to memory constraint. It's also quite compute-intense as it has 7B parameters, but we are glad that it's generating
33.
▲
by
junrushao1994
3y ago
upcoming
34.
▲
by
junrushao1994
3y ago
You no longer need a powerful latest-gen GPU to run SOTA models, plus going through complicated setups. MLC-LLM makes it possible to use GPUs from any vendors, including AMD/Apple/NV/Intel, to run LLMs at reasonable speed, at
35.
▲
MLC-LLM: GPT/Llama on consumer-class GPUs and phones
(github.com)
303 points
by
junrushao1994
3y ago
|
106 comments
36.
▲
by
junrushao1994
3y ago
Nice work! This is interesting to read the comparison between Hidet and Triton in this blog: > Hidet Script vs. Triton: Triton greatly simplifies the CUDA programming by introducing the tile-based programming model where the parallel exe
37.
▲
by
junrushao1994
3y ago
Is there any plan to support larger models than GPT-2?
38.
▲
by
junrushao1994
3y ago
To your response, the model says: > Dear [Name], > Thank you for your message. We understand that the model you are referring to is a simple and basic model. However, it is important to highlight that this model serves a specific purp
39.
▲
by
junrushao1994
3y ago
Would be nice if anyone could help us benchmark! Our primary focus though is not model performance, but to demonstrate the capability that TVM Unity generates code targeting WebGPU and allows them to run with client GPUs :-)
40.
▲
by
junrushao1994
3y ago
In LLM world, loss or perplexity may not be the best indicator of model performance :-( Perhaps HELM ( https://crfm.stanford.edu/helm/latest/ ) but we didn't take deeper look as we are not the developers of thi
41.
▲
by
junrushao1994
3y ago
To share some fun stuff, here is the response generated by this model: As an AI language model, I would respond by acknowledging that the model discussed in the message is indeed smaller than some of the larger language models like GPT-3&#x
42.
▲
by
junrushao1994
3y ago
This is unfortunately non-trivial to quantitatively evaluate the performance against ChatGPT :-( We didn't do much evaluation because there isn't much innovation on model side, but instead we are demoing the possibility of running
43.
▲
by
junrushao1994
3y ago
To clarify, running this WebLLM demo doesn't need a 3.5k MacBook Pro which costs $3.5k :-) WebGPU supports multiple backends, besides Metal on Apple Silicon, it offloads to Vulkan, DirectX, etc. It means a windows laptop with Vulkan su
44.
▲
by
junrushao1994
4y ago
Ah nice! I intentionally didn’t talk a lot about “scheduling” because 1) I’m personally heavily working on it, which potentially makes a conflict of interest, and 2) I don’t want to deviate a lot from the topic in this thread about “optimiz
45.
▲
by
junrushao1994
4y ago
I’m using my real name, so its not hard to know who I am. Unfortunately, I don’t know much about you, and actually I don’t really think there is conflict of interest if you work in Modular, because Modular is also developing compiler abstra
46.
▲
by
junrushao1994
4y ago
Yes. I am the first author of the latest generation auto-tuner in an open source deep learning compiler :-) My comment is based on my personal experience: I did lead a 2nd/3rd grade undergrad to add software pipelining support and it w
47.
▲
by
junrushao1994
4y ago
I am personally a really huge fan of cutlass, and I've almost read every single file in their `include/cutlass/` folder (haven't followed up with the `cute` stuff yet). Just like you said, really appreciate that we could
48.
▲
by
junrushao1994
4y ago
Hey I am the first author of one of the "abstractions", so I guess my words would more or less reflect my personal daily experience dealing with those lovely kernels. Well, I don't have 100 engineers working for me, unfortuna
49.
▲
by
junrushao1994
4y ago
Halide is cool and has inspired many brilliant works like Tiramisu and Exocompilation. Love the idea a lot :-) We recently have some follow-ups on this idea as well. Happy to discuss about it and I don't want to distract from this thre
50.
▲
by
junrushao1994
4y ago
Yeah cuBLAS is definitely not perfect in many cases :-(( Speaking of GEMM fusion that you mentioned, flash attention is basically GEMM fusion with online softmax right? This is something I believe really cool and can be made really easy wit
51.
▲
by
junrushao1994
4y ago
> we used tensor cores and managed to get back fp32 accuracy with 3 rounds of the things Hey are you referring to 3xTF32 ( https://github.com/NVIDIA/cutlass/tree/master/examples/28_am... )? IMO thi
52.
▲
by
junrushao1994
4y ago
This is definitely a great point! With the context of AI workloads, where critical matmuls are basically of regular large shapes, are there many cases where cutlass/Triton are worse than cuBLAS where we need to throw more GPUs at it?
53.
▲
by
junrushao1994
4y ago
One thing I really love about XLA is GSPMD which effectively allows scalable distributed training in practice. However, I was quite curious how it is related to matrix multiplication though, given XLA is more focusing on graph-level optimiz
54.
▲
by
junrushao1994
4y ago
My take: optimizing matrix multiplication is not hard on modern architecture if you have the right abstraction. The code itself could be fragmented across different programming models, which is true, but the underlying techniques are not ha
55.
▲
by
junrushao1994
4y ago
Submitted a question in Firefox support forum, and there seems some bugs blocking it from running smoothly: https://support.mozilla.org/en-US/questions/1408328
56.
▲
by
junrushao1994
4y ago
Yep it surprisingly works on my AMDGPU too, even if it’s designed only for M1/M2
57.
▲
by
junrushao1994
4y ago
I downloaded Firefox nightly a couple of days ago, turned the flags on accordingly, but it didn't work out saying: > Find an error initializing the WebGPU device TypeError: adapter.requestAdapterInfo is not a function No idea how to
58.
▲
by
junrushao1994
4y ago
Curious if this library can be integrated with WebGPU - there is a recent post on ( https://news.ycombinator.com/item?id=35191687 ) announced that WebGPU can now be used for large models
59.
▲
by
junrushao1994
4y ago
Is it possible to integrate this with [onnxruntime-web]( https://onnxruntime.ai/docs/tutorials/web/ )?
60.
▲
by
junrushao1994
4y ago
This actually works on my non-M1/M2 macbook with AMD GPUs pretty smoothly. It's a bit slow for now (~70sec), but I expect this to be faster with fp16. Any roadmap to support fp16 on WebGPUs?
More ›