Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
crowwork
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
Modern GPU Programming for MLSys
(mlc.ai)
80 points
by
crowwork
3mo ago
|
25 comments
2.
▲
Building an Open ABI and FFI for ML Systems
(tvm.apache.org)
2 points
by
crowwork
11mo ago
|
0 comments
3.
▲
Open ABI and FFI for Machine Learning Systems
(github.com)
1 points
by
crowwork
11mo ago
|
1 comments
4.
▲
by
crowwork
11mo ago
The goal of the project is to bring open ABI and FFI for machine learning systems. - Stable, minimal C ABI designed for kernels, DSLs, and runtime extensibility. - Zero-copy interop across PyTorch, JAX, and CuPy using DLPack protocol. - Com
5.
▲
by
crowwork
2y ago
Scale LLM serving with programmable cross-engine serving patterns, all in a few lines of Python
6.
▲
by
crowwork
2y ago
XGrammar is an open-source library for efficient, flexible, and portable structured generation. Bring 2x-10x speedup in grammar grammar-guided(JSON and CFG) LLM serving.
7.
▲
by
crowwork
2y ago
Comes with ability to do full structured generation with json schema also a in-browser demo https://chat.webllm.ai/
8.
▲
MLCEngine: Universal LLM Deployment to Both Cloud and Local Devices
(blog.mlc.ai)
2 points
by
crowwork
2y ago
|
0 comments
9.
▲
by
crowwork
2y ago
runs on qwen2 on iphone with 26 tok/sec and a OpenAI style swift API
10.
▲
by
crowwork
3y ago
2b model running at 20tok/sec on iphone, nice potential for future applications
11.
▲
Running Google's Gemma 2B on Android
(old.reddit.com)
2 points
by
crowwork
3y ago
|
0 comments
12.
▲
by
crowwork
3y ago
Runs Phi-2 on Samsung S23 with pretty decent speed on Google Chrome browser. LLM on browser on a phone
13.
▲
by
crowwork
3y ago
You can also try out the vulkan backend, which we know should work for windows, although speed might be slower than rocm
14.
▲
by
crowwork
3y ago
Yes, it works out of box and the blog contains a prebuilt python package that you can try out
15.
▲
by
crowwork
3y ago
There is also vulkan support which should be more universal(also included in the post), for example, the post also shows running LLM on a steamdeck APU.
16.
▲
by
crowwork
3y ago
Checkout the latest docs https://mlc.ai/mlc-llm/docs/ MLC started with demos and it evolved lately, with API integrations, documentations into an inference solution that everyone can reuse for universal deployment
17.
▲
MLC Chat: Chat with Open Language Models Locally on iPad and iPhone
(apps.apple.com)
2 points
by
crowwork
3y ago
|
0 comments
18.
▲
WebLLM NPM Package
(npmjs.com)
1 points
by
crowwork
3y ago
|
0 comments
19.
▲
Bringing Hardware Accelerated Language Models to Android Devices
(github.com)
2 points
by
crowwork
3y ago
|
0 comments
20.
▲
Bringing Hardware Accelerated Language Models to Consumer Devices
(mlc.ai)
1 points
by
crowwork
3y ago
|
0 comments
21.
▲
by
crowwork
3y ago
It certainly also involves generating code(e.g. WebGPU, vulkan) that are more akin to traditionally compiler, and more like graph and memory optimization. So indeed more than packaging. Please checkout the course if you are interested
22.
▲
by
crowwork
3y ago
There is a conda app that can be installed on macos
23.
▲
by
crowwork
3y ago
You can try out the demo and benchmark yourself
24.
▲
by
crowwork
3y ago
https://mlc.ai/mlc-llm/
25.
▲
MLC LLM – Large Language Models on iPhone GPU and Many More GPU Platforms
(mlc.ai)
2 points
by
crowwork
3y ago
|
0 comments
26.
▲
MLC LLM: Universal LLM Deployment with GPU Acceleration
(github.com)
3 points
by
crowwork
3y ago
|
1 comments
27.
▲
by
crowwork
3y ago
Supported platforms include: - Metal GPUs on iPhone and Intel/ARM MacBooks - AMD and NVIDIA GPUs via Vulkan on Windows and Linux - NVIDIA GPUs via CUDA on Windows and Linux - WebGPU on browsers (through companion project WebLLM).
28.
▲
by
crowwork
3y ago
tvm runtime is pretty decent(~700k-2M level depending on dependency included), you can checkout tvm community and bring up the question there, i think there might be some common interest. There are impl of runtime for vulkan, metal that can
29.
▲
by
crowwork
3y ago
I think instead what would be needed is a wgpu native runtime support for TVM. Like the implementations in tvm vulkan, then it will be naturally link to any runtime that provides webgpu.h Then yah the llm_chat.js would be high-level logic
30.
▲
by
crowwork
3y ago
The WGSL are generated and compiled through TVM and embedded into the wasm. I think what you mean is wgpu native support. At the moment the web gpu runtime dispatches to the js webgpu environment. Once TVM runtime comes with wgpu native sup
More ›