Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ruihangl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
PithTrain – a compact, agent-native MoE training system
(blog.mlc.ai)
4 points
by
ruihangl
4mo ago
|
0 comments
2.
▲
XGrammar: Efficient, Flexible and Portable Structured Generation for LLM
(github.com)
12 points
by
ruihangl
2y ago
|
1 comments
3.
▲
High-Throughput Low-Latency LLM Serving with MLCEngine
(blog.mlc.ai)
8 points
by
ruihangl
2y ago
|
0 comments
4.
▲
by
ruihangl
2y ago
It's good to see that it contains the standard OpenAI API interface in JavaScript which looks very convenient.
5.
▲
by
ruihangl
2y ago
A unified efficient open-source LLM deployment engine for both cloud server and local use cases. It comes with full OpenAI-compatible API that runs directly with Python, iOS, Android, browsers. Supporting deploying latest large language mod
6.
▲
Universal LLM Deployment Engine with ML Compilation
(blog.mlc.ai)
17 points
by
ruihangl
2y ago
|
7 comments
7.
▲
by
ruihangl
3y ago
Great work! I am curious that how much effort it would take to support LoRAs with different ranks?
8.
▲
Run Llama2-70B in Web Browser with WebGPU Acceleration
(webllm.mlc.ai)
9 points
by
ruihangl
3y ago
|
6 comments
9.
▲
by
ruihangl
3y ago
Purely running in web browser. Generating 6.2 tok/s on Apple M2 Ultra with 64GB of memory.
10.
▲
by
ruihangl
4y ago
Upgrading the model is pretty easy. We just need to build the new model locally in the same way we build the current model. This usually takes fewer than 2min. If people want to deploy the new version to web browser and share for others to
11.
▲
by
ruihangl
4y ago
Thanks for the pointer! As far as we know the WebGPU development on firefox is a bit lagging behind, so we use Chrome and did not develop this project on firefox.
12.
▲
by
ruihangl
4y ago
Yes of course. Optimizing and building the model to the format acceptable by ONNX web runtime will getting this in. On the other hand, we also need to enhance our own runtime (for example for better memory pool management) in the future.
13.
▲
by
ruihangl
4y ago
Thanks for your interest! Most of the existing stable diffusion demos rely on a server behind to run the image generation. It means you need to host your own GPU server to support these workloads. It is hard to have the demo run purely on w