Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chaoyu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Benchmarking LLM Inference Back Ends: VLLM, LMDeploy, MLC-LLM, TensorRT-LLM, TGI
(bentoml.com)
15 points
by
chaoyu
2y ago
|
1 comments
2.
▲
by
chaoyu
2y ago
onnx is not a good option for LLM type of autoregressive generation
3.
▲
by
chaoyu
3y ago
In my opinion, Automatic1111 is more of a development tool to experiment with different pipelines on your local GPU, or an internal tool serving a few users. OneDiffusion project aims to solve a very different problem, which is bringing a S
4.
▲
by
chaoyu
3y ago
Check out BentoML, which is the underlying serving framework used by OpenLLM, and it supports other type of models and modality such as images and videos.
5.
▲
by
chaoyu
3y ago
OpenLLM itself is under Apache 2 license, which does NOT restrict commercial use. However, OpenLLM as a framework can be extended to support other LLMs which may come with additional restrictions.
6.
▲
by
chaoyu
3y ago
OpenLLM plan to provide an OpenAI-compatible API, which allows you to even use OpenAI's python client to talk to OpenLLM, user just need to change to Base URL to point to your OpenLLM server. This feature is working-in-progress.
7.
▲
by
chaoyu
3y ago
Fine-tuning is coming up in the next release! You can actually try it out on the main branch :P
8.
▲
by
chaoyu
3y ago
Looking forward to it! OpenLLM is adding a OpenAI-compatible API layer, which will make it even easier to migrate LLM apps built around OpenAI's API spec. Feel free to join our Discord community and discuss more!
9.
▲
by
chaoyu
3y ago
Smaller models are likely more efficient to run inference and doesn't necessarily need the latest GPU. Larger language model trend to have better performance over more different type of tasks. But for a specific enterprise use case, ei
10.
▲
by
chaoyu
3y ago
The OpenLLM team is actively exploring those techniques for streamlining the fine-tuning process and making it accessible!
11.
▲
by
chaoyu
3y ago
OpenLLM in comparison focuses more on building LLM apps for production. For example, the integration with LangChain + BentoML makes it easy to run multiple LLMs in parallel across multiple GPUs/Nodes, or chain LLMs with other type of A
12.
▲
Show HN: ML Serving orchestration framework on Kubernetes
(github.com)
2 points
by
chaoyu
4y ago
|
0 comments
13.
▲
by
chaoyu
6y ago
BentoML.ai | ML Engineer, Backend Engineer | Full-time | Bay Area or Remote | Python, Kubernetes, MLOps platform, Data Infra, Tensorflow, PyTorch, etc BentoML is an open-source framework for machine learning model serving & deployment
14.
▲
by
chaoyu
6y ago
What does BentoML do? * Package models trained with any ML framework and reproduce them for model serving in production * Package once and deploy anywhere for real-time API serving or offline batch serving * High-Performance API model serve
15.
▲
BentoML: The easiest way to build Machine Learning APIs
(github.com)
4 points
by
chaoyu
6y ago
|
1 comments
16.
▲
by
chaoyu
6y ago
BentoML.ai | Open Source Evangelist / Technical Writer | San Francisco or Remote | Full time or Contract | http://docs.bentoml.org/ BentoML is an open-source platform for high-performance machine learning model serving
17.
▲
by
chaoyu
6y ago
I'm actually building a "modular open-source company/product" in the MLOps space: BentoML https://docs.bentoml.org/en/latest/
18.
▲
by
chaoyu
7y ago
BentoML( https://github.com/bentoml/BentoML ) may help you with the process of building endpoints with both Deep learning models and logistic regression/tree models, and it automatically helps you to containerize th
19.
▲
by
chaoyu
7y ago
hi Aaron, I'm one of the BentoML aurthors - great suggestion on pipreqs, will look into incorparating that into BentoML! It should be very straightforward adding support for saving/loading Statsmodels in BentoML. In fact you shoul
20.
▲
by
chaoyu
7y ago
Graphpipe solves a very unique problem when building ML model serving system, although BentoML is trying to solve a very different problem. We think it would be interesting to support GraphPipe's flatbuffer format in BentoML's RES
21.
▲
by
chaoyu
7y ago
Our quick start guide notebook on Google Colab is also a great place to get started! https://colab.research.google.com/github/bentoml/BentoML/blo...
22.
▲
by
chaoyu
7y ago
Thanks for sharing the project Kevlar1818. BentoML author here - we are building BentoML to empower Data Scientists to ship prediction services instead of delivering "models" to dev teams. We proposed a workflow that made it easy
23.
▲
Vote with Your Feet: SF Art Installation
(votewithyourfeet.org)
1 points
by
chaoyu
10y ago
|
0 comments