Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
borzunov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
borzunov
3y ago
Hi, a Petals dev here. You're right, there's no point in using Petals if your machine has enough GPU memory to fit the model and you're okay with the quantization quality. We developed Petals for people who have less GPU memo
2.
▲
by
borzunov
3y ago
Hi, a Petals dev here. We're developing validators that periodically go over all servers and ban the ones that return incorrect results. Additionally, clients can run data through multiple disjoint routes in the network and check that
3.
▲
by
borzunov
3y ago
Hi, a Petals dev here. </s> means "end of sequence" for LLMs. If a model generates it, it forgets everything and continues with an unrelated random text (I'm sorry to hear that the model generated a disturbing text
4.
▲
Talk to Falcon 180B-Chat running over Petals
(chat.petals.dev)
1 points
by
borzunov
3y ago
|
0 comments
5.
▲
Chat with Llama 2 (70B) running over the Petals P2P network
(chat.petals.dev)
4 points
by
borzunov
3y ago
|
0 comments
6.
▲
by
borzunov
3y ago
We've moved to a new domain, the chat is now at https://chat.petals.dev
7.
▲
Petals runs Llama 2 (70B) from Colab at 5 tokens/sec
(github.com)
5 points
by
borzunov
3y ago
|
3 comments
8.
▲
by
borzunov
3y ago
Chat web app: http://chat.petals.ml Colab: https://colab.research.google.com/drive/1uCphNY7gfAUkdDrTx21... Project description: https://github.com/bigscience-workshop/petals
9.
▲
by
borzunov
3y ago
Consider making a Streamlit or Gradio app and hosting it at Hugging Face Spaces: https://huggingface.co/spaces/launch This allows you to make a simple frontend using Python only, then host it for free.
10.
▲
Make a PyTorch model tensor-parallel in one line of code
(github.com)
4 points
by
borzunov
4y ago
|
0 comments
11.
▲
by
borzunov
4y ago
Sure! The point system is being developed, we'll ship it soon. Once it's ready, you'll be able to spend points on high-priority requests to increase the speed.
12.
▲
by
borzunov
4y ago
Right now most CPUs are orders of magnitude slower than GPUs for doing forward/backward passes, so you're unlikely to get a similar speed. Some kind of pruning may help though.
13.
▲
by
borzunov
4y ago
A Hivemind/Petals dev here. As far as I understand, most federated learning methods can't efficiently train very large models (with billions of parameters) because they repeat some calculations on many peers and/or involve ex
14.
▲
by
borzunov
4y ago
A Petals dev here. FlexGen is good at high-throughput inference (generating multiple sequences in parallel). During single-batch inference, it spends more than 5 sec/token in case of GPT-3/BLOOM-sized models. So, I believe 1 sec&#
15.
▲
by
borzunov
4y ago
A Petals dev here. Recent models indeed outperform BLOOM with less parameters (for English). However, the largest LLaMA still doesn't fit into one consumer-grade GPU, and these models still benefit from increasing the number of paramet
16.
▲
by
borzunov
4y ago
A Petals dev here. We say up front that "Single-batch inference runs at ≈ 1 sec per step (token)". In turn, "parallel inference" refers to the high-throughput scenario when you generate multiple sequences in parallel. Th
17.
▲
by
borzunov
4y ago
While I agree that throughput-focused scenarios exist and this work may be valuable for them, I still think that the repository can be improved to avoid "overselling". The fact that the FlexGen's single-batch generation perfo
18.
▲
by
borzunov
4y ago
I'm afraid that, unlike proprietary APIs and Petals, this system can't be used for single-batch inference of 175B models with interactive speeds - the thing you actually need for running ChatGPT and other interactive LM apps. See
19.
▲
by
borzunov
4y ago
Note that the authors report the speed of generating many sequences in parallel (per token): > The batch size is tuned to a value that maximizes the generation throughput for each system. > FlexGen cannot achieve its best throughput i
20.
▲
by
borzunov
4y ago
During the training, participants only exchange tensors (embeddings, gradients) and never send code to each other. No other peer can execute arbitrary code on your computer - they can only request you to run one of pre-defined BLOOM layers.
21.
▲
by
borzunov
4y ago
It varies from time to time. You can also switch to the few-shot mode to try machine translation, code generation, or other tasks involving longer responses
22.
▲
by
borzunov
4y ago
Chat bot interfaces are only a small part of what can be done with large LMs. You can use and fine-tune them to solve almost all existing natural language processing tasks: machine translation, recommendation/search, text classificatio
23.
▲
by
borzunov
4y ago
Theoretical best-case for RAM offloading is 5.5 sec/token, for SSD offloading - 22 sec/token. Implementations we've tested are not faster than 10 sec/token though. See details in our paper: https://arxiv.org&#
24.
▲
by
borzunov
4y ago
In case of offloading, the computations are usually still performed on GPU, but the model is hosted in RAM/SSD instead of the GPU memory (and its chunks are copied to the GPU memory when necessary).
25.
▲
by
borzunov
4y ago
BLOOM is a large LM, and Petals is a tool for running large LMs (not necessarily BLOOM). People using Petals should still follow the model's terms of use regardless of how the tool is licensed.
26.
▲
by
borzunov
4y ago
There's a lightweight HTTP API for inference: https://github.com/borzunov/chat.petals.ml#http-api-methods
27.
▲
by
borzunov
4y ago
Petals runs BLOOM, an open-source, publicly released model of the same size as GPT-3. Here's a description of the data used to train this model: https://huggingface.co/bigscience/bloom (the "Training" se
28.
▲
by
borzunov
4y ago
If someone wants to process sensitive data and is okay with 10x slowdown, it's better to use offloading. This is another, slower method for running large LMs locally without high-end GPUs, see details here: https://news.ycom
29.
▲
by
borzunov
4y ago
You're right. This comment explains offloading in more detail: https://news.ycombinator.com/item?id=34216213
30.
▲
by
borzunov
4y ago
clarification: You can also use offloading on Colab, but inference with offloading is at least 10x slower (see other comment threads). So it can't really be used for interactive inference, but may be used for fine-tuning with large bat
More ›