Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zhwu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
zhwu
11d ago
RL is not just an AI engineer problem but also a problem for AI infra: many components that can lands on different places with special requirements on scheduling during launching and recovering. We put this blog to dive into the infra chall
2.
▲
Run cloud agents on machines you manage
(cursor.com)
1 points
by
zhwu
25d ago
|
2 comments
3.
▲
by
zhwu
25d ago
It seems that cursor now supports your own sandbox
4.
▲
VRAM Ghost Busting: Who You Gonna Close()?
(hcompany.ai)
4 points
by
zhwu
3mo ago
|
0 comments
5.
▲
by
zhwu
6mo ago
The most surprising part: the agent had access to both H100s and H200s. Without being told, it noticed H200s scored better and started screening ideas on H100s, then promoting winners to H200s for validation. That strategy emerged entirely
6.
▲
A collection of reproducible LLM inference engine benchmarks: SGLang vs. vLLM
(github.com)
1 points
by
zhwu
1y ago
|
0 comments
7.
▲
by
zhwu
2y ago
Cloud services, such as autoscaling EKS or AWS Batch are mostly limited by the GPU availability in a single region. That limits the scalability of jobs that can run distributedly in a large scale. AI batch inference is one of the examples,
8.
▲
by
zhwu
2y ago
This recent blog actually looks into the case with multiple writers and the distribution for the time for a writer to take the lock: https://blog.skypilot.co/abusing-sqlite-to-handle-concurrenc...
9.
▲
Efficient GPU Resource Management for ML Workloads Using SkyPilot, Kueue on GKE
(github.com)
2 points
by
zhwu
2y ago
|
0 comments
10.
▲
by
zhwu
2y ago
Dealing with all the Kubernetes pod configs / deployments is too much for an AI engineer. Being able to focus on the real model work would be super important.
11.
▲
New Recipe: Serving Llama-2 with VLLM's OpenAI-Compatible API Server
(github.com)
1 points
by
zhwu
3y ago
|
0 comments
12.
▲
Train Your Own Vicuna on Llama-2
(github.com)
3 points
by
zhwu
3y ago
|
0 comments
13.
▲
Guide on fine-tuning your own Vicuna on Llama-2
(twitter.com)
9 points
by
zhwu
3y ago
|
0 comments
14.
▲
by
zhwu
3y ago
The finetuning can tailor the model to have more customized knowledge, just like the identity knowledge of itself shown in the blog post. If you ask the original llama model, it should know nothing about SkyPilot or Vicuña, as it is trained
15.
▲
by
zhwu
3y ago
Great reference! Just want to add about hosting your own LLM vs using ChatGPT. Cost is definitely a thing to consider, but it also depends on whether it is ok to share the requests to your product with OpenAI. Also, something you cannot do
16.
▲
by
zhwu
3y ago
It is the underlying operational guide of the latest release of Vicuna-1.5: https://twitter.com/lmsysorg/status/1686794639469371393
17.
▲
by
zhwu
3y ago
This is cool! The Llama 2-70B can be hosted in my own cloud environment.
18.
▲
Serving LLM 24x Faster on the Cloud with VLLM and SkyPilot
(blog.skypilot.co)
12 points
by
zhwu
3y ago
|
1 comments
19.
▲
by
zhwu
3y ago
It seems training the Vicuna on custom dataset could be quite easy as well, according to the following: https://github.com/skypilot-org/skypilot/tree/master/llm/vic...
20.
▲
by
zhwu
3y ago
Very interesting! Quite surprised to see PaLM-2 ranked even lower than open-sourced Vicuna.
21.
▲
Biologists are moving to the clouds with SkyPilot from UC Berkeley
(twitter.com)
5 points
by
zhwu
3y ago
|
0 comments
22.
▲
by
zhwu
3y ago
SkyPilot is actually the tool that helps you find the resources on any cloud, including AWS, GCP, Azure, IBM (comming soon) or even Lambda Clouds. It can automatically search for the spot instances across all the regions and clouds, based o
23.
▲
Vicuna releases its secrete of finding available A100s on the cloud to train it
(twitter.com)
4 points
by
zhwu
3y ago
|
2 comments
24.
▲
by
zhwu
3y ago
You need to use the transformers from the main branch instead of the pypi version, because the llama support is recently added. According to the readme of the repo, you need to install transformers with: pip3 install git+ https://
25.
▲
by
zhwu
3y ago
Wow, that is very interesting. Would you mind sharing the prompt you used to query the model?
26.
▲
by
zhwu
3y ago
Yes, you need to convert the original LLaMA model to the huggingface format, according to https://github.com/lm-sys/FastChat#vicuna-weights and https://huggingface.co/docs/transformers/main&#x
27.
▲
by
zhwu
3y ago
If you follow this command in their instruction, the delta will be automatically downloaded and applied to the base model. https://github.com/lm-sys/FastChat#vicuna-13b : `python3 -m fastchat.model.apply_delta --base &#
28.
▲
by
zhwu
3y ago
It is mainly because of the legal issues caused by the license of llama model weights. We need to figure it out with Meta's llama team before releasing.
29.
▲
by
zhwu
4y ago
They even have a eval page showing that they beat Bard by only training on ShareGPT. https://vicuna.lmsys.org/eval/