Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
za_mike157
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
za_mike157
3mo ago
Thanks for this! We did see it and it was committed long after we already implemented this functionality. However, there are a lot of edge cases that this commit doesn't deal with that users are going to have to do
2.
▲
by
za_mike157
3mo ago
Interesting! I didn't see they released this. Do you know what their benchmarks are? I know for cloud run they are pretty slow
3.
▲
by
za_mike157
3mo ago
Us and the team from Modal have been upstreaming things to the GVisor repo ( https://github.com/google/gvisor/pulls ) in order to make it compatible with cuda-checkpoint and other parts of our system. While we are b
4.
▲
by
za_mike157
3mo ago
haha you are right that the title is a bit strange - should just be "Reduce GPU cold starts with snapshotting" I can't read good ;)
5.
▲
by
za_mike157
3mo ago
No we don't use it. CRIU is used for normal checkpoint/restore of Linux processes. Since we run GVisor for container isolation we use their checkpoint/restore support for the sandboxed process state. Both approaches still nee
6.
▲
by
za_mike157
3mo ago
There are a lot of similarities. They run their snapshot agent as a Kubernetes DaemonSet, whereas our implementation runs as part of the Cerebrium container runtime path. Under the hood, both approaches rely on cuda-checkpoint, since cuda-c
7.
▲
by
za_mike157
3mo ago
Hey! Yes you are correct! We have both been upstreaming changes to the main GVisor repo. However, in order to work within our own infrastructure we had to make various changes that we explain throughout the article (Open TCP connections, mu
8.
▲
Why Kubernetes Serving Breaks Down for Real-Time AI
(cerebrium.ai)
5 points
by
za_mike157
6mo ago
|
0 comments
9.
▲
by
za_mike157
7mo ago
Glad you liked it!
10.
▲
by
za_mike157
7mo ago
You are correct! From our tests, storing model weights in the image actually isn't a preferred approach for model weights larger than ~1GB. We run a distributed, multi-layer cache system to combat this and we can load roughly 6-7GB of
11.
▲
by
za_mike157
7mo ago
A lot of AI workloads require GPUs which are expensive so customers would waste money running idle machines 24/7 with low utilisation which kills gross margins. By loading containers quickly means, means we can scale up quickly as requ
12.
▲
The 1979 Design Choice Breaking AI Workloads
(cerebrium.ai)
25 points
by
za_mike157
7mo ago
|
20 comments
13.
▲
AI Companies need to partner with Serverless compute platforms vs. K8s
(cerebrium.ai)
2 points
by
za_mike157
7mo ago
|
0 comments
14.
▲
by
za_mike157
1y ago
Hey! Founder of Cerebrium here. - Runpod is one of the cheapest but it comes at the price of reliability (critical for businesses) - We have more performant cold start performance with something special launching soon here - Iterating on yo
15.
▲
by
za_mike157
2y ago
I haven't used SkyPilot so I am unfamiliar with the experience and performance. However, some of the situations you would like to use Cerebrium over Skypilot are: - You don't want to manage you own hardware - Reduced costs: With s
16.
▲
by
za_mike157
2y ago
I think we used this UI kit: https://minimals.cc/
17.
▲
by
za_mike157
2y ago
I guess then the next question would be how quickly can they start executing your container from cold start when a workload comes in? Typically we see companies on around 30-60s
18.
▲
by
za_mike157
2y ago
Do you mean why the individual file names aren't quoted? You can see an example config file at the bottom of that link you attached - agreed we should probably make it more obvious
19.
▲
by
za_mike157
2y ago
Thanks for confirming! Our cold start, excluding model load is 2-4 seconds typically for HF models. The only time it gets much longer when companies have done a lot with very specific CUDA implementations
20.
▲
by
za_mike157
2y ago
Thanks Tom! Excited to to support you and the team as you grow
21.
▲
by
za_mike157
2y ago
Ah I see they recently cut their pricing by 40% so you are correct - sorry about that. It seems we are more expensive compared to their new pricing
22.
▲
by
za_mike157
2y ago
Thank you - appreciate the kind words! Happy to continue supporting you and the team.
23.
▲
by
za_mike157
2y ago
Thank you - updated! My team makes fun of my spelling all the time!
24.
▲
by
za_mike157
2y ago
Thanks for pointing that out!
25.
▲
by
za_mike157
2y ago
Modal is a great platform! In terms of cold starts, we seem to be very comparable from what users have mentioned and tests we have run. Easier config/setup is feedback we have gotten from users since we don't have and special synt
26.
▲
by
za_mike157
2y ago
Yes RunPod does have cheaper pricing than us however they don't allow you to specify your exact resources but rather charge you the full resource (see example of A100 above) so depending on your resource requirements our pricing could
27.
▲
by
za_mike157
2y ago
You are correct! After the first request, an image will be on a machine and it’s cached for future use. This makes subsequent container startups much faster. We also route requests to machines where the image is already cached as well as de
28.
▲
Launch HN: Cerebrium (YC W22) – Serverless Infrastructure Platform for ML/AI
48 points
by
za_mike157
2y ago
|
32 comments
29.
▲
Show HN: Turning a Live Video into Instant Purchases Using AI
(live-stream-shopper.cerebrium.ai)
3 points
by
za_mike157
2y ago
|
0 comments
30.
▲
by
za_mike157
2y ago
That is only for the data processing step that I run locally on my Mac to embed all his Youtube videos and upload to the VectorDB. Im not running Deepgram locally there
More ›