Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
2uryaa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
2uryaa
7mo ago
Thank you for the feedback. Taking note of this!
2.
▲
by
2uryaa
7mo ago
Yes, we operate on GB200s and GH200s. Usually we are cheaper for many models and can get up to double the TPS.
3.
▲
by
2uryaa
7mo ago
Yep, we are actively working on getting this down. We can meet SLAs with tuning for the real time vision workloads but trying to get rid of this compromise is our next big development task.
4.
▲
by
2uryaa
7mo ago
For consumers, we want to just pass on price to performance ratio. For enthusiasts and companies, we do see people want their own models/ ability to use the massive amounts of data they have.
5.
▲
by
2uryaa
7mo ago
That's really awesome to hear!!
6.
▲
by
2uryaa
7mo ago
Hey Jack, we use GB200s for these workloads. Feel free to check those big models out on our site! We are doing Kimi, GLM, Minimax, etc.
7.
▲
by
2uryaa
7mo ago
Also curious about this. We have a 30 day content retention policy and have to have access to your fine-tuned model/LoRa if deploying that. If there's anything we can change, happy to hear it out.
8.
▲
by
2uryaa
7mo ago
We usually charge by GPU hour for those finetunes, around 8-10 dollars depending on GPU type and volume! This is similar to Modal, but since the engine is fully ours, you don't wait ~1 min for cold starts. Ideally, we will make onboard
9.
▲
by
2uryaa
7mo ago
Haha sorry for the typo! Your F500 use case is exactly who we want to target, especially as they start serving finetunes on their own data. Thanks for the feedback!
10.
▲
by
2uryaa
7mo ago
our SLA is actually higher and we are lower priced. We are also using this as a step into serving finetuned models for much cheaper than Fireworks/Together and not having the horrible cold starts of Modal. We're essentially trying
11.
▲
by
2uryaa
7mo ago
Thank you for the feedback! I think we will definitely redo the info on the frontpage to reorg and show quantizations better. For reference, Kimi and Minimax are NVFP4. The rest are FP8. But I will make this more obvious on the site itself.
12.
▲
by
2uryaa
7mo ago
Hey Oras, thank you for the feedback! I think we definitely could list on OpenRouter but as you point out, our end goal is to host finetuned models for individuals. The IonRouter product is mostly to showcase our engine. In the backend, we