Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jrandolf
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
jrandolf
1mo ago
This is hilarious. Are you the owner?
2.
▲
Show HN: UI Inspector for Tauri
(github.com)
1 points
by
jrandolf
1mo ago
|
0 comments
3.
▲
by
jrandolf
6mo ago
This plugin uses sqlx underneath which handles prepared statement caching. Regarding migration, we just used a coding agent to migrate our database infrastructure to it. It takes <20 minutes and remember this really only helps with stati
4.
▲
Show HN: sqlc-gen-sqlx, a sqlc plugin for generating sqlx Rust code
(github.com)
1 points
by
jrandolf
6mo ago
|
2 comments
5.
▲
by
jrandolf
6mo ago
We are aware of this. There was a bug that overcounted and now it's been fixed. If you'd like for us to delete your account, please contact support@sllm.cloud.
6.
▲
by
jrandolf
6mo ago
Fixed.
7.
▲
by
jrandolf
6mo ago
First, thanks for signing up early. It means a lot. The $10/mo price needed 465 people to fill a cohort before we could turn on a single GPU. People signed up and churned while waiting, so we looked at the reservation pattern and deter
8.
▲
by
jrandolf
6mo ago
We collect emails to notify you when the cohort fills or any important information such as cancellation. No one's selling your email. Also, please read https://news.ycombinator.com/newsguidelines.html . HN is a communit
9.
▲
by
jrandolf
6mo ago
The audience here is developers buying API access. They want to see the model, the price, and the throughput, not a hero image and three paragraphs about our mission. Marketing copy between a developer and that information is friction.
10.
▲
by
jrandolf
6mo ago
See https://news.ycombinator.com/item?id=47670843
11.
▲
by
jrandolf
6mo ago
Yes.
12.
▲
by
jrandolf
6mo ago
You're right that we're less flexible than OpenRouter or Chutes. We don't let you hop between models per-request. If you want that, use those. If you want predictable cost and guaranteed throughput on one model, that's u
13.
▲
by
jrandolf
6mo ago
15-25 was a rate based on oversubscription. Now it's 60 like others :).
14.
▲
by
jrandolf
6mo ago
Thanks to everyone who shared feedback. We’re implementing it now. Here’s what’s changed: - We’ve removed the other LLMs for now and are focusing entirely on Qwen 3.5. We’ll bring back additional smaller models later, but most usage was alr
15.
▲
by
jrandolf
6mo ago
You get an API key
16.
▲
by
jrandolf
6mo ago
The problem is different. OpenRouter is a router to LLMs. It doesn't solve GPU underutilization.
17.
▲
by
jrandolf
6mo ago
20 tok/s is an average. It can be more, it can be less. If you are running off-peak I'm sure you'd get some crazy number.
18.
▲
by
jrandolf
6mo ago
Going on ChatGPT.com and using their AI for 24 hours doesn't mean you are actually using their LLM for 24 hours. It's only live for as long as the output is being generated. You reading, waiting for tool calls, etc. don't cou
19.
▲
by
jrandolf
6mo ago
There is vast.ai!
20.
▲
by
jrandolf
6mo ago
Multiplexing on a GPU cloud.
21.
▲
by
jrandolf
6mo ago
I'm feeling it Mr. Crabs.
22.
▲
by
jrandolf
6mo ago
Not if you are the only one. We have rate limits to prevent this in case, idk, you share your key with 1000 people lol.
23.
▲
by
jrandolf
6mo ago
No cohorts have been filled yet. We're still early. We are seeing reservations pick up quickly, but I'd be able to give you a more concrete estimate of fill velocity after about a week. That said, we're planning to add a 7-da
24.
▲
by
jrandolf
6mo ago
24/7 LLM for $10/month.
25.
▲
by
jrandolf
6mo ago
Yes
26.
▲
by
jrandolf
6mo ago
We implement rate-limiting and queuing to ensure fairness, but if there are a massive amount of people with huge and long queries, then there will be waits. The question is whether people will do this and more often than not users will be i
27.
▲
by
jrandolf
6mo ago
Thanks lol. I actually like Shadcn's style. It's sad that people view it as AI now.
28.
▲
by
jrandolf
6mo ago
vLLM handles GPU scheduling, not sllm. The model weights stay resident in VRAM permanently so there's no loading/unloading per request. vLLM uses continuous batching, so incoming requests are dynamically added to the running batch
29.
▲
by
jrandolf
6mo ago
OpenRouter is a little different. We are trying to experiment with maximizing a single GPU cluster.
30.
▲
by
jrandolf
6mo ago
1. It's an average. 2. We have sophisticated rate limiter.
More ›