Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cotran2
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
cotran2
1y ago
The model is compact 1.5B, most GPUs can serve it locally and has <100ms e2e latency. For L40s, its 50ms.
2.
▲
by
cotran2
1y ago
There is a case study comparing with RouteLLM in the appendix.
3.
▲
by
cotran2
1y ago
According to the post, the model is fine-tuned for routing to different tasks/domains. Classifying difficulty level is probably not the intended use case.