Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danlenton
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
danlenton
2y ago
will be soon!
32.
▲
by
danlenton
2y ago
aha maybe we should change our slogan to "Uber for LLMs"
33.
▲
by
danlenton
2y ago
Yep! Although MoE use several "expert" linear layers within a single network, and generally the "routing" is not based on high-level semantics, but token specialization, as disuccussed by Fuzhao Xue, author of OpenMoE, i
34.
▲
by
danlenton
2y ago
yeah we might need to implemenet some kind of "I am not a robot" checks soon, as well as 2FA.
35.
▲
by
danlenton
2y ago
definitely on the cards, we're keeping our options open here. Right now just focused on creating value though.
36.
▲
by
danlenton
2y ago
Thanks for sharing! These are useful toos, but they are a bit different, more based on similarity search in prompt space (a bit like semantic router: https://github.com/aurelio-labs/semantic-router ). Our router uses a
37.
▲
by
danlenton
2y ago
As outlined, DSPy is a create tool for this. Currently their focus is on optimizing in-context examples, but their broader vision includes optimising system prompts for intermediate LLM nodes in an agentic system as well. We uploaded an exp
38.
▲
by
danlenton
2y ago
Agreed, we've spoken to tons of users who reach out to us and start the conversation with "we've tried to implement this ourselves".
39.
▲
by
danlenton
2y ago
Another point here is that some users prefer to use their own API keys for the backend providers (a feature we're releasing soon). Any "discounts" would then be harder to implement. I do generally think it's much cleaner
40.
▲
by
danlenton
2y ago
that's a good point, impartiality would then be questioned
41.
▲
by
danlenton
2y ago
I certainly wouldn't complain about this lol
42.
▲
by
danlenton
2y ago
The idea is that at some point in future, we release new and improved router configurations which do take small margins, but from the user perspective they're still paying less than using a single endpoint. We don't intend to infl
43.
▲
by
danlenton
2y ago
To be honest we're thinking of moving away from this soon anyway. Open source models will soon make for perfectly good judges (or juries), with Llama3 etc.: https://arxiv.org/pdf/2404.18796
44.
▲
by
danlenton
2y ago
Super helpful feedback, thanks for going so deep! I agree that for the really heavy agentic stuff, the router in it's current form might not be the most important innovation. However, for several use cases speed is really paramount, an
45.
▲
by
danlenton
2y ago
exactly!
46.
▲
by
danlenton
2y ago
definitely similar! I'm a fan of Alex and his work on OpenRouter :) Some of the main differences would be: - we focus on performance based routing, optimizing speed, cost and quality [ https://youtu.be/ZpY6SIkBosE ] - we
47.
▲
by
danlenton
2y ago
This is a great point. With models becoming more intelligent, they're seeming to become less brittle to the subtleties in the prompts, which might mean decoupling will occur naturally anyway. With regards to customers wanting to stick
48.
▲
by
danlenton
2y ago
MoE LLMs use several "expert" fully connected layers, which are routed to during the forward pass, all trained end-to-end. This approach can also work with black-box LLMs like Opus, GPT4 etc. It's a similar concept but operat
49.
▲
by
danlenton
2y ago
Very sorry about that! I had an issue with my google calendar, I set it to "do not send emails" but for some reason some still came through. Fixed + removed now.
50.
▲
by
danlenton
2y ago
it depends on the task, but this video gives an idea :) https://www.youtube.com/watch?v=9JYqNbIEac0
51.
▲
by
danlenton
2y ago
we intend to eventually have routers which improve the speed, cost (and maybe quality) to such an extent that we can then take some margins from these best performing routers, with users still saving costs compared to individual endpoints.
52.
▲
by
danlenton
2y ago
They're not used as labels directly (we're not trainig an LLM which outputs text), they are used as an intermediate step, which is then used to compute a simple score which the neural score function is then trained on. The neural
53.
▲
by
danlenton
2y ago
Great analogy, I'm not sure tbh. I don't think we will see quite as many unique models as we see unique websites, but I do think we're going to see an increasing number of divergent and specialized models, which lend themselv
54.
▲
by
danlenton
2y ago
merged yesterday! https://github.com/run-llama/llama_index/pull/12921
55.
▲
by
danlenton
2y ago
Thanks - glad to hear the idea resonates!
56.
▲
by
danlenton
2y ago
Great question! Generally the neural network used for the router takes maybe ~20ms during inference. When deployed on prem, in your own cloud environment, then this is the only latecy. When using the public endpoints with our own intermedia
57.
▲
Show HN: Route your prompts to the best LLM
(unify.ai)
298 points
by
danlenton
2y ago
|
126 comments
58.
▲
Unify (YC W23) Is Hiring Contributors to Unify ML
(ycombinator.com)
1 points
by
danlenton
2y ago
59.
▲
Unify (YC W23) Is Hiring
(apply.unify.ai)
1 points
by
danlenton
3y ago
60.
▲
by
danlenton
3y ago
I think different providers are just trying to provide value at different areas in the market. OctoAI are consistently the most cost effective, but typically not as fast, while others are fast but come at a premium. In general, some provide
More ›