Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jiayq84
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
The missing guide to the H100 GPU market
(blog.lepton.ai)
3 points
by
jiayq84
2y ago
|
0 comments
2.
▲
by
jiayq84
2y ago
I do a startup called Lepton AI. We provide AI PaaS and fast AI runtimes as a service, so we keep a close eye on the IaaS supply chain. For the last few months we see supply chain getting better and better, so the business model that worked
3.
▲
by
jiayq84
3y ago
Full open-source code with Apache license here: https://github.com/leptonai/search_with_lepton
4.
▲
by
jiayq84
3y ago
Hi folks - Yangqing from Lepton here. The idea came from a coffee chat with a colleague on the question: how much of the RAG quality comes from the old good search engine, vs LLMs? And we figured out that the best way is to build a quick ex
5.
▲
Show HN: Conversational search in less than 500 lines of Python
(search.lepton.run)
6 points
by
jiayq84
3y ago
|
2 comments
6.
▲
Structural Decoding (Function Calling) for All Open LLMs
(leptonai.medium.com)
6 points
by
jiayq84
3y ago
|
1 comments
7.
▲
by
jiayq84
3y ago
General availability of the structured decoding capability for ALL open-source models hosted on Lepton AI. Simply provide the schema you want the LLM to produce, and all our model APIs will automatically produce outputs following the schema
8.
▲
by
jiayq84
3y ago
Super cool exhibition of what a local machine can already do in the AI frenzy!
9.
▲
by
jiayq84
3y ago
Oh wow yeah, that is a beast. Let me give it a shot.
10.
▲
by
jiayq84
3y ago
Thanks so much for the warm words!
11.
▲
by
jiayq84
3y ago
Thanks - we definitely agree that llama.cpp is great. Big fan of their optimizations. We are more or less orthogonal to the engines though - in the sense that we serve as the infra/platform to run and manage those implementations easil
12.
▲
by
jiayq84
3y ago
Thanks - the policies are listed here: https://www.lepton.ai/policies we'll put a link on our homepage. In short - we do not collect, record, or log any of your prompts and responses. They are computed in memory, retur
13.
▲
by
jiayq84
3y ago
In theory one can have 640G = 8 * 80G A100s memory and launch it. 180B Falcon with fp16 will be 360G, so there would be enough memory. It's definitely going to be very expensive indeed.
14.
▲
by
jiayq84
3y ago
Great catch! Our cloud machine encountered a cuda error (the GPU fell off PCIe) - had to restart it. It's back to normal now. All the more reason to have a managed version of services :)
15.
▲
by
jiayq84
3y ago
It's not only about "building a docker" but also maintaining multiple models, multiple environments and a lot of users. Imagine there is a group of engineers each needing to deploy their own models: one needs tensorflow 1.x,
16.
▲
by
jiayq84
3y ago
To show some actual coding examples, We have made the python library open-source at https://github.com/leptonai/leptonai/ . With it, launching a common HuggingFace model is as simple as a one liner. For example, if
17.
▲
Show HN: Running LLMs in one line of Python without Docker
(lepton.ai)
68 points
by
jiayq84
3y ago
|
26 comments
18.
▲
by
jiayq84
7y ago
I don’t want to be mean, but since you mentioned RCNN - no, you are dead wrong. RCNN was open sourced in 2014, check the repo: https://github.com/rbgirshick/rcnn Not to mention that nvidia has thrown numerous open sour
19.
▲
by
jiayq84
7y ago
Just to clarify a little bit... "At the time, very few object detection models had public implementations" - this is wrong. Almost all object detection models had public implementations starting from 2014, most notably Detectron (
20.
▲
by
jiayq84
9y ago
We are moving to Apache 2.0 in a few days. Early draft at https://github.com/Yangqing/caffe2/tree/apache pending double check to make sure we are honoring all existing contributors.
21.
▲
by
jiayq84
9y ago
Yangqing (creator and main author of Caffe/Caffe2) here. We are moving to Apache 2.0 in a few days.
22.
▲
by
jiayq84
9y ago
Yangqing (creator and main author of Caffe/Caffe2) here. We are moving to Apache 2.0 in a few days.
23.
▲
by
jiayq84
9y ago
So what we do is to keep syntax=proto2, but allow users to compile with both protobuf 2.x and protobuf 3.x libraries. Minumum need is 2.6.1. We kind of feel that this gives maximum flexibility for people who have already chosen a protobuf l
24.
▲
by
jiayq84
9y ago
Learned one more thing today!
25.
▲
by
jiayq84
9y ago
Interesting data, thanks! We chose protobuf mainly due to a good caffe adoption story and also the track record of it being compatible with many platforms (mobile, server, embedded, etc). We actually looked at thrift - which is Facebook own
26.
▲
by
jiayq84
9y ago
Thanks so much @haberman! Yep, the whole thing is a little bit confusing... We basically focused on two things: - not using extensions, as 3 does not support it - not using map, as 2 does not support it and we basically landed on restrictin
27.
▲
by
jiayq84
9y ago
Thanks! Haven't yet, but our TPMs are going to reach out for collaborations. I wish we were grad school mode where latency is <1 hour, but it pays to get things proper across multiple companies. Kindly stay tuned.
28.
▲
by
jiayq84
9y ago
Yangqing here (caffe2 and ONNX). We did use protobuf and we have an extensive discussion about its versions even, from our experience with the Caffe and Caffe2 deployment modes. Here is a snippet from the codebase: // Note [Protob
29.
▲
by
jiayq84
9y ago
Yangqing here from facebook. We consciously made it a MIT license as onnx is intended to be widely shared by a lot of participants, and MIT seems to be more widely agreeable among different parties co-owning it. It's also simpler.
30.
▲
by
jiayq84
9y ago
Yangqing here (created Caffe and Caffe2) - we are much interested in enabling this path. Historically CoreML has provided Caffe and Keras interfaces, and having ONNX / CoreML interop would help a lot for everyone to ship models more ea
More ›