5 ms·
I don't understand how so much money has been poured in to these companies? I get the why the techniques are suitable, but I just assumed who ever wants to do
by dcl 3y ago
I don't understand how so much money has been poured in to these companies?
I get the why the techniques are suitable, but I just assumed who ever wants to do this kind of retrieval can probably implement a suitable Approx. NN library themselves?
Especially so, because getting good embeddings is the hard part, not the search?
- jstx1 3y agoAnyone who wants this can implement their own library themselves? When has this worked for any problem ever? Searching efficiently is a problem, and there's several open source and proprietary solutions but I don't get how you can put it in the "everyone should roll their own" category.
- softwaredoug 3y agoAt billion vector scale, doing this yourself is pretty impossible
- deleted 3y ago[deleted]
- dmezzetti 3y agoFaiss has long discussed strategies for scaling to 1B - 1T records here - https://github.com/facebookresearch/faiss/wiki/Indexing-1G-vectors https://github.com/facebookresearch/faiss/wiki/Indexing-1G-v... There are plenty of options available to run your own local vector database, txtai is one of them. Ultimately depends if you have a sizable development team or not. But saying it is impossible is a step too far.
- deleted 3y ago[deleted]
- hiyou102 3y agoEven in that article with much smaller vectors than what GPT puts out (1536 dimensions) QPS drops below 100 if recall@1 is more than 0.4. That's to say nothing of cost of regenerating this index using incremental updates. I don't get why people on HN are so adamant on the idea that no one needs scale beyond 1 machine ever.
- dmezzetti 3y agoThe comment said that having an instance with 1B+ vectors yourself is impossible. Clearly that's not the case.
- quickthrower2 3y agoIf you have a billion vectors, is “yourself” a large tech company who does stuff like roll their own browsers, programming languages, invents kubernetes etc. Probably could roll this! And indeed sell this.
- deleted 3y ago[deleted]
- gpderetta 3y agoLast time I had to deal with vector representation of documents was more than 10 years ago, so I'm a bit rusty, but billion vector scale sound relatively trivial.
- esafak 3y agoWith retrieval time in the milliseconds? The entries may be ads, or something else user facing. Your users are not going to sit around while you leisurely retrieve them.
- VHRanger 3y agonot particularly? 1B vectors * 300dimensions * float32 (4 Bytes) ~= 1.2TB This pretty much still runs on consumer hardware. Just run that on a 4TB nvme ssd, or a RAID array of ssd's if you're frisky.
- linuxdude314 3y agoYou do realize you have to query an index of all of that data for every single query your use makes right? Computing that index is not entirely trivial, nor is the operation of partitioning the data so it fits in ram across a pool of nodes. Sure, role your own, but don’t act like making a highly scalable database is a weekend project.
- ndriscoll 3y agoI'm not familiar with the index part, but you can get at least 2TB on a single CPU socket these days. You shouldn't need multiple machines to fit in RAM. Depending on what QPS you need to handle, you might also be fine to not have the whole thing fit in RAM.
- dbthrowfu 3y agoConsumer hardware can still handle that with 1TB RAM + ThreadRipper Pro. > You do realize you have to query an index of all of that data for every single query your use makes right? Computing that index is not entirely trivial, nor is the operation of partitioning the data so it fits in ram across a pool of nodes. I don't know what any of this means -- and it sounds like you're slapping a bunch of terminology together, rather than communicating a well-thought-out idea. Yes, in the general case you're going to have to use an index. Computing an index or a key to that index? Computing the index is a solved problem, that does not have a hard real-time component -- you can do it outside of normal query executions. Computing the key to the index on each query is also a solved problem. Have dimensions stored in columnar format, generate a sparse primary index on said columns, and then use binary search to quickly find the blocks of interest to do a sequential search on viz. distance function. Or you could even just use regular old SS trees, SR trees, or M Trees for high-dimensional indexing -- they're not expensive to use at all. There, you can easily run a query on a single dimension (1 billion entries) under a second. You want 300 dimensions? Ok, parallelize it. 128 threads, easy. At most this will take 3 seconds if everything is configured properly (big IF, that seems like few can get right). This is literally a weekend project. Anyone can build something like this, but not everyone has the integrity to be upfront about how they're reinventing the wheel, and spinning it like they've just broken ground in database R&D.
- ramoz 3y agolol, not true. Even for huge vectors (1000 page docs), today you can do this with enough disk storage with something like leveldb on a single node, and in memory with something like ScaNN for nearest neighbor.
- deleted 3y ago[deleted]
- hiyou102 3y agoWhat kind of QPS are you getting and how fast are incremental index updates? That's the hard part.
- heipei 3y agoFor the same reason you have money going to various SQL-as-a-service companies that run Postgres / MySQL for you as a service: Some folks would rather eat the network latency, give up control of their data and complicate their compliance process than operating a database themselves.
- pantulis 3y agoThe difference with SQL is that it's not like storing vector embeddings outside your perimeter suppose a big compliance issue --at least in security or legal terms. Giving up control of their data and network latency are legit concerns, that's for sure.
- gk1 3y ago(I'm from Pinecone) While Pinecone isn't available as a self-hosted option (see many comments with alternatives), we do offer the option of running Pinecone for you on a managed VPC, and we do have SOC2 compliance, and we do pass enterprise-level security reviews regularly. Whether that's sufficient is up to you of course.
- hiyou102 3y agoA lot of people use some form of managed services if they are in the cloud. Be it S3 or Dynamo DB. Generally cheaper than running things yourself and operationally much easier too.
- itsoktocry 3y ago>I don't understand how so much money has been poured in to these companies? First time here? Just kidding. But not. You have to separate the VC hype with the product, because the VCs always need something to overhype. Half these people were pumping money in to crypto and whatever-the-hell-web3-is/was just a couple months ago, this is just the next thing they like. Half these companies probably aren't remotely good companies. The VC money hardly ever makes sense.
- carimura 3y agoBecause if AI is the gold rush VC's want to find the Levi's and Wells Fargo's.
- jderiksen 3y agoI am on a small team that initially rolled our own semantic search system. We quickly ran into issues around scaling, maintenance, and performance. Since we want to focus on delivering features and not turning into a DevOps team, we switched to Pinecone and it has met our needs pretty well. We would like to see auto-scaling and I believe that this feature is in the works. Support has been very responsive and helpful when we do have questions and issues. There are plenty of LLMs to choose from with regard to finding sources of embeddings. Some free, some for money.
- ShamelessC 3y agoIt is at the intersection of technology investors "know" (databases) and technology investors don't know, but have been told is about to blow up (ML). It is also effectively "roll your own Google/Shazam/whatever", which probably makes for a fancy demo to those who don't know how trivial it is to implement. Basically investors are morons on average.