4 ms·
I'm sure the database is nice for other reasons, but also without really understanding the problem it sounds like you could represent two separate things here:
by evmar 3y ago
I'm sure the database is nice for other reasons, but also without really understanding the problem it sounds like you could represent two separate things here: (1) metadata about photos and their metadata sync state, (2) embeddings.
You could store a dense mapping to integers in (1) and then represent (2) as a separate file that is literally just packed floats. Photo number n has its embeddings start at file offset n * 512 * sizeof(float). Updating any given photo's embeddings is an easy seek+write.
Given that the use case of similarity search is to stream through the embeddings and not any complex seeking/rewriting/indexing/etc, feels like a relatively simple bit of code.
The post talks about serializing 100k photos. With 32-bit floats each embedding vector is 4*512=2048 bytes, so that's ~205mb of data, much less than the >500mb quoted in the post.