3 ms·
SQLite supports raw binary BLOBs, and I would store embeddings as a simple 512*4 = 2048 bytes blob representing just the float32 embeddings. Then you can acces
by lovasoa 3y ago
SQLite supports raw binary BLOBs, and I would store embeddings as a simple 512*4 = 2048 bytes blob representing just the float32 embeddings. Then you can access the data from the buffer directly in dart: https://api.dart.dev/stable/2.8.1/dart-typed_data/ByteBuffer/asFloat32List.html https://api.dart.dev/stable/2.8.1/dart-typed_data/ByteBuffer...
- vishnumohandas 3y agoThis[1] has improved the situation considerably! ``` [log] SQLite: 100000 embeddings inserted in 2635 ms [log] SQLite: 100000 embeddings retrieved in 561 ms ``` Isar is still ~2x faster, but there's room for optimization. Thanks a bunch for sharing this, I will update the post. [1]: https://github.com/ente-io/edge-db-benchmarks/commit/51ec4965fd4138a02b2ab544191fa26e47e99e3c https://github.com/ente-io/edge-db-benchmarks/commit/51ec496...
- lqet 3y agoWow, kudos for actually implementing the proposed fix and documenting it here so quickly!
- lovasoa 3y agoLooking at the rest of your code, it looks like you are doing a lot of unnecessary processing on these embeddings. They even seem to be doubly json encoded at some point: https://github.com/ente-io/clip-ggml/blob/main/lib/clip_ggml.dart#L67 https://github.com/ente-io/clip-ggml/blob/main/lib/clip_ggml... In my opinion, they should probably just be a memory buffer representing the raw floats all the way down: from the output of the model to the database. They should never be encoded, neither in json, nor as a dart List<double>.
- vishnumohandas 3y agoType-conversions with FFI turned out to be non-trivial, so `{"embedding":[...]}` was a way out at the cost of a small performance hit (when compared to the time spent on inference). We'll take another look.
- lovasoa 3y agoThe result of clip.[...].call(...) is already a buffer, isn't it? You just have not to touch it at all. And on the cpp side, remove the json encoding and just return a raw buffer.
- CarefreeCrayon 3y agoI found this interesting and have been profiling to understand why the increase in performance was so significant. It looks to me that the main culprit was Protobufs + Garbage collection. Serializing to protos performs a lot of allocations. Using just the Float32View skips all that. ``` I/scudo ( 641): Stats: SizeClassAllocator64: 572M mapped (0M rss) in 11986660 allocations; remains 257629 I/scudo ( 641): 00 ( 64): mapped: 1024K popped: 506106 pushed: 491660 inuse: 14446 total: 15044 rss: 0K releases: 0 last released: 0K region: 0x7ceae87000 (0x7ceae86000) I/scudo ( 641): 01 ( 32): mapped: 1024K popped: 92137 pushed: 73047 inuse: 19090 total: 26708 rss: 0K releases: 0 last released: 0K region: 0x7cfae8c000 (0x7cfae86000) ``` I think this is because during the proto encoding/decoding stage the protobuf lib ended up creating a bunch of objects to support the process ``` "Class","Library","Total Instances","Total Size","Total Dart Heap Size","Total External Size","New Space Instances","New Space Size","New Space Dart Heap Size","New Space External Size","Old Space Instances","Old Space Size","Old Space Dart Heap Size","Old Space External Size" _FieldSet,package:protobuf/protobuf.dart,100010,4800480,4800480,0,0,0,0,0,100010,4800480,4800480,0 PbList,package:protobuf/protobuf.dart,108536,3473152,3473152,0,0,0,0,0,108536,3473152,3473152,0 Embedding,package:edge_db_benchmarks/models/embedding.dart,108535,3473120,3473120,0,0,0,0,0,108535,3473120,3473120,0 EmbeddingProto,package:edge_db_benchmarks/models/embedding.pb.dart,100010,1600160,1600160,0,0,0,0,0,100010,1600160,1600160,0 ``` What's missing here is that these have to be copied over to the database isolate as well.
- lovasoa 3y agoMaybe that's an opportunity for a pull request to the dart protobuf library?