4 ms·
Yeah, I also found that for ultra low footprint models ORT is a big portion of the total payload, because it contains logic for general ONNX graph operations. I
by salamo 3mo ago
Yeah, I also found that for ultra low footprint models ORT is a big portion of the total payload, because it contains logic for general ONNX graph operations. In my case I found that ORT alone was 3.4MB over the wire, so I swapped it out for a tiny wasm that was 850x smaller and only contained the operations I needed: https://blog.lukesalamone.com/posts/creating-tiny-semantic-search/ https://blog.lukesalamone.com/posts/creating-tiny-semantic-s...
- scoriiu 3mo agodid you skip simd just because the model's tiny? naive conv perf is honestly the only reason i haven't done exactly this for the cnn
- salamo 3mo agoYeah, the model is small enough that inference is already basically instant for my usecase (only 6 transformer layers for the blog search).