2 ms·
Plus one, using the exact setup to make it scale. If Azure Doc Intelligence gets too expensive, VLMs also work great
by IndieCoder 2y ago
Plus one, using the exact setup to make it scale. If Azure Doc Intelligence gets too expensive, VLMs also work great
- vinothgopi 2y agoWhat is a VLM?
- saharhash 2y agoVision Language Model like Qwen VL https://github.com/QwenLM/Qwen2-VL https://github.com/QwenLM/Qwen2-VL or CoPali https://huggingface.co/blog/manu/colpali https://huggingface.co/blog/manu/colpali
- sidmo 2y agoVLMs are cool - they generate embeddings of the images themselves (as a collection of patches) and you can see query matching displayed as a heatmap over the document. Picks up text that OCR misses. Here's an open-source API demo I built if you want to try it out: https://github.com/DataFog/vlm-api https://github.com/DataFog/vlm-api