3 ms·
Is that not just traditional OCR applied on top of LLM?
Is that not just traditional OCR applied on top of LLM?
No it’s not, it’s a multimodal transformer model.
It's possible they have a software layer that does that. But I was assuming they don't, because the open source multimodal models don't.