4 ms·
Are you running everything through a single end-to-end vision model, or do you dynamically dispatch to specialized OCR, detection, and segmentation backends?
by kernel33 1y ago
Are you running everything through a single end-to-end vision model, or do you dynamically dispatch to specialized OCR, detection, and segmentation backends?
- fzysingularity 1y agoThis demo showcases the latter approach with tool-calling - essentially filling in the gaps of current VLMs. That said, we're of course interested in folding all these capabilities into a single model, but that's going to take a bit more work. What makes this approach interesting is that our VLMs need to able to understand intermediate results (sometimes in the form of images themselves), and then delegate to other specialized tools whenever it can't perform a specific action.