3 ms·
Unified Vision-Language Agents – Detect, Segment, OCR, Generate and More
- fzysingularity 10mo agoHere's a short cookbook exploring an agentic approach to vision–language tasks: detection, segmentation, OCR, generation, and combining classical CV tools with VLM reasoning. Happy to run examples if you leave a comment. [1] IPython notebook: https://github.com/vlm-run/vlmrun-cookbook/blob/main/notebooks/12_orion_image_understanding.ipynb https://github.com/vlm-run/vlmrun-cookbook/blob/main/noteboo... [2] Colab: https://colab.research.google.com/github/vlm-run/vlmrun-cookbook/blob/main/notebooks/12_orion_image_understanding.ipynb https://colab.research.google.com/github/vlm-run/vlmrun-cook...