4 ms·
What’s the best library for fine-tuning VLMs at the moment and do they support this architecture or that for the IBM Granite vision models? Document understandi
by bugglebeetle 2y ago
What’s the best library for fine-tuning VLMs at the moment and do they support this architecture or that for the IBM Granite vision models? Document understanding tasks seem in special need of fine-tuning.
- _ea1k 2y agoIt looks like the model itself is here: https://huggingface.co/ds4sd/SmolDocling-256M-preview https://huggingface.co/ds4sd/SmolDocling-256M-preview It was fine tuned from this: https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct There's an example of fine tuning the base that would likely be applicable to this one as well.