5 ms·
There are advantages to smaller models, namely you can process a lot more data, with a lot less vram. I think the intent here from the Google team is for a task
by llama_person 2y ago
There are advantages to smaller models, namely you can process a lot more data, with a lot less vram. I think the intent here from the Google team is for a task-specific VLM that you fine-tune on your data, rather than a general purpose assistant.
From my own experimentation I have found it to really pack a punch for its weight. Another small model which has been very good has been https://github.com/vikhyat/moondream https://github.com/vikhyat/moondream .
- whimsicalism 2y agofinetuning is easily within reach for llava-mistral or something like that, just rent an a100 or two for ~$20 bucks and you'll have your finetuned model
- simonw 2y agoHave you seen any good documentation anywhere on how to do that?
- llama_person 2y agoHere's a tutorial https://wandb.ai/byyoung3/ml-news/reports/How-to-Fine-Tune-LLaVA-on-a-Custom-Dataset--Vmlldzo2NjUwNTc1 https://wandb.ai/byyoung3/ml-news/reports/How-to-Fine-Tune-L... There's not really a super easy to use software solution yet, but a few different ones have cropped up. Right now you'll have to read papers to get the training recipes. - https://github.com/haotian-liu/LLaVA/blob/main/scripts/finetune_lora.sh https://github.com/haotian-liu/LLaVA/blob/main/scripts/finet... - https://github.com/InternLM/xtuner/tree/main https://github.com/InternLM/xtuner/tree/main - https://github.com/TinyLLaVA/TinyLLaVA_Factory https://github.com/TinyLLaVA/TinyLLaVA_Factory Is a pointer in the right direction, along with: https://arxiv.org/abs/2304.08485 https://arxiv.org/abs/2304.08485
- simonw 2y agoThanks!
- whimsicalism 2y agothe linked resources are good, also the llava repo has pretty good guides on how to train which you can adapt to SFT
- codester1000 2y agoLLaVA is even less open than PaliGemma, it is trained on CC-BY-NC4.0 data so it can't be used commercially. I emailed the team about it. At least with Pali-Gemma the base pt models are available to be used commercially if you fine-tune them yourself
- whimsicalism 2y agollava-mistral v1.6 is apache, my friend :)
- codester1000 2y agoyes their code is, but not their dataset if you want to use the model weights without training yourself
- codester1000 2y agoWhere did you get this info? because I would love to use it, but I went through their info and it says they use a lot of the same data as 1.5 and the acknowledgement section of their site says: "The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes" Would love it if I could use LLaVA, but don't want to spend the money on like 18 A100s for 24 hrs that they use for training it. A lot of the models using CC BY NC 4.0 datasets, like VILA, thats not available for commercial use unless you train the model yourself. This is the first time at least a research or company has been open with this info, they specifically say: only the pt models can be used with fine-tuning for commercial use.
- llama_person 2y agoyou can try out https://huggingface.co/datasets/BAAI/SVIT https://huggingface.co/datasets/BAAI/SVIT which appears to support commercial. I've not tried it yet, but it seems to be an option. If you build a smaller model, you should need much less than 18 A100s for 24 hours, though I don't disagree you'll need at least a few.
- 2y ago