10 ms·
PaliGemma: Open-Source Multimodal Model by Google
- gliched_robot 2y agoIf anyone wants to try this out, here is the hugging-face spaces demo: https://huggingface.co/spaces/google/paligemma https://huggingface.co/spaces/google/paligemma
- KaoruAoiShiho 2y agoIsn't this outdated already. It's 2 models slapped together rather than natively multimodal.
- maciejgryka 2y agoIsn't "two models slapped together" basically how all of these things work, starting with CLIP? Not sure about GPT4o, obviously, I don't think they released any underlying architecture details?
- whimsicalism 2y agoYour understanding is correct, even GPT4o will have an encoder model.
- tkellogg 2y agowhat even is “a model”? I’m not sure there is a technical definition that corresponds to how it’s used by the tech public - single interconnected neural network (LLM attention layers break this, autoencoders complicate this) - single training pass (LLMs have multiple passes, GANs have a single but produce multiple models)
- whimsicalism 2y ago> - single training pass (LLMs have multiple passes, GANs have a single but produce multiple models) LLMs have multiple passes? wdym?
- whimsicalism 2y agoNo, that's how all of these models work.
- JeremyHerrman 2y agoFrom TFA, PaliGemma is competitive to GPT-4o and even beats it in terms of speed and OCR accuracy. It also can do object detection (bounding boxes) and segmentation which GPT-4V/o and Claude 3 Opus can't do at all. Not to mention it's built to be fine tuned and commercially permissive!
- resource_waste 2y agoDid I read that its a 3B model? That means it can run on my 3060 laptop? Also, I imagine there isnt some oobabooga/automatic1111 for this?
- seaal 2y agoLM Studio is pretty nice and simple to setup. https://lmstudio.ai https://lmstudio.ai
- sroussey 2y agoCan probably run on your iPhone
- resource_waste 2y agoCPU isnt realistic, regardless what the marketers tell you.
- simonw 2y agoiPhones have GPUs. I've successfully run Mistral on mine using this app: https://llm.mlc.ai/docs/deploy/ios.html https://llm.mlc.ai/docs/deploy/ios.html
- resource_waste 2y agoPlease don't say 'what is technically possible' It blurs what is realistically useful. Also 'GPUs' No. This is a major red flag man. They do not. The marketers told you this, and you believed them.
- simonw 2y agoI don't get it. The iPhone does have a GPU, and it can be used to accelerate AI workloads. This isn't theoretical. You can see this for yourself by installing https://apps.apple.com/us/app/mlc-chat/id6448482937 https://apps.apple.com/us/app/mlc-chat/id6448482937 and turning off your wifi and running the quite capable Mistral 7B Instruct LLM on your phone. I've even used it to answer simple questions while I was offline and only had my phone with me. What am I missing here?
- abrichr 2y agoExcited to test how this performs compared to MiniCPMv2, especially when analyzing GUI images: https://github.com/OpenAdaptAI/OpenAdapt/issues/637 https://github.com/OpenAdaptAI/OpenAdapt/issues/637
- maciejgryka 2y agoI haven’t tried this yet, excited to see how it can do segmentation by outputting series of coordinates! That's something I just assumed transformers will generally be bad at.
- yeldarb 2y agoHow it does this is really cool. It’s got a VAE decoder. Reminds me a lot of how SAM works.
- ChrisArchitect 2y agoOfficial docs page shared yesterday: https://news.ycombinator.com/item?id=40358461 https://news.ycombinator.com/item?id=40358461 Hugging Face blog post similar to OP: https://huggingface.co/blog/paligemma https://huggingface.co/blog/paligemma
- jackienotchan 2y agoEverybody trashed Google yesterday, but I actually think they are catching up in the AI race. They will now start fully leveraging their distribution advantage across products and platforms.
- mrbungie 2y agoDon't hold your breath. Gemini 1.5 Pro/Ultra was meant to be the hottest shit ever, and we know how it ended up mere days later.
- mi_lk 2y agoSpeaking for yourself. I have accesses to both and to my surprise I use Gemini more than GPT4
- simonw 2y agoHow did it end up?
- Jensson 2y agoGemini 1.5 is just a few points behind gpt-4 on chatbot arena but has way larger context size and is free to use and you can disable the censorship if you want, to me that is pretty awesome. For my uses it is much better than any other model available today, and it is free on top.
- joaogui1 2y agoGemini 1.5 Ultra was never announced
- iFire 2y agoSince it took me 5 minutes to find, the terms of use of PaliGemma are not FOSS therefore not open-source. Like not on the list on https://opensource.org/licenses https://opensource.org/licenses. https://ai.google.dev/gemma/terms https://ai.google.dev/gemma/terms I was hopeful it was a change of license to be FOSS.
- cma 2y ago>(e) "Model Derivatives" means all (i) modifications to Gemma, (ii) works based on Gemma, or (iii) any other machine learning model which is created by transfer of patterns of the weights, parameters, operations, or Output of Gemma, to that model in order to cause that model to perform similarly to Gemma, including distillation methods that use intermediate data representations or methods based on the generation of synthetic data Outputs by Gemma for training that model. For clarity, Outputs are not deemed Model Derivatives. What's the difference between "Outputs of Gemma"/"Outputs By Gemma" (both included in "Model Derivatives") and Outputs ("not deemed Model Derivatives").
- joaogui1 2y agoI think (iii) is about models trained using Gemma output, while the "For clarity" part says that the Output itself is not a Model Derivative
- Hizonner 2y agoWhat is with these ML idiots and their compulsive abuse of the phrase "open source", anyway?
- surajrmal 2y agoOSI's definition of open source is one definition of it, and not necessarily one everyone agrees with. Folks claiming their work is open source are not beholden to this one organization's opinion about what can and cannot be considered open source. That's not to say I agree with the choice of using it in this instance, but I respect the fact that the term isn't precise enough to say it's wrong.
- visarga 2y agoDoesn't know JSON and has many OCR errors on documents.
- JeremyHerrman 2y agoHave you tested PaliGemma's OCR abilities? The article says it does well: "In average accuracy, we saw 85.84%, beating all other OCR models except for Anthropic’s Claude 3 Opus."
- yeldarb 2y agoIt’s very good. And the cool thing is it’s made for fine tuning also. Excited to see how fine-tuned OCR models do.
- whimsicalism 2y agoThere are plenty of better OSS vLLMs already released :) unless you need a 3b, I wouldn't overhype this.
- llama_person 2y agoThere are advantages to smaller models, namely you can process a lot more data, with a lot less vram. I think the intent here from the Google team is for a task-specific VLM that you fine-tune on your data, rather than a general purpose assistant. From my own experimentation I have found it to really pack a punch for its weight. Another small model which has been very good has been https://github.com/vikhyat/moondream https://github.com/vikhyat/moondream .
- whimsicalism 2y agofinetuning is easily within reach for llava-mistral or something like that, just rent an a100 or two for ~$20 bucks and you'll have your finetuned model
- simonw 2y agoHave you seen any good documentation anywhere on how to do that?
- llama_person 2y agoHere's a tutorial https://wandb.ai/byyoung3/ml-news/reports/How-to-Fine-Tune-LLaVA-on-a-Custom-Dataset--Vmlldzo2NjUwNTc1 https://wandb.ai/byyoung3/ml-news/reports/How-to-Fine-Tune-L... There's not really a super easy to use software solution yet, but a few different ones have cropped up. Right now you'll have to read papers to get the training recipes. - https://github.com/haotian-liu/LLaVA/blob/main/scripts/finetune_lora.sh https://github.com/haotian-liu/LLaVA/blob/main/scripts/finet... - https://github.com/InternLM/xtuner/tree/main https://github.com/InternLM/xtuner/tree/main - https://github.com/TinyLLaVA/TinyLLaVA_Factory https://github.com/TinyLLaVA/TinyLLaVA_Factory Is a pointer in the right direction, along with: https://arxiv.org/abs/2304.08485 https://arxiv.org/abs/2304.08485
- simonw 2y agoI don't understand the dog segmentation example, where the prompt was "segment dog". The article shows a screenshot with a red overlay on the dog - how was that data returned by the model, did it return a co-ordinate polygon of some sort? UPDATE: Figured it out using this tool, the mask looks like this: https://huggingface.co/spaces/google/paligemma https://huggingface.co/spaces/google/paligemma <loc0099><loc0092><loc0874><loc0926><seg014><seg009><seg123><seg126><seg004><seg074><seg092><seg112><seg000><seg021><seg099><seg015><seg096><seg043><seg012><seg019>
- sigmoid10 2y agoYou could have also just looked at the documentation linked in the original post...
- simonw 2y agoHave you seen any documents that help explain how to interpret those tokens?
- zerojames 2y ago(I work at Roboflow) We're actively working on this! Our ML team has noted that the segmentation masks in particular from PaLiGemma are a bit tedious and unintuitive to decode. We should be pushing out (more!) open source software that uses this model in the coming days. Look forward to an easy way to fine-tune PaLiGemma and broader support for its task types in `inference`, the package used in the blog post.
- llama_person 2y agohttps://huggingface.co/spaces/google/paligemma/blob/main/paligemma_parse.py#L132 https://huggingface.co/spaces/google/paligemma/blob/main/pal... the blog post details it but essentially to convert from PaliGemma tokens to bbox: y0 / 1024 * h x0 / 1024 * w y1 / 1024 * h x1 / 1024 * w have not played with segmentation yet.
- sigmoid10 2y ago
- simonw 2y agoI get really nervous about using LLMs for OCR due to the risk of both prompt injection (instructions in the OCR'd text causing the model to behave differently) and also the risk from safety filters - I don't want an OCR tool that refuses to output text from a document if that text contains offensive language, for example.
- actionfromafar 2y agoWell if it refuses, that's at least easier to spot and handle than if direction of the output is "re-directed" so to speak.
- croemer 2y agoThis is not the official announcement. It's from some blog with horrible picture quality: https://blog.roboflow.com/content/images/2024/05/image-11.png https://blog.roboflow.com/content/images/2024/05/image-11.pn...
- deleted 2y ago[deleted]
- monkeydust 2y agoAnyone else challenged with keeping up with all the releases over last week? All I need is a simple guide that tells me for task X model Y is best given benchmark Z. Where can I find that?
- swyx 2y agohuggingface spaces that serve as leaderboards. then you will need to pay a lot more for actual experts to tell you why benchmark Z is bullshit and model Y2 is actually better for the task you're actually trying to do and btw would you like to develop your own because that's a moat.
- Jensson 2y ago> then you will need to pay a lot more for actual experts to tell you why benchmark Z is bullshit and model Y2 is actually better for the task Or you get that for free here on HN.
- mrKnowNothing 2y agoDoes anyone have any idea how to interpret the 'segment' part? I ran example 'segment cat' from the example provided in the PaliGemma demo and it responded this <loc0055><loc0115><loc1023><loc1023><seg063><seg108><seg045><seg028><seg056><seg052><seg114><seg005><seg042><seg023><seg084><seg064><seg086><seg077><seg090><seg054> no documentation explanation on how to interpret segment token,