5 ms·
The wait is finally over. One or two iterations, and I’ll be happy to say that language models are more than fulfilling my most common needs when self-hosting.
by originalvichy 6mo ago
The wait is finally over. One or two iterations, and I’ll be happy to say that language models are more than fulfilling my most common needs when self-hosting. Thanks to the Gemma team!
- adamtaylor_13 6mo agoWhat sort of tasks are you using self-hosting for? Just curious as I've been watching the scene but not experimenting with self-hosting.
- irishcoffee 6mo agoI would personally be much more interested in using LLMs if I didn’t need to depend on an internet connection and spending money on tokens.
- vunderba 6mo agoNot OP but one example is that recent VL models are more than sufficient for analyzing your local photo albums/images for creating metadata / descriptions / captions to help better organize your library.
- kejaed 6mo agoAny pointers on some local VLMs to start with?
- canyon289 6mo agoYou could try Gemma4 :D
- vunderba 6mo agoThe easiest way to get started is probably to use something like Ollama and use the `qwen3-vl:8b` 4‑bit quantized model [1]. It's a good balance between accuracy and memory, though in my experience, it's slower than older model architectures such as Llava. Just be aware Qwen-VL tends to be a bit verbose [2], and you can’t really control that reliably with token limits - it'll just cut off abruptly. You can ask it to be more concise but it can be hit or miss. What I often end up doing and I admit it's a bit ridiculous is letting Qwen-VL generate its full detailed output, and then passing that to a different LLM to summarize. - [1] https://ollama.com/library/qwen3-vl:8b https://ollama.com/library/qwen3-vl:8b - [2] https://mordenstar.com/other/vlm-xkcd https://mordenstar.com/other/vlm-xkcd
- BoredPositron 6mo agoI use local models for auto complete in simple coding tasks, cli auto complete, formatter, grammarly replacement, translation (it/de/fr -> en), ocr, simple web research, dataset tagging, file sorting, email sorting, validating configs or creating boilerplates of well known tools and much more basically anything that I would have used the old mini models of OpenAI for.
- ktimespi 6mo agoFor me, receipt scanning and tagging documents and parts of speech in my personal notes. It's a lot of manual labour and I'd like to automate it if possible.
- mentalgear 6mo agoAdding to the Q: Any good small open-source model with a high correctness of reading/extracting Tables and/of PDFs with more uncommon layouts.
- mh- 6mo agoI haven't tried it yet, but I bookmarked this recently: https://github.com/opendataloader-project/opendataloader-pdf https://github.com/opendataloader-project/opendataloader-pdf
- mentalgear 6mo agoThank you, looks great!
- vunderba 6mo agoStrongly agree. Gemma3:27b and Qwen3-vl:30b-a3b are among my favorite local LLMs and handle the vast majority of translation, classification, and categorization work that I throw at them.
- misiti3780 6mo agowhat HW are you running them on ? are you using OLLAMA ?
- vunderba 6mo agoI'm using the default llama-server that is part of Gerganov's LLM inference system running on a headless machine with an nVidia 16GB GPU, but Ollama's a bit easier to ease into since they have a preset model library. https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp
- curioussquirrel 6mo agoGive Gemma 31B a shot for translation, it does a very good job at that given its size.
- kolja005 6mo agoI would be inclined to agree with this except that my "most common needs" keeps expanding and increasing in difficulty each year. In 2023 and 2024, most of my needs were asking models simple questions and getting a response. They were a drop-in replacement for Stack Overflow. I think the best open source models today that I can run on my laptop serve that need. Now that coding agents are a thing my frame of reference has shifted to where I now consider a model that can be that my most common need. And unfortunately open models today cannot do that reliably. They might, like you said, be able to in a year or two, but by then the cloud models will have a new capability that I will come to regard as a basic necessity for doing software development. All that said this looks like a great release and I'm looking forward to playing around with it.
- dakolli 6mo ago[flagged]
- originalvichy 6mo agoTake a walk outside.
- dakolli 6mo agoI do more than most, that's why I'm not saying stuff like "The wait is finally over, just two more iterations"