4 ms·
The smaller models have been creeping upward. They don't make headlines because they aren't leapfrogging the mainline models from the big companies, but they ar
by strangescript 1y ago
The smaller models have been creeping upward. They don't make headlines because they aren't leapfrogging the mainline models from the big companies, but they are all very capable.
I loaded up a random 12B model on ollama the other day and couldn't believe how good it competent it seemed and how fast it was given the machine I was on. A year or so ago, that would have not been the case.
- apples_oranges 1y agoexactly, it seems to validate my assumption from some time ago, that we will mostly use local models for everyday tasks.
- pzo 1y agoyeah especially that this simplifies e.g. doing mobile app for 3rd party developers - not extra cost, no need to setup proxy server, monitoring usage to detect abuse, don't need to make complicated subscription plan per usage. We just need Google or Apple to provide their own equivalent of both: Ollama and OpenRouter so user either use inference for free with local models or BringYourOwnKey and pay themself for tokens/electricity bill. We then just charge smaller fee for renting or buying our cars.
- wg0 1y agoBut who will keep them updated and what incentive they would have? That's I can't imagine. Bit vague.
- ebiester 1y agoEventually? Microsoft and Copilot, and Apple and Siri - even if they have to outsource their model making. It will be a challenge to desktop Linux.
- WorldPeas 1y agoI figure this will take the same shape as package distribution. If you have ever used a linux distribution you’ll always see a couple .edu domains serving you packages. Big tech might be able to have specialized models, but following the linux paradigm, it will likely have more cutting edge but temperamental models from university research
- cruzcampo 1y agoWho keeps open source projects maintained and what incentive do they have?
- jsheard 1y agoMost open source projects don't need the kinds of resources that ML development does. Access to huge GPU clusters is the obvious one, but it's easy to forget that the big players are also using huge amounts of soulcrushing human labor for data acquisition, cleaning, labeling and fine tuning, and begrudgingly paying for data they can't scrape. People coding in their free time won't get very far without that supporting infrastructure. I think ML is more akin to open source hardware, in the sense that even when there are people with the relevent skills willing to donate their time for free, the cost of actually realizing their ideas is still so high that it's rarely feasible to keep up with commercial projects.
- cruzcampo 1y agoThat's a fair point. I think GPU clusters are the big one, the rest sounds like a good fit for volunteer work.
- wg0 1y agoOr sharing GPU compute. Crowd sourcing.
- cruzcampo 1y agoOoooh I can see a Seti@Home setup working
- jsheard 1y agoEasier said than done, training is usually done on "big iron" GPUs which are a cut above any hardware that consumers have lying around, and the clusters run on multi-hundred-gigabit networks. Even if you scaled it down to run on gaming cards, and gathered enough volunteers, the low bandwidth and high latency of the internet would still be a problem.
- jillesvangurp 1y agoIncluding figuring out which more expensive models to use when needed instead of doing that by default. Early LLMs were not great at reasoning and not great at using tools. And also not great at reproducing knowledge. Small models are too small to reliably reproduce knowledge but when trained properly they are decent enough for simple reasoning tasks. Like deciding whether to use a smarter/slower/more expensive model.
- mring33621 1y agostrong agree my employer talks about spending 10s of millions on AI but, even at this early stage, my experiments indicate that the smaller, locally-run models are just fine for a lot of tech and business tasks this approach has definite privacy advantages and likely has cost advantages, vs pay-per-use LLM over API.
- lyu07282 1y agoI spend a lot of time working with smaller models, I often had to split the problem into smaller subtasks to make it give acceptable accuracy. With the big models in the cloud you can often get things working much faster, it seems like a tradeoff in engineering time. What was your experience?
- AustinDev 1y agoNot just local models but bespoke apps. The number of bespoke apps I've created shot up dramatically in the last 6 months. I use one to do my recipes/meal plan every week. I have one that goes through all my email addresses and summarizes everything daily. I just finished an intelligent planner / scheduler for my irrigation system that takes into account weather forecast and soil moisture levels. If something is annoying and there is no commercial solution or open-source solution that has the features I want I just make it now and it's fantastic. I've had friends/family ask to use some of them; I declined. I don't want to do support / feature requests.
- the_pwner224 1y agoAs someone who hasn't used AI for "real" app development (mainly just getting ChatGPT to generate small functions & scripts), do you have any recommendations on what tools or resources I should use to get started with this?
- AustinDev 1y agoCursor/Cline/Windsurf are my recommendations for clients. For models stay away from Sonnet 3.7. I find it just lies to you. I'd rather you a slightly less capable model like Sonnet 3.5 where I know it will just make mistakes that won't compile. I do my planning with a combination of Grok3, and higher power OpenAI models. Once I have plan of what I want to build, I create an implemenation_plan.md with all the steps to build my solution. (Generated by the higher power models) I carefully review this plan and if it looks good, I throw it into agent mode and get to work.
- justlikereddit 1y agoLast time I did that I was also impressed, for a start. Problem was that of a top ten book recommendations only the first 3 existed and the rest was a casually blended hallucination delivered in perfect English without skipping a beat. "You like magic? Try reading the Harlew Porthouse series by JRR Marrow, following the orphan magicians adventures in Hogwesteros" And the further towards the context limit it goes the deeper this descent into creative derivative madness it goes. It's entertaining but limited in usefulness.
- omnimus 1y agoLLMs are not search engines…
- mirekrusin 1y agoExactly, I think all those base models should be weeded out from this nonsense, kardashian-like labyrinths of knowledge complexities that just makes them dumber by taking space and compute time. If you can google out some nonsense news, it should stay there in search engines for retrieval. Models should be good at using search tools, not at trying to replicate their results. They should start from logic, math, programming, physics and so on, similar to how education system is suppose to equip you with. IMHO small models can give this speed advantage (faster to experiment ie. with parallel diverging results, ability to munch through more data etc). Stripped to this bare minimum they can likely be much smaller with impressive results, tunable, allow for huge context etc.
- Philpax 1y agoAn interesting development to look forward to will be hooking them up to search engines. The proprietary models already do this, and the open equivalents are not far behind; the recent Qwen models are not as great at knowledge, but are some of the best at agentic functionality. Exciting times ahead!
- hedgehog 1y agoIf you use something like Open Web UI today the search integration works reasonably well.
- nickip 1y agoWhat model? I have been using api's mostly since ollama was too slow for me.
- estsauver 1y agoQwen3 and some of the smaller gemma's are pretty good and fast. I have a gist with my benchmark #'s here on my m4 pro max (with a whole ton of ram, but most small models will fit on a well spec'ed dev mac.) https://gist.github.com/estsauver/a70c929398479f3166f3d69bcededac3 https://gist.github.com/estsauver/a70c929398479f3166f3d69bce...
- patates 1y agoI really like Gemma 3. Some quantized version of the 27B will be good enough for a lot of things. You can also take some abliterated version[0] with zero (like zero zero) guardrails and make it write you a very interesting crime story without having to deal with the infamous "sorry but I'm a friendly and safe model and cannot do that and also think about the children" response. [0]: https://huggingface.co/mlabonne/gemma-3-12b-it-abliterated https://huggingface.co/mlabonne/gemma-3-12b-it-abliterated
- djmips 1y agoWhich model?