7 ms·
Liquid AI reveals 8B-A1B MoE trained on 38T
- adityashankar 4mo agoThis is super interesting, I'm particularly excited for this one as it may allow teams to scale this architecture for VLAs (vision language action models), and having sparser models means more real-time actions on a locally hosted model demo link for anyone that wants to try this out https://playground.liquid.ai/chat?model=cmppnbgse000004l4bc8df3wx https://playground.liquid.ai/chat?model=cmppnbgse000004l4bc8...
- HappMacDonald 4mo agoNeither of the VL models work for me in playground though, they just error out
- HenryMulligan 4mo agoWhy does this not have (day-one) support for Ollama? The previous model is on there? Is it related to the ongoing refactor work or are people abandoning Ollama for other LLM engines?
- gmuslera 4mo agoHomeopathic AI
- nickpsecurity 4mo agoI'd normally call that a low-effort, troll comment. But, thinking on it, you may have a great metaphor. They keep promising great performance out of models whose key ingredient (parameters) they are diluting. Many seem to be in a competition saying they're getting smaller and higher performance at the same time. Then, the homeopathic models don't perform as well as real models when independently tested. Again, spot on.
- elorant 4mo agoWow, this is fucking phenomenal. I fed it a long transcript asking it to create a summary and it executed it extremely well. For an 8B model this is quite impressive.
- SubiculumCode 4mo agoI gave it a 2000 line python code that does some fairly sophisticated geodesic calculations on surfaces, and asked to review the code. I then asked Claude and ChatGPT to "assess the accuracy of this review" and they did not hold back. That said, its a very small model, and very fast.
- ValdikSS 4mo agoBad at translation, at least to Russian. Very fast though, about 2x faster than Gemma 4 e2b on my CPU.
- ramshanker 4mo agoGuess we can run this even on CPU!
- bee_rider 4mo agoThey seem… much better than all the models they compared against? What’s the catch?
- FuckButtons 4mo agoThey only showed the benchmarks where they outperformed?
- andai 4mo agoIt's twice the size?
- mlmonkey 4mo agoQuestion: I have a dirty car and the car wash is just 50 meters away. Should I walk or drive to the carwash? Answer: . . . . So, unless you have a compelling reason not to, walk to the car wash.
- cwnyth 4mo agoI'm surprised these models haven't picked this up yet in the training data. Both Claude and ChatGPT missed that one when I posed the question to them last year.
- tingletech 4mo agoWhy would a model know that one washes cars at a car wash? We don't clean our bodies at the body wash or clean the kitchen at the kitchen wash.
- shepardrtc 4mo agoThere's meaning in the term "car wash" that it understands. But I don't suspect anyone has taught it that for 99.9% of people, going to car wash ONLY means that you're going to wash your car and that it should make that implicit assumption. What if you're the car wash owner? Or a maintenance technician? Pretty easy to just walk over there if you're just 50ft away.
- jjtheblunt 4mo agoto your point, when my Aussie friends first mentioned a "car park" to my north american born self, i wondered _momentarily_ what that was, then realized it's sort of a fun name for what i would call a parking lot.
- nl 4mo agoI've never thought of it as a fun term before. We use "park" as "I will park the car" not park as in "amusement park"
- chabes 4mo agoThe small models are getting really impressive. I recently realized that Qwen3.5:4B is way more capable than I thought a model that size could be. Combine that with the work Liquid puts into RL and fine tuning, and you get models that perform extremely well on minimal hardware. Combine that with your own fine tuning, and you get a specialized tool that is fast, private, and doesn’t require internet connection.
- r0b05 4mo agoWhat did you use qwen3.5 4b for?
- sroussey 4mo agoI find it works well in the browser.
- cjtrowbridge 4mo agoits really good at agentic tasks
- steve_adams_86 4mo agoI use it for triaging my messages and emails and reminding me how all of it ties together. It uses Obsidian to know where to put stuff and how to connect information. It isn't perfect. It's very slow (using a 32GB M2 Max) but fast enough for my needs. A good example of how it's helpful is that it will make certain things relatively frictionless. Like, I need to pay property taxes. I hate this stuff. I got the email reminder from my municipality and it made an entry in my TODOs which points to page with instructions to pay the taxes, including my folio and access numbers for when I log in. That was taken from the email and a document which contains past property tax information. I have it all there, but it compiles relevant data into dedicated TODO pages. I'm so bad at doing all of this myself. I really don't enjoy it. Send me to buy a carrot at the store and I'll happily walk 30 minutes there and back to do it. It isn't the effort so to speak; it's how unrewarding, inefficient, and bureaucratic it all is. I'm allergic to it. Why isn't it baked into my income taxes? Why are we still doing this? Sometimes it does a really bad job of making TODOs. Like my wife messaged me about what our dinner plan was, so Qwen went ahead and made a plan for chicken meatball soup based on messages from a week earlier. It totally fabricated the recipe. Yet, I don't know, it was still helpful to be reminded that I'm in charge of dinner. It's probably best at scaffolding responses to emails I don't want to send. I will write it, but I appreciate basic information being fleshed out so I can write it without jumping around looking for files or numbers or whatever constantly. I use it with a custom harness. It could be a lot better. Everything about it could be better. The model is remarkably good for its size and price, though. Letting Sonnet 4.6 do it instead always yields much better results, much faster, but it's kind of like using a new phone vs a super old one. They can both get you there. The sound quality and camera might be worse, it doesn't look as fancy, but the new one is $1200 and the old one is free on marketplace if you're handy with a screwdriver and a fresh battery. Sounds great to me Worth noting: this was all vibe-coded using Opus 4.6 and 4.7. It's the only project I've built that is strictly vibe-coded. It's simultaneously exciting and disgusting. I'm not sure if I'll ever 'software engineer' it, or I'll just let it be slop. It works.
- SubiculumCode 4mo agoAnybody use their localcowork [1] before? That is where the demo lives. Or not? [1] https://github.com/Liquid4All/cookbook/tree/main/examples/localcowork https://github.com/Liquid4All/cookbook/tree/main/examples/lo...
- Ifkaluva 4mo agoLiquid does amazing work, but I kinda feel like they are overtraining their models. 38T tokens seems like a lot for an 8B model
- andai 4mo agoWhat's the downside? Don't they stop when they hit diminishing returns?
- irthomasthomas 4mo agoWoah, chinchilla scaling is 20 x active_params. I think mistral was 2 x Chinchilla. This is 1800 x
- zmmmmm 4mo agoNo vision support?
- kilroy123 4mo agoHmm, I asked it who made it, and it says Google?
- pure_magic 4mo agoMany such cases. Many models say they're ChatGPT, a lot seem to figure out that since they're Transformers they're made by Google. Doesn't really tell you a lot. Perhaps a pretraining / midtraining artifact.
- jauntywundrkind 4mo agoI really love how fast it is! Their press release comparing it on Strix Halo and M5 Max are impressive. It going twice as fast at GPU benchmarks even more so!
- onlyrealcuzzo 4mo agoI just tested this on a bug fixing benchmark I'm working on. It did not perform as well as I expected. Qwen2.5-Coder-3B (2 years old) outperformed it by a wide range -> fixing ~50% of bugs whereas this model only fixed ~12%. Granted, it's not a coder specific model, but given its benchmark performance to Gemma models, and that it's two years newer, and that it's an MoE with 8B total params, I expected it to be more competitive.
- HanClinto 4mo agoSome of the coding-specific fine-tunes were really impressive boosts. Qwen2.5-3B-Instruct is also available [0] -- if it's not too much to ask, I'd be curious how more general models stack up in your benchmark? [0] - https://huggingface.co/Qwen/Qwen2.5-3B-Instruct https://huggingface.co/Qwen/Qwen2.5-3B-Instruct
- debazel 4mo agoI tried it with OpenCode and it is borderline incapable of using tool calls, so that might be why it is doing so bad on your test.
- peder 4mo agoI just did the same. Absolutely awful. I assume OpenCode's heavy context is a problem, and it's probably better to use Liquid's own OpenCode alternative for this.
- solarkraft 4mo agoWhere can I find that agent harness? A look at their Docs and asking Gemini yielded no results. Edit: Is it this? https://github.com/Liquid4All/cookbook/tree/main/examples/localcowork https://github.com/Liquid4All/cookbook/tree/main/examples/lo... FYI: Opencode is very well tuned for Qwen models, but I haven’t found it that rare for niche models to perform badly in it.
- XCSme 4mo agoI will test it when it's accessible via OpenRouter, but the previous LFM2 model (lfm-2-24b-a2b) didn't do well on my tests, it got only 1/20 questions/tasks right, way below Gemma 31B or Qwen 35b-a3b (those get like 10/20 right)
- frankdlc222 4mo agoLook at the accuracy numbers and these things clearly don't know much yet, and I'm not about to hand one my hardest work. But you can see where it's going. As quantization and the MoE stuff keeps getting better, "good enough to just run on my own machine" keeps eating into more of what I'm currently paying a frontier lab for. Once a local model can handle like 80% of what I need, the math stops making sense for the subscription.
- feelingsonice 4mo agoIs Liquid AI still using the liquid neural network architecture?
- 2001zhaozhao 4mo agoAt some point we have to be running into some inherent mathematical limits of knowledge compression, right? No way the knowledge benchmarks on these 8B models will keep getting better without overfitting on these benchmarks
- yorwba 4mo agoIf you give the model access to specialized tools (e.g. web search for question answering) the knowledge doesn't have to be stored in the model weights, which leaves some room for improvement. You'd still be overfitting to benchmarks (since different tasks might require different tools) but not necessarily to specific benchmark questions, so within-domain generalization could be quite good. As an example for a similar approach, Teapot AI has trained very small models https://teapotai.com/models https://teapotai.com/models to only answer questions where the answer can be found within the context window, and although not perfect, they do quite well at this compared to larger, more general models.
- geek_at 4mo agogood point I have the feeling larger models (20b+) rely too much about their stored knowledge and sometimes fail to use tools because they think they know the answer. smaller specialized tool calling models could be the smart route for the future
- Woodi 4mo agoYea, it's strange all that all possible books stealing movement and then lobbying for law prohibiting... something. Humans train "thinking methodology" first and then know how to use it while accessing data and to build knowledge. Humans do not memorize at once all text in existence, that's totally stupid. Already thinking humans specialize in disciplines: math, chemistry, IT, cooking, etc while still using new data. All of that computing is local- on the LAN of the brain. So if some "agents" wants to help then there is zero need for computation outside of home/corporation/car local area network. Licenses ??
- grigio 4mo agoI tested the previous model from Liquid, unfortunatly big claim but poor real performance
- asb 4mo agoBeware the license. They misleadingly state on the blog post "Open-weight — Download, fine-tune, and deploy without restrictions". But if you read their license <https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICE...> it has significant restrictions for any org with other $10M in revenue.