6 ms·
Desert Ant Labs: local, fast models that run on device
- deleted 23d ago[deleted]
- hnx0rqy49u 23d ago[dead]
- mailonce 23d ago[dead]
- sipjca 23d agoat first i got very excited about a new fast transcription model (voz) but turns out its just parakeet v3 with some new inference code which is macOS/iOS specific
- gorgmah 23d agoI was also thinking that this is almost too good to be true
- pveugen 23d agoIt's an ANE optimized version of Parakeet, with our own inference, which enabled us to push performance to about 300x realtime speed on an iPhone 16/17. Our next gen Voz model is trained from scratch and will be at least twice as fast. Android and other platforms will land soon.
- sipjca 23d agoI understand. And inference optimizations are great. It's just that it's not a new model, and that's what I was excited about. I work on cross-platform transcription inference and was hoping for something more than just a re-badged parakeet v3 Look forward to the next model
- Muromec 23d agoNice, now I know what I actually need.
- splittydev 23d agoThis seems to be a common theme with these "European sovereign AI" companies. Mostly built on top of other open source work and slightly adjusted, often to collect grant money (although I'm not sure about that last part in this case).
- yencabulator 22d agoEven worse, they made their additions proprietary: > How much does Voz cost? > Every model is free up to 100k monthly active devices per SDK. Unlimited inference per user. Contact us for custom licenses.
- nullbio 23d agoThis is a cool idea. The most useful one for me would be something that can process pdf files into a json schema. Title and tag generation from a post would also be useful. I'm interested in web app though.
- pveugen 23d agoOCR on steroids. Our Schemer model will soon be available (free form text to structured JSON). Once that lands, we want to jump into image to JSON.
- capevace 23d agoBut Schemer will be text-only, correct? Once image support drops this could be a cool addition to https://struktur.sh https://struktur.sh. Will it be on OpenRouter too or will all inference have to be self-managed? I’m thinking about server-side use cases, where budgets are low, so small models shine. edit: ah I saw mainly iOS for now. But an integration would be possible on macOS then, right?
- pveugen 23d agoSchemer is text only. But we're working on sort like models for images to data too. We aim to make all our models available cross-platform. There's a subset currently only iOS/macOS, because they take some more effort to port over to Android and web. Check out the CLI for easy experimentation: https://github.com/Desert-Ant-Labs/desert-ant-cli#desert-ant-cli https://github.com/Desert-Ant-Labs/desert-ant-cli#desert-ant...
- ironsmoke 23d agoYou should check out IBM's Docling.
- ashenke 23d agoA lot of the models would be useful in a web context, to improve on the CMS we're making for clients. But they look like most of them are iOS only, few have a node package or something other, and all the benchmark are running it on modern iPhones so I doubt it would be that fast on a 20$ VPS.
- pveugen 23d agoJust a few are iOS first (pure practical timing/sequencing). We plan to make all models available cross-platform in the coming weeks.
- library8848 23d agoShiny layer of marketing and proprietary code on top of open models? Voz is Parakeet 0.6B v3 Clear is DeepFilterNet 3 Ear is the language predictor from whisper-tiny ...
- sipjca 23d agoseems to be…
- Muromec 23d agoAI beige theme and obnoxious AI writing signaled as much.
- sudb 23d agoIt is now also reasonably straightforward if you have access to frontier LLMs, a recent-ish mac and a recent-ish iPhone to point them at the job of porting a given model to run on the ANE - it's a reasonably easy task to hill-climb at this point!
- deleted 23d ago[deleted]
- joshuat 23d agobut they're European!
- illright 23d agoI wonder why they only support Apple platforms, citing CoreML. Doesn't Android have a similar framework, ML Kit?
- pveugen 23d agoWe plan to make most our models available for Android and web too. Some are a bit harder to port to the different platforms and will take a bit longer to properly land on Android or web. Mostly sequencing (Voz, Clips, Title). Soon!
- 1dom 23d agoThis is a cool way of approaching AI models. I'm a big fan of local LLMs, local specific models like this should be even more powerful. > Every model is free up to 100k monthly active devices. No tokens, no logins. I dunno about the business model though. Cloud LLM billing makes sense: you're getting another computer to do work with each request, and using their compute via their gateway that they bill you. These local models are like old school software. They're producing the weights, and then giving them to people. If I'm happy with the weights you've given me, and I'm not using your compute for inference and not wanting or needing any updates off you, why should you continue getting money off me and my customers? The whole "but we need to keep it updated for your security" doesn't really work as well for software designed to run fully offline like these local models are. I'm not saying they shouldn't get paid, but I guess I feel a personal sadness that it's less obvious how to successfully monetise such a sincerely useful and beneficial approach towards AI models.
- donavanm 23d ago> why should you continue getting money off me and my customers? Welcome to the concept of fair market value. Less snarkily you're conflating the concepts of price and cost; theyre not the same thing and theyre not the same for you or the seller.
- Muromec 23d agoHow does piracy of digital content and software factor into this concept?
- handfuloflight 23d agoHow is it not obvious and fair that they are asking you to pay them when you see success (defined as >100K MAU)? How more aligned can you and them be besides this?
- Muromec 23d agoWhen I buy I chair I don't pay a share of my income to the furniture shop when I get rich. I would buy another chair at some point too, maybe fancier one (or the same). Because chairs are commodity. I do however pay taxes to the government based on my income because it keeps doing ongoing maintenance on everything. Everyone wants to be paid forever for something they produced once is some kind of a mind virus. Make me a better chair and maybe I will buy it, but don't expect to become a trillion dollar company. It's deeply unfair to everyone who wants to be a trillion dollar company of course. What is really funny to me -- the ones that do make it to collect the rent indefinitely also decrease their own taxes paid to the government AND also decrease the amount of contribution to society by making less and shittier stuff.
- mtlynch 23d agoI love this idea and hope to see more on-device models. How do they make money, though? I tried out their demo for Clear, the audio quality improvement model.[0] I'm not sure if it's just I don't have refined enough an ear or their demo is broken, but the "raw" and "enhanced" versions sounded exactly the same to me. [0] https://desertant.com/models/clear/ https://desertant.com/models/clear/
- sudb 23d agoThe underlying model (DFNet3) is not particularly great but it is very small and fast - imo the best commercially usable denoising model is MossFormer2 (no affiliation - it's just excellent) with one drawback in that it can't remove reverb. Nvidia's RE-USE model can do what MossFormer2 does _and_ can remove reverb, but it is non-commercial licensed.
- myshapeprotocol 23d ago[dead]
- bronlund 23d agoThe website looks amazing.
- pveugen 23d agoThanks!
- wuisce 23d agoHow was it built?
- pveugen 23d agoWe design in Figma. Build a design system out that with components for all our public facing products, demos and marketing assets. The website is built with the components, and we have a set of skills and tools to keep both in sync. We also use skills to sync between HuggingFace, our SDK on GitHub and the website. We write and review most our core copy by hand and use that to help expand into different pages. We prob do a write-up about this on our blog later.
- tuxguy 23d agobrilliant ! Would love to see a deep dive on your component design system & copy process. Here is to more beauty in the world - both physical & digital :)
- markdog12 23d ago> opinionated on-device intelligence > Hate speech triage. On-device moderation that flags hateful, abusive and threatening text What could go wrong here?
- lemome 23d agoI don't think you understood. It means these are specialized models. Their toxic model could be ideal for video game lobbies without investing a ton of money if you're an indie dev This also could be ideal if you want your child to play online to have auto-censorship
- ch_sm 23d agoI don’t know if i want to expose my child to auto-censorship (or online gaming anyway).
- Muromec 23d ago1. Not having human in the loop to review it because humans are expensive. 2. Having human in the loop to review it and subject said human to the worst other humans produce.
- ai_for_everyone 23d ago[flagged]
- nater5000 23d agoI definitely think there's a lot to be done with small models dedicated to specific tasks. I've always thought the REAL value is in having large models be able to easily build small models for custom tasks (which I know is kind of a thing), but perhaps just providing the small models directly is the more accessible approach. >accessible via one SDK for Swift, Kotlin, and JavaScript Lol well let me know when there's a Python SDK and I'll give it a try then. Obviously this isn't a deal breaker if you have a real case, but as someone who is willing to spin something up and try it out if there's a quick "pip install" command, this is getting put back on the shelf for now.
- embedding-shape 23d agoThe founders and website seems to talk a lot about mobile devices and how many are being sold/shipped, that's the "unused compute power" they're targeting. So obviously mobile-first platforms and SDKs seem to be the focus first, hence those languages. Understandable, given their target, but just like you I wish there was an easier way to give it a try on a desktop computer.
- ChickeNES 23d agoSame, just wish I could afford the computer to train my own ngl
- lukevp 23d agoI would love to use Voz and Ear, but I’d need a version that is competitive with other audio transcription LLMs for platform availability - meaning macOS, Windows and Linux, and supporting GPUs if available.
- faangguyindia 23d agoCool! is there a local model for LLM command approval?
- init0 23d agoAwaiting web version...
- viccis 23d agoI wish we could pop a tiny one into my phone so that when I type "Will see you" and swipe the word "later" it chooses that instead of "lasso"
- asamadx 23d ago[flagged]
- agcat 23d agoHonestly this is a really cool idea and i am glad that there are companies being built in this space. This is closest to the vision of what i want to do next.
- PreownedPlaid 23d ago[flagged]
- ricardobeat 23d ago> Ranks a transcript's best non-overlapping moments: each clip gets scored and ranked. Build strong selections or unique editing features to pick the best sentences in video or audio recordings. This is very impressive for a 248MB model. I wonder how good the results are, as an LLM 10x the size is still quite bad at that.
- Dwedit 23d agoTesting out Tongue, it detected "馬鹿外人" as Chinese.
- eproxus 23d agoIt does say this: "Certain by script: these characters belong to only one language, so the model never ran." So I think it's some buggy code in their UI that even prevents the model from running.
- deleted 23d ago[deleted]
- anigbrowl 23d agoI like the concept and the development choices seem sensible even if they're not my favorites (although I think missing Python is a mistake). The text feels very LLM generated though, and I reflexively discount the value of anything presented with this writing style.
- momojo 23d ago> The world ships more than a billion capable phones, tablets, and laptops a year, most with a chip built for exactly this work, paid for and idle most of the day. Run the model there and the economics flip: no per-call cost, no round-trip, and nothing leaves the device. This. I run small models (>50MB) for bio-imaging/biotech applications, it feels like every README implies that you need a discrete GPU to get started. While some do, many, especially the most useful ones, do not. Sure it matters if you're also going to do fine-tuning, but I believe your typical user just wants to detect some nuclei and get some cell-body ratios. The laptop on your desk won't be running Meta's SAM, but it has more than enough compute to crunch 100's of your H&E slides overnight.
- hypercube33 22d agoI'm curious what models you're running and what hardware if you'd be open to sharing that info.
- momojo 22d agoCellpose[0] and Stardist[1] are the two you'll see most heavily run and talked about on the forums[2]. They're classic CNN's, but it just feels like they (and any modern models) are swept up in "Go out and buy a 4090.But these guys will give you plenty of mileage before you have to reach for the bigger ones. [0] https://github.com/mouseland/cellpose https://github.com/mouseland/cellpose [1] https://github.com/stardist/stardist https://github.com/stardist/stardist [2] https://forum.image.sc/ https://forum.image.sc/
- suqingfu 23d ago[flagged]
- shelled 22d agoI recently had to use dictation for a few weeks and I was pleasantly surprised that many of the apps (in use/vogue) did support models on my 2021 16GB M1 Pro mac (many of those even supported connecting to a remote or local model endpoint) and at the same time for any worthwhile STT enhancement the model size was hitting higher I would have comfortably wanted. Even though I don't necessarily need dictation any more I intend to keep a custom fully offline setup and try these models (not sure they support live/streaming STT). If any of you are interested there are apps like https://github.com/altic-dev/FluidVoice https://github.com/altic-dev/FluidVoice (this one's a great app) and this https://sam-pop.github.io/WhisperDictation https://sam-pop.github.io/WhisperDictation. The latter, even though it has just 7 stars right now, seems to be more "intuitive". I just hope they expose a way to "connect" to available models on the machine or remotely)
- hefu_hk 22d ago[flagged]
- MisterMunchkin 22d agoThat’s fun! I’ve been looking for something tiny I could embed in a webapp. Not a full genius model, just something light which could enhance the product without requiring ongoing cost. Most organisations give their users terrible hardware, so anything which requires 32GB of RAM or a MacBook Pro won’t work if it’s a government or large organisation.
- devenquan 22d ago[dead]
- YavenTeam 22d ago[flagged]
- noxxider 22d ago[dead]