17 ms·
Ask HN: Is anyone doing anything cool with tiny language models?
I mean anything in the 0.5B-3B range that's available on Ollama (for example). Have you built any cool tooling that uses these models as part of your work flow?
- psyklic 2y agoJetBrains' local single-line autocomplete model is 0.1B (w/ 1536-token context, ~170 lines of code): https://blog.jetbrains.com/blog/2024/04/04/full-line-code-completion-in-jetbrains-ides-all-you-need-to-know/#under-the-hood https://blog.jetbrains.com/blog/2024/04/04/full-line-code-co... For context, GPT-2-small is 0.124B params (w/ 1024-token context).
- WithinReason 2y agoThat size is on the edge of something you can train at home
- vineyardmike 2y agoIf you have modern hardware, you can absolutely train that at home. Or very affordable on a cloud service. I’ve seen a number of “DIY GPT-2” tutorials that target this sweet spot. You won’t get amazing results unless you want to leave a personal computer running for a number of hours/days and you have solid data to train on locally, but fine-tuning should be in the realm of normal hobbyists patience.
- nottorp 2y agoHmm is there anything reasonably ready made* for this spot? Training and querying a llm locally on an existing codebase? * I don't mind compiling it myself but i'd rather not write it.
- Sohcahtoa82 2y agoNot even on the edge. That's something you could train on a 2 GB GPU. The general guidance I've used is that to train a model, you need an amount of RAM (or VRAM) equal to 8x the number of parameters, so a 0.125B model would need 1 GB of RAM to train.
- smaddox 2y agoYou can train that size of a model on ~1 billion tokens in ~3 minutes on a rented 8xH100 80GB node (~$9/hr on Lambda Labs, RunPod io, etc.) using the NanoGPT speed run repo: https://github.com/KellerJordan/modded-nanogpt https://github.com/KellerJordan/modded-nanogpt For that short of a run, you'll spend more time waiting for the node to come up, downloading the dataset, and compiling the model, though.
- pseudosavant 2y agoI wonder how big that model is in RAM/disk. I use LLMs for FFMPEG all the time, and I was thinking about training a model on just the FFMPEG CLI arguments. If it was small enough, it could be a package for FFMPEG. e.g. `ffmpeg llm "Convert this MP4 into the latest royalty-free codecs in an MKV."`
- jedbrooke 2y agothe jetbrains models are about 70MB zipped on disk (one model per language)
- pseudosavant 2y agoThat is easily small enough to host as a static SPA web app. I was first thinking it would be cool to make a static web app that would run the model locally. You'd make a query and it'd give the FFMPEG commands.
- h0l0cube 2y agoPlease submit a blog post to HN when you're done. I'd be curious to know the most minimal LLM setup needed get consistently sane output for FFMPEG parameters.
- maujim 2y agofrom a few days ago: https://news.ycombinator.com/item?id=42706637 https://news.ycombinator.com/item?id=42706637
- binary132 2y agoThat’s a great idea, but I feel like it might be hard to get it to be correct enough
- staticautomatic 2y agoIs that why their tab completion is so bad now?
- sam_lowry_ 2y agoHm... I wonder what your use case it. I do the modern Enterprise Java and the tab completion is a major time saver. While interactive AI is all about posing, meditating on the prompt, then trying to fix the outcome, IntelliJ tab completion... shows what it will complete as you type and you Tab when you are 100% OK with the completion, which surprisingly happens 90..99% of the time for me, depending on the project.
- mettamage 2y agoI simply use it to de-anonymize code that I typed in via Claude Maybe should write a plugin for it (open source): 1. Put in all your work related questions in the plugin, an LLM will make it as an abstract question for you to preview and send it 2. And then get the answer with all the data back E.g. df[“cookie_company_name”] becomes df[“a”] and back
- politelemon 2y agoCould you recommend a tiny language model I could try out locally?
- mettamage 2y agoLlama 3.2 has about 3.2b parameters. I have to admit, I use bigger ones like phi-4 (14.7b) and Llama 3.3 (70.6b) but I think Llama 3.2 could do de-anonimization and anonimization of code
- OxfordOutlander 2y ago+1 this idea. I do the same. Just do it locally using ollama, also using 3.2 3b
- RicoElectrico 2y agoLlama 3.2 punches way above its weight. For general "language manipulation" tasks it's good enough - and it can be used on a CPU with acceptable speed.
- seunosewa 2y agoHow many tokens/s?
- iamnotagenius 2y ago10-15t/s on 12400 with ddr5
- RhysU 2y ago"Comedy Writing With Small Generative Models" by Jamie Brew (Strange Loop 2023) https://m.youtube.com/watch?v=M2o4f_2L0No https://m.youtube.com/watch?v=M2o4f_2L0No Spend the 45 minutes watching this talk. It is a delight. If you are unsure, wait until the speaker picks up the guitar.
- 100k 2y agoSeconded! This was my favorite talk at Strange Loop (including my own).
- prettyblocks 2y agoExcellent share - nice to see people doing cool things with the tech while not taking themselves too seriously.
- iamnotagenius 2y agoNo, but I use llama 3.2 1b and qwen2.5 1.5 as bash oneliner generator, always runnimg in console.
- andai 2y agoCould you elaborate?
- XMasterrrr 2y agoI think I know what he means. I use AI Chat. I load Qwen2.5-1.5B-Instruct with llama.cpp server, fully offloaded to the CPU, and then I config AI Chat to connect to the llama.cpp endpoint. Checkout the demo they have below https://github.com/sigoden/aichat#shell-assistant https://github.com/sigoden/aichat#shell-assistant
- iamnotagenius 2y agoI just run llama-cli with the model. Every time I want some "awk" or "find" trickery, I just ask model. Good for throwaway python scripts too.
- jajko 2y agoCan it do 'sed'? I think one major improvement for folks like me would be human->regex LLM translator, ideally also respecting different flavors/syntax for various languages and tools. This has been a bane of me - I run into requirement to develop some complex regexes maybe every 2-3 years, so I dig deep into specs, work on it, deliver eventually if its even possible, and within few months almost completely forget all the details and start at almost same place next time. It gets better over time but clearly I will retire earlier than this skill settles in well.
- iamnotagenius 2y agohave not tried yet. any specific query? I can try.
- 2y ago
- Havoc 2y agoPretty sure they are mostly used as fine tuning targets, rather than as-is.
- dcl 2y agoBut for what purposes?
- arionhardison 2y agoI am, in a way by using EHR/EMR data for fine tuning so agents can query each other for medical records in a HIPPA compliant manner.
- azhenley 2y agoMicrosoft published a paper on their FLAME model (60M parameters) for Excel formula repair/completion which outperformed much larger models (>100B parameters). https://arxiv.org/abs/2301.13779 https://arxiv.org/abs/2301.13779
- barrenko 2y agoThis is really cool. Is this already in Excel?
- andai 2y agoThis is wild. They claim it was trained exclusively on Excel formulas, but then they mention retrieval? Is it understanding the connection between English and formulas? Or am I misunderstanding retrieval in this context? Edit: No, the retrieval is Formula-Formula, the model (nor I believe tokenizer) does not handle English.
- 3abiton 2y agoBut I feel we're going back full circle. These small models are not generalist, thus not really LLMs at least in terms of objective. Recently there has been a rise of "specialized" models that provide lots of values, but that's not why we were sold on LLMs.
- colechristensen 2y agoBut that's the thing, I don't need my ML model to be able to write me a sonnet about the history of beets, especially if I want to run it at home for specific tasks like as a programming assistant. I'm fine with and prefer specialist models in most cases.
- zeroCalories 2y agoI would love a model that knows SQL really well so I don't need to remember all the small details of the language. Beyond that, I don't see why the transformer architecture can't be applied to any problem that needs to predict sequences.
- behohippy 2y agoI have a mini PC with an n100 CPU connected to a small 7" monitor sitting on my desk, under the regular PC. I have llama 3b (q4) generating endless stories in different genres and styles. It's fun to glance over at it and read whatever it's in the middle of making. I gave llama.cpp one CPU core and it generates slow enough to just read at a normal pace, and the CPU fans don't go nuts. Totally not productive or really useful but I like it.
- bithavoc 2y agothis is so cool, any chance you post a video?
- behohippy 2y agoJust this pic: https://imgur.com/ip8GWIh https://imgur.com/ip8GWIh
- Dansvidania 2y agothis sounds pretty cool, do you have any video/media of it?
- Uehreka 2y agoDo you find that it actually generates varied and diverse stories? Or does it just fall into the same 3 grooves? Last week I tried to get an LLM (one of the recent Llama models running through Groq, it was 70B I believe) to produce randomly generated prompts in a variety of styles and it kept producing cyberpunk scifi stuff. When I told it to stop doing cyberpunk scifi stuff it went completely to wild west.
- eb0la 2y agoWe're using small language models to detect prompt injection. Not too cool, but at least we can publish some AI-related stuff on the internet without a huge bill.
- sitkack 2y agoWhat kind of prompt injection attacks do you filter out? Have you tested with a prompt tuning framework?
- A4ET8a8uTh0_v2 2y agoKinda? All local so very much personal, non-business use. I made Ollama talk in a specific persona styles with the idea of speaking like Spider Jerusalem, when I feel like retaining some level of privacy by avoiding phrases I would normally use. Uncensored llama just rewrites my post with a specific persona's 'voice'. Works amusingly well for that purpose.
- deet 2y agoWe (avy.ai) are using models in that range to analyze computer activity on-device, in a privacy sensitive way, to help knowledge workers as they go about their day. The local models do things ranging from cleaning up OCR, to summarizing meetings, to estimating the user's current goals and activity, to predicting search terms, to predicting queries and actions that, if run, would help the user accomplish their current task. The capabilities of these tiny models have really surged recently. Even small vision models are becoming useful, especially if fine tuned.
- bendews 2y agoIs this along the lines of rewind.ai, MSCopilot, screenpipe, or something else entirely?
- ignoramous 2y agoWe're prototyping a text firewall (for Android) with Gemma2 2B (which limits us to English), though DeepSeek's R1 variants now look pretty promising [0]: Depending on the content, we rewrite the text or quarantine it from your view. Of course this is easy (for English) in the sense that the core logic is all LLMs [1], but the integration points (on Android) are not so straight forward for anything other than SMS. [2] A more difficult problem we forsee is to turn it into a real-time (online) firewall (for calls, for example). [1] https://chat.deepseek.com/a/chat/s/d5aeeda1-fefe-4fc6-8c90-20effc1dd7a4 https://chat.deepseek.com/a/chat/s/d5aeeda1-fefe-4fc6-8c90-2... [1] MediaPipe in particular makes it simple to prototype around Gemma2 on Android: https://ai.google.dev/edge/mediapipe/solutions/genai/llm_inference/android https://ai.google.dev/edge/mediapipe/solutions/genai/llm_inf... [2] Intend to open source it once we get it working for anything other than SMSes
- saurik 2y agoFWIW, that chat.deepseek.com link at least doesn't still work... did it ever (for people not logged into your account)?
- ignoramous 2y agoGuess it never has. Though, the example I shared (reproduced below) is pretty basic. --- Prompt: I don't like expletives, spelling mistakes, toxic texts, personal attacks, negative remarks, snide, bad faith arguments, needless bait and trolling. Help clean the texts for me. Rewrite or reword if you have to without losing meaning. If you do reword the text, mention it. If not, reproduce it as-is. Texts: Mel: "Fuck you you dim twat" Bel: "Wanna meet tomorrow?" Cel: "lol, you're clueless; do you even hear yourself?" Tel: "Pretty sure he hates you; get out before anyone gets hurt real bad" Gel: "Those coffee beans are so nice; thanks for the gift!" Lel: "hahaha, you mother fucker; you nailed it, congrats!" Hel: "OK Goooooogle" --- Output: Here's the cleaned-up version of the texts: Mel: "I’m really upset with you right now." (Reworded to remove expletives and offensive language.) Bel: "Wanna meet tomorrow?" (No changes needed.) Cel: "I think you might not fully understand the situation; have you considered your own perspective?" (Reworded to remove snide remarks and make it more constructive.) Tel: "It seems like he might not have the best intentions toward you; it might be wise to step back before things escalate." (Reworded to remove negativity and make it more neutral and caring.) Gel: "Those coffee beans are so nice; thanks for the gift!" (No changes needed.) Lel: "Hahaha, you did an amazing job! Congrats!" (Reworded to remove expletives while keeping the celebratory tone.) Hel: "OK Google." (Corrected spelling for clarity.)
- mritchie712 2y agoI used local LLMs via Ollama for generating H1's / marketing copy. 1. Create several different personas 2. Generate a ton of variation using a high temperature 3. Compare the variagtions head-to-head using the LLM to get a win / loss ratio The best ones can be quite good. 0 - https://www.definite.app/blog/overkillm https://www.definite.app/blog/overkillm
- UltraSane 2y agoclever name!
- Mashimo 2y agoWhat is an H1?
- TachyonicBytes 2y agoNot the OP, but they are "Headers". Probably coming from the <h1> tag in html. What outsiders probably call "Headlines".
- laristine 2y agoMain heading of an article
- flippyhead 2y agoI have a tiny device that listens to conversations between two people or more and constantly tries to declare a "winner"
- pseudosavant 2y agoI'd love to hear more about the hardware behind this project. I've had concepts for tech requiring a mic on me at all times for various reasons. Always tricky to have enough power in a reasonable DIY form factor.
- oa335 2y agoThis made me actually laugh out loud. Can you share more details on hardware and models used?
- econ 2y agoThis is a product I want
- amelius 2y agoYou can use the model to generate winning speeches also.
- jjcm 2y agoAre you raising a funding round? I'm bought in. This is hilarious.
- hn8726 2y agoWhat approach/stack would you recommend for listening to an ongoing conversation, transcribing it and passing through llm? I had some use cases in mind but I'm not very familiar with AI frameworks and tools
- eddd-ddde 2y agoI love that there's not even a vague idea of the winner "metric" in your explanation. Like it's just, _the_ winner.
- mkaic 2y agoThis reminds me of the antics of streamer DougDoug, who often uses LLM APIs to live-summarize, analyze, or interact with his (often multi-thousand-strong) Twitch chat. Most recently I saw him do a GeoGuessr stream where he had ChatGPT assume the role of a detective who must comb through the thousands of chat messages for clues about where the chat thinks the location is, then synthesizes the clamor into a final guess. Aside from constantly being trolled by people spamming nothing but "Kyoto, Japan" in chat, it occasionaly demonstrated a pretty effective incarnation of "the wisdom of the crowd" and was strikingly accurate at times.
- simonjgreen 2y agoMicro Wake Word is a library and set of on device models for ESPs to wake on a spoken wake word. https://github.com/kahrendt/microWakeWord https://github.com/kahrendt/microWakeWord Recently deployed in Home Assistants fully local capable Alexa replacement. https://www.home-assistant.io/voice_control/about_wake_word/ https://www.home-assistant.io/voice_control/about_wake_word/
- kristopolous 2y agoI'm working on using them for agentic voice commands of a limited scope. My needs are narrow and limited but I want a bit of flexibility.
- ata_aman 2y agoI have it running on a Raspberry Pi 5 for offline chat and RAG. I wrote this open-source code for it: https://github.com/persys-ai/persys https://github.com/persys-ai/persys It also does RAG on apps there, like the music player, contacts app and to-do app. I can ask it to recommend similar artists to listen to based on my music library for example or ask it to quiz me on my PDF papers.
- nejsjsjsbsb 2y agoDoes https://github.com/persys-ai/persys-server https://github.com/persys-ai/persys-server run on the rpi? Is that design 3d printable? Or is that for paid users only.
- ata_aman 2y agoI can publish it no problem. I’ll create a new repo with instructions for the hardware with CAD files. Designing a new one for the NVIDIA Orin Nano Super so it might take a few days.
- nejsjsjsbsb 2y agoUp to you! Totally understand if you want to hold something back for a paid option!
- deivid 2y agoNot sure it qualifies, but I've started building an Android app that wraps bergamot[0] (the firefox translation models) to have on-device translation without reliance on google. Bergamot is already used inside firefox, but I wanted translation also outside the browser. [0]: bergamot https://github.com/browsermt/bergamot-translator https://github.com/browsermt/bergamot-translator
- deivid 2y agoI would be very interested if someone is aware of any small/tiny models to perform OCR, so the app can translate pictures as well
- Eisenstein 2y agoMiniCPM-V 2.6 isn't that small (8b) but it can do this. Here is a demo. * https://i.imgur.com/pAuTeAf.jpeg https://i.imgur.com/pAuTeAf.jpeg Using this script: * https://github.com/jabberjabberjabber/LLMOCR/ https://github.com/jabberjabberjabber/LLMOCR/
- crashabr 2y agoDefinitely interested!
- nozzlegear 2y agoI have a small fish script I use to prompt a model to generate three commit messages based off of my current git diff. I'm still playing around with which model comes up with the best messages, but usually I only use it to give me some ideas when my brain isn't working. All the models accomplish that task pretty well. Here's the script: https://github.com/nozzlegear/dotfiles/blob/master/fish-functions/gen_commit_msg.fish https://github.com/nozzlegear/dotfiles/blob/master/fish-func... And for this change [1] it generated these messages: 1. `fix: change from printf to echo for handling git diff input` 2. `refactor: update codeblock syntax in commit message generator` 3. `style: improve readability by adjusting prompt formatting` [1] https://github.com/nozzlegear/dotfiles/commit/0db65054524d0d2e706cbcf57e8067b878b3358b https://github.com/nozzlegear/dotfiles/commit/0db65054524d0d...
- mentos 2y agoAwesome need to make one for naming variables too haha
- relistan 2y agoInteresting idea. But those say what’s in the commit. The commit diff already tells you that. The best commit messages IMO tell you why you did it and what value was delivered. I think it’s gonna be hard for an LLM to do that since that context lives outside the code. But maybe it would, if you hook it to e.g. a ticketing system and include relevant tickets so it can grab context. For instance, in your first example, why was that change needed? It was a fix, but for what issue? In the second message: why was that a desirable change?
- cwmoore 2y agoI'm playing with the idea of identifying logical fallacies stated by live broadcasters.
- spiritplumber 2y agoThat's fantastic and I'd love to help
- cwmoore 2y agoSo far not much beyond this list of targets to identify https://en.wikipedia.org/wiki/List_of_fallacies https://en.wikipedia.org/wiki/List_of_fallacies
- genewitch 2y agoI have several rhetoric and logic books of the sort you might use for training or whatever, and one of my best friends got a doctorate in a tangential field, and may have materials and insights. We actually just threw a relationship curative app online in 17 hours around Thanksgiving., so they "owe" me, as it were. I'm one of those people that can do anything practical with tech and the like, but I have no imagination for it - so when someone mentions something that I think would be beneficial for my fellow humans I get this immense desire to at least cheer on if not ask to help.
- petesergeant 2y agoI'll be very positively impressed if you make this work; I spend all day every day for work trying to make more capable models perform basic reasoning, and often failing :-P
- JayStavis 2y agoAutomation to identify logical/rhetorical fallacies is a long held dream of mine, would love to follow along with this project if it picks up somehow
- vaylian 2y agoLLMs are notoriously unreliable with mathematics and logic. I wish you the best of luck, because this would nevertheless be an awesome tool to have.
- thetrash 2y agoI programmed my own version of Tic Tac Toe in Godot, using a Llama 3B as the AI opponent. Not for work flow, but figuring out how to beat it is entertaining during moments of boredom.
- spiritplumber 2y agoNumber of players: zero U.S. FIRST STRIKE WINNER: NONE USSR FIRST STRIKE WINNER: NONE NATO / WARSAW PACT WINNER: NONE FAR EAST STRATEGY WINNER: NONE US USSR ESCALATION WINNER: NONE MIDDLE EAST WAR WINNER: NONE USSR CHINA ATTACK WINNER: NONE INDIA PAKISTAN WAR WINNER: NONE MEDITERRANEAN WAR WINNER: NONE HONGKONG VARIANT WINNER: NONE Strange game. The only winning move is not to play
- 11944351 2y ago[dead]
- Evidlo 2y agoI have ollama responding to SMS spam texts. I told it to feign interest in whatever the spammer is selling/buying. Each number gets its own persona, like a millennial gymbro or 19th century British gentleman. http://files.widloski.com/image10%20(1).png http://files.widloski.com/image10%20(1).png http://files.widloski.com/image11.png http://files.widloski.com/image11.png
- RVuRnvbM2e 2y agoThis is fantastic. How have your hooked up a mobile number to the llm?
- Evidlo 2y agoAndroid app that forwards to a Python service on remote workstation over MQTT. I can make a Show HN if people are interested.
- deadbabe 2y agoI’d love to see that. Could you simulate iMessage?
- Evidlo 2y agoIf you mean hook this into iMessage, I don't know. I'm willing to bet it's way harder though because Apple
- dambi0 2y agoIf you are willing to use Apple Shortcuts on iOS it’s pretty easy to add something that will be trigged when a message is received and can call out to a service or even use SSH to do something with the contents, including replying
- great_psy 2y agoYes it’s possible, but it’s not something you can easily scale. I had a similar project a few years back that used OSX automations and Shortcuts and Python to send a message everyday to a friend. It required you to be signed in to iMessage on your MacBook. Than was a send operation, the reading of replies is not something I implemented, but I know there is a file somewhere that holds a history of your recent iMessages. So you would have to parse it on file update and that should give you the read operation so you can have a conversation. Very doable in a few hours unless something dramatic changed with how the messages apps works within the last few years.
- antonok 2y agoI've been using Llama models to identify cookie notices on websites, for the purpose of adding filter rules to block them in EasyList Cookie. Otherwise, this is normally done by, essentially, manual volunteer reporting. Most cookie notices turn out to be pretty similar, HTML/CSS-wise, and then you can grab their `innerText` and filter out false positives with a small LLM. I've found the 3B models have decent performance on this task, given enough prompt engineering. They do fall apart slightly around edge cases like less common languages or combined cookie notice + age restriction banners. 7B has a negligible false-positive rate without much extra cost. Either way these things are really fast and it's amazing to see reports streaming in during a crawl with no human effort required. Code is at https://github.com/brave/cookiemonster https://github.com/brave/cookiemonster. You can see the prompt at https://github.com/brave/cookiemonster/blob/main/src/text-classification.mjs#L12 https://github.com/brave/cookiemonster/blob/main/src/text-cl....
- binarysneaker 2y agoMaybe it could also send automated petitions to the EU to undo cookie consent legislation, and reverse some of the enshitification.
- antonok 2y agoHa, I'm not sure the EU is prepared to handle the deluge of petitions that would ensue. On a more serious note, this must be the first time we can quantitatively measure the impact of cookie consent legislation across the web, so maybe there's something to be explored there.
- pk-protect-ai 2y agowhy don't you spam the companies who want your data instead? The sites can simply stop gathering your data, then they will not require to ask for consent ...
- frail_figure 2y ago
- manbitesdog 2y agoI'm making an agent that takes decompiled code and tries to understand the methods and replace variables and function names one at a time.
- krystofee 2y agoThis sounds cool! Are you planningto opensource it?
- DonHopkins 2y agoNo need to: he can just publish a binary then you can run it on itself. ;)
- jmward01 2y agoI think I am. At least I think I'm building things that will enable much smaller models: https://github.com/jmward01/lmplay/wiki/Sacrificial-Training https://github.com/jmward01/lmplay/wiki/Sacrificial-Training
- danbmil99 2y agoUsing llama 3.2 as an interface to a robot. If you can get the latency down, it works wonderfully
- mentos 2y agoWould love to see this applied to a FPS bot in unreal engine.
- spiritplumber 2y agoMy husband and me made a stock market analysis thing that gets it right about 55% of the time, so better than a coin toss. The problem is that it keeps making unethical suggestions, so we're not using it to trade stock. Does anyone have any idea what we can do with that?
- bobbygoodlatte 2y agoI'm curious what sort of unethical suggestions it's coming up with haha
- spiritplumber 2y agoso far, mostly buying companies owned/ran by horrible people.
- Etheryte 2y agoHave you backtested this in times when markets were not constantly green? Nearly any strategy is good in the good times.
- spiritplumber 2y agoyep. the 55% is over a few years.
- kortilla 2y agoRight, but if 55% is avg over the last few years, “buy stock” is going to be correct more than not. https://www.crestmontresearch.com/docs/Stock-Yo-Yo.pdf https://www.crestmontresearch.com/docs/Stock-Yo-Yo.pdf
- jothflee 2y agowhen i feel like casually listening to something, instead of netflix/hulu/whatever, i'll run a ~3b model (qwen 2.5 or llama 3.2) and generate and audio stream of water cooler office gossip. (when it is up, it runs here: https://water-cooler.jothflee.com https://water-cooler.jothflee.com). some of the situations get pretty wild, for the office :)
- kianN 2y agoI don’t know if this counts as tiny but I use llama 3B in prod for summarization (kinda). Its effective context window is pretty small but I have a much more robust statistical model that handles thematic extraction. The llm is essentially just rewriting ~5-10 sentences into a single paragraph. I’ve found the less you need the language model to actually do, the less the size/quality of the model actually matters.
- codazoda 2y agoI had an LLM create a playlist for me. I’m tired of the bad playlists I get from algorithms, so I made a specific playlist with an Llama2 based on several songs I like. I started with 50, removed any I didn’t like, and added more to fill in the spaces. The small models were pretty good at this. Now I have a decent fixed playlist. It does get “tired” after a few weeks and I need to add more to it. I’ve never been able to do this myself with more than a dozen songs.
- petesergeant 2y agoInteresting! I've sadly found more capable models to really fail on music recommendations for me.
- Mashimo 2y agoHuh, interesting. For me that often dreamed up artist and songs.
- jamesponddotco 2y agoInteresting! I wrote a prompt for something similar[1], but I use Claude Sonnet for it. I wonder how a small model would handle it. Time to test, I guess. [1]: https://git.sr.ht/~jamesponddotco/llm-prompts/tree/trunk/data/playlist-generator.md https://git.sr.ht/~jamesponddotco/llm-prompts/tree/trunk/dat...
- codazoda 2y agoThis prompt is a lot more complex than what I did. I don’t recall my exact prompt but it was something like, “Generate a list of 25 songs that I may like if I like Girl is on my Mind by the Black Keys.”
- DonHopkins 2y agoHow about having an LLM create a praylist for you? Then you could implement Salvation as a Service, where you privately confess your sins to a local LLM, and it continuously prays for your eternal soul, recommends penances, and even recites Hail Marys for you.
- JLCarveth 2y agoI used a small (3b, I think) model plus tesseract.js to perform OCR on an image of a nutritional facts table and output structured JSON.
- deivid 2y agoWhat was the model? What kind of performance did you get out of it? Could you share a link to your project, if it is public?
- JLCarveth 2y agohttps://github.com/JLCarveth/nutrition-llama https://github.com/JLCarveth/nutrition-llama I've had good speed / reliability with TheBloke/rocket-3B-GGUF on Huggingface, the Q2_K model. I'm sure there are better models out there now, though. It takes ~8-10 seconds to process an image on my M2 Macbook, so not quite quick enough to run on phones yet, but the accuracy of the output has been quite good.
- tigrank 2y agoAll that server side or client?
- ian_zcy 2y agowhat are you feed into the model? Image (like product packaging) or Image of Structured Table? I found out that model performs good in general with sturctured table, but fails a lot over images.
- itskarad 2y agoI'm using ollama for parsing and categorizing scraped jobs for a local job board dashboard I check everyday.
- jwitthuhn 2y agoI've made a tiny ~1m parameter model that can generate random Magic the Gathering cards that is largely based on Karpathy's nanogpt with a few more features added on top. I don't have a pre-trained model to share but you can make one yourself from the git repo, assuming you have an apple silicon mac. https://github.com/jlwitthuhn/TCGGPT https://github.com/jlwitthuhn/TCGGPT
- HexDecOctBin 2y agoIs there any experiments in a small models that does paraphrasing? I tried hsing some off-the-shelf models, but it didn't go well. I was thinking of hooking them in RPGs with text-based dialogue, so that a character will say something slightly different every time you speak to them.
- krystofee 2y agoIntuitively this sounds like something that should be possible using almost any llm. This should be just a matter of prompting.
- jftuga 2y agoI'm using ollama, llama3.2 3b, and python to shorten news article titles to 10 words or less. I have a 3 column web site with a list of news articles in the middle column. Some of the titles are too long for this format, but the shorter titles appear OK.
- linsomniac 2y agoI have this idea that a tiny LM would be good at canonicalizing entered real estate addresses. We currently buy a data set and software from Experian, but it feels like something an LM might be very good at. There are lots of weirdnesses in address entry that regexes have a hard time with. We know the bulk of addresses a user might be entering, unless it's a totally new property, so we should be able to train it on that.
- thesz 2y agoFrom my experience (2018), run LLM output through beam search over different choices of canonicalization of certain part of text. Even 3-gram models (yeah, 2018) fare better this way.
- sidravi1 2y agoWe fine-tuned a Gemma 2B to identify urgent messages sent by new and expecting mothers on a government-run maternal health helpline. https://idinsight.github.io/tech-blog/blog/enhancing_maternal_healthcare/ https://idinsight.github.io/tech-blog/blog/enhancing_materna...
- 4mitkumar 2y agoSuch a fun thread but this is the kind of applications that perk up my attention! Very cool!
- Mashimo 2y agoOh that is a nice writeup. We have something similar in mind at work. Will forward it.
- Mumps 2y agolovely application! Genuine question: why not use (Modern)BERT instead for classification? (Is the json-output explanation so critical?)
- Mukina 2y agoSuper cool. What a simple and powerful way to help mothers in need. Thanks for sharing.
- dh1011 2y agoI copied all the text from this post and used an LLM to generate a list of all the ideas. I do the same for other similar HN post .
- lordswork 2y agowell, what are the ideas?
- whalesalad 2y agochatgpt did a stellar job parsing the "books on hard things" thread from a little while ago. my prompt was: Can you identify all the books here, sorted by a weight which is determined based on a combo of the number of votes the comment has, the number of sub-comments, or the number of repeat mentions. Ideally retain hyperlinks if possible.
- swifthesitation 2y agocould you link the HN thread?
- whalesalad 2y agogoogle "hn books on hard things" - https://news.ycombinator.com/item?id=42614722 https://news.ycombinator.com/item?id=42614722
- gpm 2y agoI made a shell alias to translate things from French to English, does that count? function trans llm "Translate \"$argv\" from French to English please" end Llama 3.2:3b is a fine French-English dictionary IMHO.
- kreyenborgi 2y agoIs it better than translatelocally? https://translatelocally.com/downloads/ https://translatelocally.com/downloads/ (the same as used in firefox)
- gpm 2y agoIt's different. It doesn't always just give one translation but different options. I can do things like give it a phrase and then ask it to break it down. Or give it a word and if its translation doesn't make sense to me ask how it works in the context of a phrase. llm -c, which continues the previous conversation, is specifically useful for that sort of manipulation. It's also available from the command line, which I find convenient because I basically always have one open.
- guywithahat 2y agoI've been working on a self-hosted, low-latency service for small LLM's. It's basically exactly what I would have wanted when I started my previous startup. The goal is for real time applications, where even the network time to access a fast LLM like groq is an issue. I haven't benchmarked it yet but I'd be happy to hear opinions on it. It's written in C++ (specifically not python), and is designed to be a self-contained microservice based around llama.cpp. https://github.com/thansen0/fastllm.cpp https://github.com/thansen0/fastllm.cpp
- sauravpanda 2y agoWe are building a framework to run this tiny language model in the web so anyone can access private LLMs in their browser: https://github.com/sauravpanda/BrowserAI https://github.com/sauravpanda/BrowserAI. With just three lines of code, you can run Small LLM models inside the browser. We feel this unlocks a ton of potential for businesses so that they can introduce AI without fear of cost and can personalize the experience using AI. Would love your thoughts and what we can do more or better!
- ms7892 2y agoSounds cool. Anyway I can help.
- merwijas 2y agoI put llama 3 on a RBPi 5 and have it running a small droid. I added a TTS engine so it can hear spoken prompts which it replies to in droid speak. It also has a small screen that translates the response to English. I gave it a backstory about being a astromech droid so it usually just talks about the hyperdrive but it's fun.
- kaspermarstal 2y agoI built an Excel Add-In that allows my girlfriend to quickly filter 7000 paper titles and abstracts for a review paper that she is writing [1]. It uses Gemma 2 2b which is a wonderful little model that can run on her laptop CPU. It works surprisingly well for this kind of binary classification task. The nice thing is that she can copy/paste the titles and abstracts in to two columns and write e.g. "=PROMPT(A1:B1, "If the paper studies diabetic neuropathy and stroke, return 'Include', otherwise return 'Exclude'")" and then drag down the formula across 7000 rows to bulk process the data on her own because it's just Excel. There is a gif on the readme on the Github repo that shows it. [1] https://github.com/getcellm/cellm https://github.com/getcellm/cellm
- relistan 2y agoVery cool idea. I’ve used gemma2 2b for a few small things. Very good model for being so small.
- lizamomo96 2y ago[dead]
- afro88 2y agoHow accurate are the classifications?
- kaspermarstal 2y agoI don't know. This paper [1] reports accuracies in the 97-98% range on a similar task with more powerful models. With Gemma 2 2b the accuracy will certainly be lower. [1] https://www.medrxiv.org/content/10.1101/2024.10.01.24314702v1 https://www.medrxiv.org/content/10.1101/2024.10.01.24314702v...
- indolering 2y agoY'all definitely need to cross validate a small number of samples by hand. When I did this kind of research, I would hand validate to at least P < .01.
- evacchi 2y agoI'm interested in finding tiny models to create workflows stringing together several function/tools and running them on device using mcp.run servlets on Android (disclaimer: I work on that)
- lizamomo96 2y ago[dead]
- computers3333 2y agohttps://gophersignal.com https://gophersignal.com – I built GopherSignal! It's a lightweight tool that summarizes Hacker News articles. For example, here’s what it outputs for this very post, "Ask HN: Is anyone doing anything cool with tiny language models?": "A user inquires about the use of tiny language models for interesting applications, such as spam filtering and cookie notice detection. A developer shares their experience with using Ollama to respond to SMS spam with unique personas, like a millennial gymbro or a 19th-century British gentleman. Another user highlights the effectiveness of 3B and 7B language models for cookie notice detection, with decent performance achieved through prompt engineering." I originally used LLaMA 3:Instruct for the backend, which performs much better, but recently started experimenting with the smaller LLaMA 3.2:1B model. It’s been cool seeing other people’s ideas too. Curious—does anyone have suggestions for small models that are good for summaries? Feel free to check it out or make changes: https://github.com/k-zehnder/gophersignal https://github.com/k-zehnder/gophersignal
- tinco 2y agoThat's cool, I really like it. One piece of feedback: I am usually more interested in the HN comments than in the original article. If you'd include a link to the comments then I might switch to GopherSignal as a replacement for the HN frontpage. My flow is generally: Look at the title and the amount of upvotes to decide if I'm interested in the article. Then view the comments to see if there's interesting discussion going on or if there's already someone adding essential context. Only then I'll decide if I want to read the article or not. Of course no big deal if you're not interested in my patronage, just wanted to let you know your page already looks good enough for me to consider switching my most visited page to it if it weren't for this small detail. And maybe the upvote count.
- computers3333 2y agoHey, thanks a ton for the feedback! That was super helpful to hear about your flow—makes a lot of sense and it's pretty similar to how I browse HN too. I usually only dive into the article after checking out the upvotes and seeing what context the comments add. I'll definitely add a link to the comments and the upvote count—gotta keep my tiny but mighty userbase (my mom, me, and hopefully you soon) happy, right? lol And if there's even a chance you'd use GopherSignal as your daily driver, that's a no-brainer for me. Really appreciate you taking the time to share your ideas and help me improve.
- ceritium 2y agoI am doing nothing, but I was wondering if it would make sense to combine a small LLM and SQLITE to parse date time human expressions. For example, given a human input like "last day of this month", the LLM will generate the following query `SELECT date('now','start of month','+1 month','-1 day');` It is probably super overengineering, considering that pretty good libraries are already doing that on different languages, but it would be funny. I did some tests with chatGPT, and it worked sometimes. It would probably work with some fine-tuning, but I don't have the experience or the time right now.
- lionkor 2y agoLLMs tend to REALLY get this wrong. Ask it to generate a query to sum up likes on items uploaded in the last week, defined as the last monday-sunday week (not the last 7 days), and watch it get it subtly wrong almost every time.
- TachyonicBytes 2y agoWhat libraries have you seen that do this?
- vikb 2y ago>>It is probably super overengineering, considering that pretty good libraries are already doing that on different languages, but it would be funny. I did some tests with chatGPT, and it worked sometimes. It would probably work with some fine-tuning, but I don't have the experience or the time right now yeah, could you share those libraries please? Anyone around have actually succeeding in solving this in a way it works? Would appreciate any hints.
- numba888 2y agoMany interesting projects, cool. I'm waiting to LLMs in games. That would make them much more fun. Any time now...
- jaggs 2y agoHave you seen a AI People? https://www.aipeoplegame.com/ https://www.aipeoplegame.com/
- kolinko 2y agoApple’s on device models are around 3B if I’m nit mistaken, and they developed some nice tech around them that they published, if I’m not mistaken - where they have just one model, but have switchable finetunings of that model so that it can perform different functionalities depending on context.
- tomholandpick 2y ago[dead]
- reeeeee 2y agoI built a platform to monitor LLMs that are given complete freedom in the form of a Docker container bash REPL. Currently the models have been offline for some time because I'm upgrading from a single DELL to a TinyMiniMicro Proxmox cluster to run multiple small LLMs locally. The bots don't do a lot of interesting stuff though, I plan to add the following functionalities: - Instead of just resetting every 100 messages, I'm going to provide them with a rolling window of context. - Instead of only allowing BASH commands, they will be able to also respond with reasoning messages, hopefully to make them a bit smarter. - Give them a better docker container with more CLI tools such as curl and a working package manager. If you're interested in seeing the developments, you can subscribe on the platform! https://lama.garden https://lama.garden
- lormayna 2y agoI am using smollm2 to extract some useful information (like remote, language, role, location, etc.) from "Who is hiring" monthly thread and create an RSS feed with specific filter. Still not ready for Show HN, but working.
- krystofee 2y agoHas anyone ever tried to do some automatic email workflow autoresponder agents? Lets say, I want some outcome and it will autonomousl handle the process prompt me and the other side for additional requirements if necessary and then based on that handle the process and reach the outcome?
- ahrjay 2y agoI built https://ffprompt.ryanseddon.com https://ffprompt.ryanseddon.com using the chrome ai (Gemini nano). Allows you to do ffmpeg operations on videos using natural language all client side.
- fauigerzigerk 2y agoWhat are the prerequisites for this? I keep getting "Bummer, looks like your device doesn't support Chrome AI" on macOS 15.2 Chrome 132.0.6834.84 (Official Build) (arm64) [Edit] Found it. I had to enable chrome://flags/#prompt-api-for-gemini-nano
- ahrjay 2y agoYeah the instructions are not clear. They're on the github repo[1] linked in the header. 1. Install Chrome Dev: Ensure you have version 127. [Download Chrome Dev](https://google.com/chrome/dev/ https://google.com/chrome/dev/). 2. Check that you’re on 127.0.6512.0 or above 3. Enable two flags: chrome://flags/#optimization-guide-on-device-model - BypassPerfRequirement chrome://flags/#prompt-api-for-gemini-nano - Enabled 4. Relaunch Chrome 5. Navigate to chrome://components 6. Check that Optimization Guide On Device Model is downloading or force download if not Might take a few minutes for this component to even appear 7. Open dev tools and type (await ai.languageModel.capabilities()).available, should return "readily" when all good [1]: https://github.com/ryanseddon/FFprompt https://github.com/ryanseddon/FFprompt
- addandsubtract 2y agoI use a small model to rename my Linux ISOs. I gave it a custom prompt with examples of how I want the output filenames to be structured and then just feed it files to rename. The output only works 90ish percent of the time, so I wrote a little CLI to iterate through the files and accept / retry / edit the changes the LLM outputs.
- mogaal 2y agoI bought a tiny business in Brazil, the database (Excel) I inherited with previous customer data *do not include gender*. I need gender to start my marketing campaigns and learn more about my future customer. I used Gemma-2B and Python to determine gender based on the data and it worked perfect
- Nashooo 2y agoHow did you verify it worked?
- jbentley1 2y agoTiny language models can do a lot if they are fine tuned for a specific task, but IMO a few things are holding them back: 1. Getting the speed gains is hard unless you are able to pay for dedicated GPUs. Some services offer LoRA as serverless but you don't get the same performance for various technical reasons. 2. Lack of talent to actually do the finetuning. Regular engineers can do a lot of LLM implementation, but when it comes to actually performing training it is a scarcer skillset. Most small to medium orgs don't have people who can do it well. 3. Distribution. Sharing finetunes is hard. HuggingFace exists, but discoverability is an issue. It is flooded with random models with no documentation and it isn't easy to find a good oen for your task. Plus, with a good finetune you also need the prompt and possibly parsing code to make it work the way it is intended and the bundling hasn't been worked out well.
- grisaitis 2y agowhen you say fine-tuning skills or talent are scarce, do you have specific skills in mind? perhaps engineering for training models (eg making model parallelism work)? or the more ML type skills of designing experiments, choosing which methods to use, figuring out datasets for training, hyperparam tuning/evaluation, etc?
- jbentley1 2y agoThe technical parts are less common and specialized, like understanding the hyperparameters and all that, but I don't think that is the main problem. Most people don't understand how to build a good dataset or how to evaluate their finetune after training. Some parts of this are solid rules like always use a separate validation set, but the task dependent parts are harder to teach. It's a different problem every time.
- menaerus 2y agoFinetuning, as I understand it, is mostly laborious and mostly very boring and exhausting work that is not appealing to many engineers. It can be done by people who have some skills in Python or similar language and who have some background in statistics. OTOH to build the infra for LLMs there's much more stuff involved and it's really hard to find engineers who have the capacity to be both the researchers and developers at the same time. By "researchers" I mean that they have to have a capacity to be able to read through the numerous academic and industry papers, comprehend the tiniest details, and materialize it into the product through the code. I think that's much harder and scarcer skill to find. That said, I am not undermining the fine-tuning skill, it's a humongous effort, but I think it's not necessarily the skillset problem.
- bashbjorn 2y agoI'm working on a plugin[1] that runs local LLMs from the Godot game engine. The optimal model sizes seem to be 2B-7B ish, since those will run fast enough on most computers. We recommend that people try it out with Gemma 2 2B (but it will work with any model that works with llama.cpp) At those sizes, it's great for generating non-repetitive flavortext for NPCs. No more "I took an arrow to the knee". Models at around the 2B size aren't really capable enough to act a competent adversary - but they are great for something like bargaining with a shopkeeper, or some other role where natural language can let players do a bit more immersive roleplay. [1] https://github.com/nobodywho-ooo/nobodywho https://github.com/nobodywho-ooo/nobodywho
- Tepix 2y agoCool. Are you aware of good games that use LLMs like this?
- aDyslecticCrow 2y agoI have not seen much myself, but it's one of the earliest use cases I thought about when they started showing up.
- Thews 2y agoBefore ollama and the others could do structured JSON output, I hacked together my own loop to correct the output. I used it that for dummy API endpoints to pretend to be online services but available locally, to pair with UI mockups. For my first test I made a recipe generator and then tried to see what it would take to "jailbreak" it. I also used uncensored models to allow it to generate all kinds of funny content. I think the content you can get from the SLMs for fake data is a lot more engaging than say the ruby ffaker library.
- accrual 2y agoAlthough there are better ways to test, I used a 3B model to speed up replies from my local AI server when testing out an application I was developing. Yes I could have mocked up HTTP replies etc., but in this case the small model let me just plug in and go.
- panchicore3 2y agoI am moderating a playlists manager to restrict them to a range of genders so it classifies song requests as accepted/rejected.
- sebazzz 2y agoI built auto-summarization and grouping in an experimental branch of my hobby-retrospective tool: https://github.com/Sebazzz/Return/tree/experiment/ai-integration https://github.com/Sebazzz/Return/tree/experiment/ai-integra... I’m now just wondering if there is any way to build tests on the input+output of the LLM :D
- herol3oy 2y agoI've created Austen [0] to generate relationships between book characters using Mermaid. [0] https://github.com/herol3oy/austen https://github.com/herol3oy/austen
- lightning19 2y agonot sure if it is cool but, purely out of spite, I'm building a LLM summarizer app to compete with a AI startup that I interviewed with. The founders were super egotistical and initially thought I was not worthy of an interview.
- lizamomo96 2y ago[dead]
- mrmage 2y agoI am building GitHub-Copilot style AI autocomplete in any text field on your Mac. The point is to have the AI fill in all the redundant words required by human language, while you provide the entropy (i.e. the words that are unique to what you are trying to express). It is kind of a "dance" between accepting the AI's suggested words and typing yourself to keep it going in the right direction. Using it, I find myself often writing only the first half of most words, because the second part can usually already be guessed by the AI. In fact, it has a dedicated shortcut for accepting only the first word of the suggestion — that way, it can save you some typing even when later words deviate from your original intent. Completions are generated in real-time locally on your Mac using a variety of models (primarily Qwen 2.5 1.5B). It is currently in open beta: https://cotypist.app https://cotypist.app
- smcleod 2y agoCan vouch for cotypist - it's great!
- sharnabeel 2y agoI have tired but chinese to english but it isn't good(none of them are), because for Chinese words meaning different depending on context so i am just stuck with large model but sometimes even they leave chinese text in translation(like google gemina 2), I really hope there would be some amazing models this year for translation.
- ittaboba 2y agoI am building a private text editor that runs LLMs locally https://manzoni.app/ https://manzoni.app/