4 ms·
I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for
by mavamaarten 8d ago
I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.
I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.
E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.
- apimade 8d agoFor TTS I launched this like this week based on a Reddit thread of recommendations, added a new one to it.. I want to say on Wednesday, but this week has been a blur. Problem I’ve found with similar sites is I can’t run a lot of the models, or the results are beyond stale. But this is only stuff I can run locally, or it’s a cloud model. So.. This is good for right now! https://apimade.com/audio-compare.html https://apimade.com/audio-compare.html
- Systemerror7A69 8d agoMy approach to this problem is to just...not try them all. As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe. So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem. If google/gemma-4-31b works, you don't need to overthink it.
- bmordue 8d agosatisficing instead of optimising
- blensor 8d agoThis 100%
- deleted 8d ago[deleted]
- insightfulornot 8d agoStarting with GLM-5.3 Flash was a pretty decent first try! I started with other models, and ended up settling on this exact one because all the others were either too slow or unreliable for my tasks. Qwen 3.8 didn't do it for me, whatever tweaks I added to my harness. Where I'm getting at is you did start with an incredible model in the first place, which greatly helps sticking to it.
- alfiedotwtf 7d agoGLM-5.3 Flash has been my goto since it came out. Only failed once when it lost context, but I'm assuming that was my fault rather than the model. If models never make it past today's close-to-frontier for the rest of my life, I wouldn't complain.
- epolanski 8d ago> If google/gemma-4-31b works, you don't need to overthink it. Up until recently I had a gemini flash 2.0 api deployed that did summarization and translation of news articles/corporate statements fast and cheap and had no reason to update it. If it works fine, this chase of the latest LLM is bit pointless.
- Lalabadie 8d agoOpenAI, Anthropic and peers deeply fear that everyone would eventually come to that same conclusion. (I think you're right)
- killingtime74 8d agoIt's the same problem as trying to buy a car or choose what clothes to buy. You just have to read about options, try things out.
- sisve 8d agoHave you tried openrouters auto model? Tries to give you the best model based on prompt and price https://openrouter.ai/docs/cookbook/coding-agents/openclaw-integration#using-auto-model-for-cost-optimization https://openrouter.ai/docs/cookbook/coding-agents/openclaw-i...
- ComputerGuru 8d agoThat’s if you have disparate prompts and don’t want to actively select a model. If you’re developing a pipeline, it’s a terrible idea. You want to choose a model, validate it, then stick to it.
- mavamaarten 8d agoI have, but honestly that was exactly what I am not looking for. Sometimes it picked a model for a Dutch email that totally does not support Dutch. Other times it would work fine. It's just a layer of indeterminism I wasn't looking for.
- croon 8d agoI use claude for work, and opencode go at home, but i do dabble with openrouter from time to time, and then i just browse the model catalog and filter/sort by recent popularity, price, context size or whatever matters for the task. Mostly just popularity trends, hoping that there's some wisdom in the crowd.
- deedree 8d agoI use costgoat to be a cheapass https://costgoat.com/compare/llm-api#pricing-guide https://costgoat.com/compare/llm-api#pricing-guide
- bonoboTP 8d agoIt's not like new releases come with fully mapped out capability scores for exactly the aspects that you're interested in. There are benchmarks, but reality is often different. It's simply unknown to humanity how well each model will perform in your own bespoke context unless you just try them. You can read experiences and vibes by others but often they will use them in different ways or have different preferences etc. They generally all try to make them good at everything, it's not like they'd declare "this model is not made for task X".
- edude03 8d agoFor cost you can route the same task via openrouter and see what it costs. Literally a for m on models do task loop and look at the cost.
- ledak 8d ago>I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try. Ask one of the top-tier models to do deep research on it.
- rudicjd2747 8d agoIf you use the Chinese ones at least the energy comes from solar - aside from that, pick one and see if it solves your problems
- yorwba 8d agoIf you use a model hosted in China while it's night there, the energy obviously doesn't come from solar, as there's not enough storage capacity. Even if you use it during daytime, most of the energy still won't come from solar, because the ideal solar power locations in the sparsely-populated west are far from the ideal data center locations in the densely-populated east and there's not enough transmission capacity between them. Additionally, using a Chinese provider doesn't mean the model will be hosted in China, e.g. for Qwen Omni here, the supported regions are: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). https://www.alibabacloud.com/help/en/model-studio/qwen-omni#qwen3-8-omni-flash https://www.alibabacloud.com/help/en/model-studio/qwen-omni#...
- boltzmann-brain 8d agopower capacity coming from solar means that at peak, no new non-solar power plants are needed. it does not mean that at nadir everything is solar. legacy power plants continue working around the clock, it either goes to waste or it goes to things that run at night. that aside, solar also includes storage of energy that comes from solar. you should read up on energy networks.
- idiotsecant 8d agoif you mine 4 tons of coal, anywhere on earth, 1 ton of it will be going to china to be burned to make electricity. China's energy production is extraordinarily dirty, the cleanest part of it is the PR. More than half of the energy they produce is from burning coal. They certainly want to integrate more renewables but they have the same problems with that everyone else does - storage and transmission are expensive and essential for a renewable heavy grid.
- paimapi 8d agoI doubted this but it does look like there's an actual government initiative that mandates 80% clean energy for all new data-center builds: https://www.fastcompany.com/91578780/how-china-is-powering-new-data-centers-with-clean-energy https://www.fastcompany.com/91578780/how-china-is-powering-n... that said, China is also rapidly scaling up coal-fired plants: https://apnews.com/article/china-coal-power-plant-carbon-climate-change-ba86e7584e3afe1826eed5cffa25354a https://apnews.com/article/china-coal-power-plant-carbon-cli... those presumably support all of the surrounding infrastructure + people + manufacturing so it's not as if it's truly solar-powered. but it's still handily better than the state-by-state abandonment of clean energy goals here in the US - I lay this out a bit here: https://news.ycombinator.com/item?id=49700743 https://news.ycombinator.com/item?id=49700743
- testycool 8d agoI use models.dev's CLI tool, which I think gets data from OpenRouter, and ArtificialAnalysis so your coding agent can help you narrow it down. <sidenote> Similarly, HuggingFace has a CLI + a few skills, and they are very useful. I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store. It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either. We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo. </sidenote>
- apples_oranges 8d agoWhy not just ask Google AI or ChatGPT to help with choosing?
- stevenhubertron 8d agoPick the cheapest model with good speed. ZDR, and price. If it works great. If it doesn’t pick one a bit more expensive till you get what you need.
- greazy 7d agoThis is the approach recommended for starting a hobby as well. Buy the cheapest set of tools and if they break or are not meeting your needs.
- bushido 8d agoThe approach I normally take is: I have small benchmarks for myself which test for things I care about. And that has any models that I'm considering through that.
- maxyurk 7d agoI bumped into openrouter's `ori eval`, didn't try it though https://openrouter.ai/ori/eval https://openrouter.ai/ori/eval