5 ms·
I recently upgraded a large portion of my pipeline from gpt-4.1-mini to gpt-5-mini. The performance was horrible - after some research I decided to move everyth
by barrell 1y ago
I recently upgraded a large portion of my pipeline from gpt-4.1-mini to gpt-5-mini. The performance was horrible - after some research I decided to move everything to mistral-medium-0525.
Same price, but dramatically better results, way more reliable, and 10x faster. The only downside is when it does fail, it seems to fail much harder. Where gpt-5-mini would disregard the formatting in the prompt 70% of the time, mistral-medium follows it 99% of the time, but the other 1% of the time inserts random characters (for whatever reason, normally backticks... which then causes it's own formatting issues).
Still, very happy with Mistral so far!
- brcmthrowaway 1y agoWhat are you actually making
- barrell 1y agohttps://phrasing.app https://phrasing.app I’m making an app to learn multiple languages. This portion of the pipeline is about explaining everything I can determine about a work in a sentence in specifically formatted prose. Example: https://x.com/barrelltech/status/1963684443006066772?s=46&t=aEq34ARsgqYIBlyTDH7Tkg https://x.com/barrelltech/status/1963684443006066772?s=46&t=...
- dotancohen 1y agoThis looks great. I invest over an hour a day on Anki, so I'm probably your target audience. I'd love to try Phrasing, but there is no way that I'm going to give my credit card information if I've never seen how it works. I'm not asking for special treatment. I think the 4 Euro 14 day trial should be free. Not because I don't want to pay, but because I don't want my credit card details shared with every service I ever try. Yes, I understand that your credit card provider is safe and that is not my worry. It's just that handing over that information is a very large hurdle, and one that users get over only after you built trust with them. If we can never see what we're getting, we have no reason to get over that hurdle and develop that trust.
- barrell 1y agoThe onboarding flow is already updated with better examples. It’s been my main task this past week — it still needs quite a bit of QA, but it’s been better than the previous flow, so I decided to ship it early. Any bugs will get ironed out in the next day or two. There’s now pretty clear examples of everything My main task this week (besides QAing the new setup flow) is updating examples on the home page and in the blog. It’s all a one person show, so things move a little slowly, but I’m working as fast as I can :) The paid trial gets you credits at cost — Phrasing is bootstrapped by me, and there’s no way I can afford a $4 loss leader for everyone who signs up. I would have to 100x the subscription prices (or more) in order to afford that. There is a way to use the new setup process to “try” the application without spending money, but that’s not an avenue I want to intentionally funnel users through, so you won’t see any copy around it. Feel free to get in touch via the app though if you need more details. I have plans for a free trial eventually, but its blocked by a few key product developments, so it will be a year at least before that’s an option!
- thijsverreck 1y agoAny chance at fixing it with regex parsing or redoing inference when the results are below a certain treshold?
- barrell 1y agoIt’s user facing, so will just have an option for users to regenerate the explanation. It happens rarely enough that it’s not a huge issue, and normally doesn’t effect content (I think once I saw it go a little wonky and end the sentence with a few random words). Just sometimes switches to mono space font in the middle of a paragraph, or it “spells” a word wrong (spell is in quotes because it will spell `chien` as `chi§en`). It’s pretty rare though. Really solid model, just a few quirks
- mark_l_watson 1y agoIt is such a common pattern for LLMs to surround generated JSON with ```json … ``` that I check for this at the application level and fix it. Ten years ago I would do the same sort of sanity checks on formatting when I used LSTMs to generate synthetic data.
- Alifatisk 1y agoI think this is the first time I stumped upon someone who actually mentions LSTM in a practical way instead of just theory. Cool! Would you like to elaborate further on how the experience was with it? What was your approach for using it? How did you generate synthetic data? How did it perform?
- p1esk 1y ago10 years ago I used LSTMs for music generation. Worked pretty well for short MIDI snippets (30-60 seconds).
- barrell 1y agoYeah, that’s infuriating. They’re getting better now with structured data, but it’s going to be a never ending battle getting reliable data structures from an LLM. This is maybe more maybe less insidious. It will literally just insert a random character into the middle of a word. I work with an app that supports 120+ languages though. I give the LLM translations, transliterations, grammar features etc and ask it to explain it in plain English. So it’s constantly switching between multiple real, and sometimes fake (transliterations) languages. I don’t think most users would experience this
- Alifatisk 1y agoI do use backticks a lot when sharing examples in different format when using LLMs and I have instructed them to do likewise, I also upvote whenever they respond in that matter. I got this format from writing markdown files, it’s a nice way to share examples and also specify which format it is.
- viridian 1y agoI'm sure the reason is the plethora of markdown data is was trained on. I personally use ``` stuff.txt ``` extremely frequently, in a variety of places. In slack/teams I do it with anything someone might copy and paste to ensure that the chat client doesn't do something horrendous like replace my ascii double quotes with the fancy unicode ones that cause syntax errors. In readme files any example path, code, yaml, or json is wrapped in code quotes. In my personal (text file) notes I also use ``` {} ``` to denote a code block I'd like to remember, just out of habit from the other two above.
- fkyoureadthedoc 1y agoSame, my project has a step that selects between many options when a user is trying to do some tasks. The test set for the workflow that supports this has a better success rate by about 7% on gpt-4.1-mini vs gpt-5 and gpt-5-mini (with minimal thinking)
- viridian 1y agoI'm curious what your prompts look like, as this is the opposite of my experience. I use lmarena for many of the random one shot questions I have, and I've noticed that mistral-medium is almost always the worse of the two after I blind vote. Feels like it consistently takes losses from qwen, llama, gemini, gpt, you name it. I find it overwhelmingly the most likely to produce factually untrue information to an inquiry. Would you be willing to share an example prompt? I'm curious to see what it'sesponding well to.
- barrell 1y agoI provide it with data and ask it to convert it to prose in specific formats. Mistral medium is ranked #8 on lmsys arena IIRC, so it’s probably just not your style? I’m also comparing this to gpt-5-mini, not the big boy
- viridian 1y agoI think input strategy probably accounts for the difference. Usually I'm just asking a short question with no additional context, and usually it's not the sort of thing that has one well defined answer. I'm really asking it to summarize the wisdom of the crowd, so to speak. For example, I ask, what are the most common targets of removal in magic: the gathering? Mistral's answer is so-so, including a slew of cards you would prioritize removing, but also several you typically wouldn't, including things like mox amber, a 0 cost mana rock. Gemini flash gave far fewer examples, one for each major card type type, but all of them are definitely priority targets that often defined an entire metagame, like Tarmogoyf.
- barrell 1y agoAh yeah. I’m only grading it on its prose, formatting, ability to interpret data, and instruction following. I do not use it as a store of knowledge
- epolanski 1y agoI had a similar experience on my pipeline. Was looking to both decrease costs and experiment out of OpenAI offering and ended up using Mistral Small on summarization and Large for the final analysis step and I'm super happy. They have also a very generous free tier which helps in creating PoCs and demos.
- noreplydev 1y agomistral speed is amazing
- VeryNosy 1y ago[flagged]
- neya 1y agoI'm sorry, I don't see anywhere that they work for Mistral? Is there something I'm missing here?
- barrell 1y agoI’m curious too. I’m building https://phrasing.app https://phrasing.app. A quick glance at my profile should be enough to disprove the claim that I work for mistral. I have no affiliation with mistral, I just have recent experience with them and wanted to share Also @mistral hmu if you want to arrange something!
- neya 1y agoNice to meet a fellow Elixir dev, all the best for your app, looks super cool :) Are you on Reddit? (I'm usually more active there).
- barrell 1y agoThanks for the kind words! You can find me at https://reddit.com/u/phrasingapp https://reddit.com/u/phrasingapp and https://x.com/barrelltech https://x.com/barrelltech I don’t write much about elixir (or clojure) as I’m terrible at technical writing, but I am a die hard fan of both languages
- FranklinMaillot 1y agoYou may be aware of that, but they released mistral-medium-2508 a few days ago.
- barrell 1y agoI did not! It’s not on azure yet and I’ve still got some credits to burn. That’s exciting though, hopefully it will iron out this weird ghost character issue.
- WhitneyLand 1y agoWere you using structured output with gpt-5 mini? Is there an example you can show that tended to fail? I’m curious how token constraint could have strayed so far from your desired format.
- barrell 1y agoHere is an example of the formatting I desired: https://x.com/barrelltech/status/1963684443006066772?s=46&t=aEq34ARsgqYIBlyTDH7Tkg https://x.com/barrelltech/status/1963684443006066772?s=46&t=... Yes I use(d) structured output. I gave it very specific instructions and data for every paragraph, and asked it to generate paragraphs for each one using this specific format. For the formatting, I have a large portion of the system prompt detailing it exactly, with dozens of examples. gpt-5-mini would normally use this formatting maybe once, and then just kinda do whatever it wanted for the rest of the time. It also would freestyle and put all sorts of things in the various bold and italic sections (using the language name instead of the translation was one of its favorites) that I’ve never seen mistral do in the thousands of paragraphs I’ve read. It also would fail in some other truly spectacular ways, but to go into all of them would just be bashing on gpt-5-mini. Switched it over to mistral, and with a bit of tweaking, it’s nearly perfect (as perfect as I would expect from an LLM, which is only really 90% sufficient XD)
- siva7 1y agoI thought i was the only one experiencing this slowness. I can't comprehend why something called gpt mini is actually slower than their non-mini counterpart.
- barrell 1y agoNooo you are definitely not alone. gpt-5-nano even is slowest model I’ve used since like 2023, second only to gpt-5-mini