4 ms·
Open-weight models are not one month behind. In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.
by user43928 21d ago
Open-weight models are not one month behind.
In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.
- cgio 21d agoFor practical uses they are there. Arguably the frontier models are worse for some of these practical tasks. And keep in mind, people will use maybe frontier for 1/10th of the work, planning and review, and go open source for rest. The question is if they manage to impose outside us. If not, they are losing competitiveness.
- user43928 21d agoI always thought switching from a SOTA model to a dumber model after planning was a terrible idea. Mostly I heard this from people who I got the impression have little experience in developing greenfield software with agentic AI. Often the same people who talk about spec frameworks. I fundamentally disagree with the approach. I believe the ability to autonomously evaluate, test, and adjust during long horizon tasks is critical to using AI efficiently.
- rescbr 21d agoWell, it's the enterprise software house pipeline... The software architect writes the spec, hands it down to the implementation team, senior leads, junior devs or offshore teams codes it. I also disagree with the approach, this is cargo-culting the existing ways of working.
- jurgenburgen 21d ago> this is cargo-culting the existing ways of working. I think you mean the existing anti-patterns. They don’t call them ivory tower architects for nothing.
- lukewarm707 21d agoopen models are ahead in speed. they complete tasks as fast as you choose to scale compute. they are more efficient and require less compute for the same thing. they are ahead in specialized tasks. they are ahead in areas closed models refuse to answer. they are ahead in emotional intelligence.
- user43928 21d agoI doubt most of your claims. Maybe the guardrails and emotional intelligence is true. For speed and efficiency, you are most likely wrong. Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot serve a single DeepSeek Flash 4.1 stream at 750 t/s, plus the model is less intelligent as seen on newer benchmarks. I believe OpenAI and Anhropic are at the frontier of efficiency too. There were numerous reports about their breakthroughs and associated API price cuts. The idea that open-weight models are more efficient seems unfounded.
- gajjanag 21d ago> The idea that open-weight models are more efficient seems unfounded. https://artificialanalysis.ai/ https://artificialanalysis.ai/ intelligence vs cost per task disagrees with this statement.
- lukewarm707 21d agoi can provide some sources. 5.6 sol ultrafast on cerebras is 750tps, open models readily exceed this. just by using a smaller model cerebras serves qwen 3.8 27b at 1850tps. or even larger models, mimo 2.5 pro was served for a while at 1000tps. and so on. [https://inference-docs.cerebras.ai/models/choose-a-model https://inference-docs.cerebras.ai/models/choose-a-model] the chinese ai companies have 10% of the total compute resources of the US ones. since the USA tries to stop them from buying nvidia gpus. they maxed out the efficiency. deepseek v4.1 has engram architecture. it has 550b params instead of 5T+ for astra/fable. it has 8b active instead of potentially hundreds active for astra/fable. compare input/output/cache: $0.15/$0.60/$0.003 for v4.1 to $10.00/$50.00/$1.00 for astra and $10.00/$50.00/$0.25 for fable. astra cache reads are over 330 times more expensive. at the artificial analysis 7:2:1 ratio, deepseek is $0.18/m, fable is $7.18/m, astra is $7.7/m. but what about intelligence? AA would rate deepseek v4.1 at AA 40, astra is AA 53. so it cost 4,180% more for 32% more intelligence. they are serving that at over 250tps at baseten. to get close to that on astra API you are paying double the cost for fast mode. so it is now 8456% more expensive for a similar speed and 32% more intelligence. 84 times more expensive.
- mistercow 21d agoI’m not sure how you can really make either statement work anymore. Now that smaller models are actually broadly usable, “behindness” is no longer a scalar and at the tails, where no open lab seems to be trying to compete at the >10T scale and no closed lab seems to care about <400B anymore, it’s just apples to oranges. It’s like talking about whether Qualcomm is “behind” Nvidia.
- epolanski 21d ago1. Mythos wasn't released in February. Let's stick to only public-facing models. 2. For public-facing models, the differences are really minor with some occasional model (like Fable or Astra) showing some better performance in specific benchmarks for the span of some weeks or few months before open ones catch it. 3. Being bleeding edge is overblown anyway in the real world, besides the occasional "very latest fresh model did this task which previous one couldn't", and the number of those tasks is increasingly small and far from mundane corporate needs.
- jurgenburgen 21d agoStill waiting for our org to roll out Mythos. I guess it was too expensive so we’re stuck on the previous model until the internal team can figure out self-hosting open models.