3 ms·
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Oops did they just out GPT-5.6 sol’s parameter count?
by reilly3000 1mo ago
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Oops did they just out GPT-5.6 sol’s parameter count?
- whatever1 1mo agoI mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
- verdverm 1mo agosave qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence https://artificialanalysis.ai/models#intelligence
- Vax- 1mo agoI wonder why they removed DeepSWE from their incorporates evaluations
- verdverm 1mo agoThey didn't afaict https://artificialanalysis.ai/agents/coding-agents?coding-agents-performance-chart=deep-swe https://artificialanalysis.ai/agents/coding-agents?coding-ag... It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either
- sho 1mo agoSol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10
- nozzlegear 1mo agoRumors and allegations aren't worth much. Why don't they just tell us mere mortals?
- brookst 1mo agoWhy would they? What the upside, for them?
- eigenspace 1mo agoYeah, its not like this js some sort of Open AI company. That'd be ridiculous.
- brookst 1mo agoYou think their name means releasing competitive details would be good for them? I’ve got sone bad news about Federal Express.
- nozzlegear 1mo agoIndeed, what is the upside of transparency?
- brookst 1mo agoFor Linux? Assuring stakeholders that they can both control and observe development and how it works. For a charity? Assuring donors that funds are being managed appropriately. For Anthropic? No benefit at all. Turns out "transparency" is like "weight" or "velocity" in that it has no intrinsic value, and can be positive or negative depending on context.
- WinstonSmith84 1mo agoBecause that would reveal their edge to investors, or the lack thereof. If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn't likely to be the case ...
- kanwisher 1mo agocerebras model are different size then the original models