5 ms·
Cursor raised 900M, are losing market share to claude code(resorting to poaching 2 leads from there [1]), AND they're decreasing the value of their product? Hug
by deepdarkforest 1y ago
Cursor raised 900M, are losing market share to claude code(resorting to poaching 2 leads from there [1]), AND they're decreasing the value of their product? Huge red flag. They should be able to burn cash like no tomorrow. Also, the PR language on this post, and the timing(midnight on a US holiday) is not ideal.
This news coupled with google raising the new gemini flash cost by 5x, azure dropping their startup credits, and 2-3 others(papers showing RL has also hit a wall for distilling or improving models), are now solid signals that despite what Sam altman says, intelligence will NOT be soon too cheap to meter. I think we are starting to see the squeeze from the big players. Interesting. I wonder how many startups are betting on models becoming 5-10x cheaper for their business models. If on device models don't get good, I bet a lot of them are in big trouble
[1] https://www.investing.com/news/economy-news/anysphere-hires-anthropics-claude-code-leaders-amid-ai-talent-race--the-information-93CH-4119909 https://www.investing.com/news/economy-news/anysphere-hires-...
- exclipy 1y agoWhat happened to wafer inference hardware like Cerebras? Why isn't Claude being served from that if it's so much faster and energy efficient?
- totaa 1y agoCurrently Cerebras, although faster, is more expensive than the traditional alternatives. Cursor's use case doesn't benefit from instant, users are happy to wait the few seconds (and watching the magic may even be beneficial)
- niux 1y agoHow is it more expensive?
- recursivecaveat 1y agoFancy hardware with bespoke production process, smaller economies of scale, utilization probably not that great since they are user-speed positioning and purportedly under-invested in their compiler, which has a hard job compiling for such an arch anyways. Ignoring for the moment the cost for their bespoke software stack, which they can probably amortize away eventually.
- totaa 1y agoaccording to OpenRouter, Cerebras charges $0.65/$0.85 for 1m input/output tokens for Llama 4 Scout. Google charges $0.25/$0.70; lambda.ai charges $0.08/$0.30 for the same model.
- avianlyric 1y agoI doubt Cerebras has even close to the scale to be a major player in this area. Nvidia sold $35B of just datacenter GPUs last year. Of which the vast majority will be used for AI. Cerebra entire revenue last year was only $78M. That’s three orders of magnitude smaller than Nvidia datacenter GPU business. Scaling a company 10X in a year is a pretty hard thing to do, and it’s not a question of money, it’s a question of people and organisation. So much stuff in a business breaks when it scales 10X, that it take months to years to fix enough stuff to support another 10x growth spurt without everything just imploding.
- tom_m 1y agoAnd also if they can keep up. Imagine not just selling that many GPUs, but selling that many new GPUs every few years for the same amount of money or more. Where the previous generation hardware becomes almost worthless. The insane thing here is that $35B worth of GPUs will be worth more like $350m in a few years. Or less. Who can keep up with that???
- anonthrowawy 1y ago> papers showing RL also hitting a wall any reference for this?
- jonplackett 1y agoWhat’s funny is even electricity (nuclear in particular) isn’t ’too cheap to meter’ as originally promised. It’s actually the most expensive.
- avianlyric 1y ago> I think we are starting to see the squeeze from the big players. I’m not convinced that these price increases represent an attempt to squeeze more profit out of a saturated market. To me they look an awful lot like people realising that the sheer compute cost associated with modern models makes the historical zero-marginal cost model of software impossible. API calls to LLM models have far more in common with making calls to EC2 or Lambda for compute, than they do a standard API calls for SaSS. A lot of early LLM based business models seemed to assume that the historical near zero-marginal cost of delivery for software would somehow apply to hosted LLM models which clearly isn’t the case. You mix that in with rising datacenter costs, driven by lack of available electricity infrastructure to meet their demands, plus everyone trying to grab as much LLM land as possible, which requires more datacenters, more faster. And the result is rapidly increasing base costs for compute. Which we’re now seeing reflected in LLM pricing. For me the thing that stands out about LLMs, is that their compute costs are easily 100-10000x greater per API call than a traditional SaSS API call. That fact alone should be enough for people to realise that the historically bottomless VC money that normal funds this stuff, isn’t quite a bottomless as it needs to be to meaningfully subsidise consumer pricing.
- patapong 1y agoVery insightful. I think the payment model would have worked out just fine if the state of the art was the optimization of GPT-4 class models to bring down the cost over time, which would have made the services profitable eventually. Instead, newer models are getting larger and more resource heavy through reasoning, meaning costs per request are going up instead of down.
- avianlyric 1y agoI think we’ll start seeing people focus on optimisation, we already see companies like Apple focus on it. LLM are still to new, and still advancing to quickly for optimisation to take place. It’s like we’re back in the MHz wars of old between CPU manufacturers. The goal is just more performance, regardless of cost, because it was clear that even in the consumer space, people wanted more performance. Then we hit a kind of plateau in last 10 years, where basic compute is so powerful that your average consumer is not longer upgrading every year for better performance. A 5 year old machine has enough performance for most people. Then the focus on energy efficiency kicked in, because people didn’t want faster computers, they wanted battery life and cheaper computers. No doubt we’ll see the same with LLM, possibly quite soon. Claude Sonnet 4 and similar class models have enough reasoning performance, that agentic systems can be quite reliable. Which means we hit the base level of “reasoning” performance needed, and we can extend that “performance” in domain specific ways by lightly customising the agentic framework, with no need to fine tuning. The elimination of fine tuning to build domain specific agents is a huge game changer. But it also means that putting together a 10x or 100x efficient model, with “reasoning” performance equivalent to current gen LLM would also be a huge game changer. It opens up the possibility to apply this tech into spaces that currently require either lots of specialists knowledge to fine tune an LLM, or a huge amount of on tap compute to allow the agents to take enough turns to slowly “reason” they’re way through problems. But a Claude Sonnet 4 that runs on a iPhone for example. That would make Apple’s complete failure to improve Siri look like a genius level move. Why bother with small incremental improvements using current tech, when waiting a few years, and just stuffing a full fat LLM and agent system into an iPhone will basically give you the ultimate Siri.
- msgodel 1y agoHeh. Qwen3 still works on my machine.
- tom_m 1y agoIt was obvious that they over raised. It's insane to raise that kinda money to go sell an open-source and otherwise free code editor that wraps an LLM that you don't own or host. So you're not providing the service, you don't have to code much of a product because you're front loaded 95% of the product with open source...you have no secret sauce... And you're going to raise that kinda money for what exactly? In hopes you can fool a bunch of people? Do they think they're secret sauce is UX? There's better editors out there now too. You want to know what the hype train of Cursor was for? It was marketing for LLMs.