3 ms·
It's a limit on input tokens. So that's 3 50k requests per minute. At Cerebras speeds, that's about 5 seconds of usage per minute. I was very excited last year
by wild_egg 1mo ago
It's a limit on input tokens. So that's 3 50k requests per minute. At Cerebras speeds, that's about 5 seconds of usage per minute.
I was very excited last year for their coding plan but seeing a burst of requests pulse and then sitting there watching the cooldown reset is really not a great time.
Even though each individual request was fast, the sessions were only maybe 10% faster on wall clock time since there was so much waiting time.
- amelius 1mo agoCan't you do something with multiple accounts?
- sandworm101 1mo agoOr just buy a 5060. This will run on most any 16gb card. Slower for sure but far cheaper than another subscription.
- embedding-shape 1mo agoOr buy a raspberry pi with a SSD, about the same difference, if you're giving up on the 1500 tokens/s anyways.
- ma2kx 1mo agoThats not the point if you choose Cerebras as provider.
- interactivecode 1mo agoThe whole point is the speed
- nateb2022 1mo ago[dead]
- jychang 1mo agoYou would lose caching (if they cache)
- kristjansson 1mo agoThey made the coding plan a bit better toward the end, but it was pretty tough to use throughout. Seems like an Amdahl’s law of inference economics? there’s so much compute relative to SRAM on the chip and shoreline bandwidth onto the chip that caching buys ~nothing? The contended resource is SRAM and a given token of context needs just as much as another.
- vlovich123 1mo agoThat’s not what caching is for. Caching lets you resume with a pre computed KV cache saving you from having to ingest everything in the chat history as input on every single round trip. You still need caching regardless of SRAM or not as it saves a huge amount (and ever growing) of compute ingesting the preceding history every time you want a completion. I don’t know why they don’t give a price discount. Maybe their hardware is incapable for some reason of saving/restoring the state? Or maybe they just haven’t built the infrastructure to do it?
- jmalicki 29d ago> Maybe their hardware is incapable I am unsure if it is incapable, but it sounds hard. They have tons of cores with 64k of SRAM each, and relatively slow paths in/out. On a GPU you can leave it in SRAM. On Cerebras, you have to send it out of the system which is a giant bottleneck.