3 ms·
Hey! Banana cofounder here. Firstly, thank you so much for trying our service, we'll do our best to meet performance expectations and win you back! Re: #1 and
by edunteman 3y ago
Hey! Banana cofounder here.
Firstly, thank you so much for trying our service, we'll do our best to meet performance expectations and win you back!
Re: #1 and #2, cold boots are the most vital thing for us to solve, because it fixes #1 directly and helps #2 indirectly. As we drive cold starts exponentially toward 0s (obv impossible but there's a near-zero asymptote to what's possible, limited only by disk read throughput), it makes it more viable to actually run serverless and scale from 0->1 for each user call. Our goal RN is to hit 1s, as that's generally the sweetspot for LLM app builders to stop feeling pain from the cold boot. We're getting close (2-5s) for most models with Turboboot (shill warning https://www.banana.dev/blog/turboboot https://www.banana.dev/blog/turboboot). Glorious future would be more like 100ms, but depending on where the size of models end up, 1s+ cold boots may just be the cost of doing business.
Re: #3, if I understand, you're looking for interactive compute (IE you ssh in, mess around, do training runs, attach a jupyter notebook). For my personal ML training I use and suggest Lambda Labs or Brev.dev. Have heard great things of Coreweave, and users of Mosaic have seemed quite satisfied but I don't believe it's as interactive as you may want. Banana has no plans to support interactive GPU sessions, to conserve focus toward being best at cold boots.
You're definitely not alone in this wishlist, so I validate you. If I were building applications on top of a provider, I'd expect the same things. Big gnarly challenge with all the tools being a few years old at best! Fastest route to dependable tools is intense focus, us on cold boots.
AMA, if anyone is interested in digging
- r3trohack3r 3y ago> cold boots are the most vital thing for us to solve ... Our goal RN is to hit 1s Did a bit of package management and experimenting with optimizing bits over the wire for serverless scaling on NFLX's internal serverless platform. I'm selfishly interested in learning more about what you're doing to optimize meeting an incoming request with a live instance, but also might be able to help. Know you're busy but if you, or your engineering team, have time to connect I'd love to chat: hn@blankenship.io
- edunteman 3y agoNeat! Bit too busy now, but we'll hopefully put some technical blogs over time to explain these things. We don't do any predictive scaling yet; only when a call hits the queue do we scale replicas. Replicas cold boot (pod scheduling + application loading models into GPU memory) then subscribe to the queue. Moving away from this replica + queue design very soon; using K8s primitives has gotten us this far but impossible to hit 1s cold boots with orchestration overhead. Building our own orchestrator and python runtime now.
- r3trohack3r 3y agoSome questions, no obligation to answer: Are you colocating your storage plane and GPUs? What’s ingress/egress to a node and are those links near saturated (with comfy room for returning model output, but I’m assuming moving the models dwarfs I/O from customer workloads)? Do you see high reusability across workloads? Have you explored chunking/hashing your workloads IPFS style (do these models radically change, or is there a high chance that two models that share an ancestor also share 50% of their bits. If you’re chunking your models and colocating the storage plane with GPUs, can you distribute chunks to increase the hit-rate of a chunk being on-node? Is your scheduler aware of the existing distribution of chunks across nodes? Given the workload patterns you see, and the shared bits between models, is it even practical to try and chase a local cache hit rate to reduce bits-over-wire? If you have a cache miss, what’s the path to getting those bits to the node with the GPU? How does the cost of that path compare to the cost of the scheduler making a decision?
- edunteman 3y agoCan only publicly answer one of these: - reusability of workloads: yes, introducing the community templates feature (https://banana.dev/templates https://banana.dev/templates) for common models has dramatically cut back on storage requirements and transfers. We're still majority custom code, but it's helped prevent us from exploding storage over people running the same "model of the week" As for caching / chunking, sounds like you're thinking on our wavelength, perhaps even ahead of us, so maybe I should take you up on the offer to chat! Will reach out.