24 ms·
We were wrong about GPUs
- doctorpangloss 2y ago> GPUs terrified our security team. Ha ha, it didn't terrify Modal. It ships with all those security problems, and pretends it doesn't have them. Sorry Eric.
- arccy 2y agoif your security team is 0 people, do you terrify all or none of them?
- deleted 2y ago[deleted]
- tptacek 2y agoWe wrote all sorts of stuff this week and this is what gets to the front page. :P
- Rzor 2y agoSign of the times.
- ethbr1 2y agoSounds like you might have been wrong in spending time on the other stuff. ;) Who even knows what the customer is ever going to want? Pivot. Pivot. Pivot. PS: And pouring one out for the engineering hours that went into shipping GPUs. Sometimes it's a fine product, but just doesn't fit.
- transpute 2y ago> We burned months trying (and ultimately failing) to get Nvidia’s host drivers working to map virtualized GPUs into Intel Cloud Hypervisor... We think there’s probably a market for users doing lightweight ML work getting tiny GPUs. This is what Nvidia MIG does, slicing a big GPU into arbitrarily small virtual GPUs. But for fully-virtualized workloads, it’s not baked; we can’t use it. Near as we can tell, MIG gives you a UUID to talk to the host driver, not a PCI device. Apparently this is technically possible, if you can find the right person at Nvidia to talk about vGPU licensing and magic incantations. Hopefully someone reading this HN front page story can make the introduction.
- tptacek 2y agoWe (I) spent a lot of time talking to several different teams at Nvidia about this. We were able to get VFIO vGPUs to the point where guest libraries would recognize them, but the process fell apart in the guest/host licensing dance, and we weren't really OK with the idea that there'd be a phone-home licensing dance every time a Fly Machine started. Unlike GPU enablement at GCP or AWS, the core DX of a Fly Machine is that it stop/start very quickly; think of it as a midpoint in the design space between Lambda and Fargate. This is what we're talking about when we say it's hard to fit GPUs into our DX.
- transpute 2y ago> phone-home licensing dance every time a Fly Machine started To userspace Nvidia license server (a) in each host, (b) for entire Fly cloud, or (c) over WAN to Nvidia cloud?
- tptacek 2y agoIIRC, (a) and (c), which I think sort of implies (b)? Really what we'd have wanted to do would have been to give Fly Machines MIG slices. But to the best of my understanding, MIG is paravirtualized; it doesn't give you SR-IOV-style PCI addresses for the slices, but rather a reference Nvidia's userland libraries pass to the kernel driver, which is a dance you can't do across VM boundaries unless your hypervisor does it deliberately.
- transpute 2y agoHypothetical scenario for Nvidia licensing of fast-start microVMs: 1. Instead of blocking VM start for license validation, convert that step into non-blocking async submission of usage telemetry, allowing every VM to start instantly. For PoC purposes, Nvidia's existing stack could be binary patched to proxy the license request to a script that isn't blocking VM start, pending step 2 negotiation. 2. Reconcile aggregate vGPU usage telemetry from Nvidia Fly-wide license server (Step 1) with aggregate vGPU usage reports from Fly's orchestration/control plane, which already has that data for VM usage accounting. In theory, Fly orchestration has more awareness of vGPU guest workload context than Nvidia's VM-start gatekeeping license agent, so there might be mutual interest in trading instant VM start for async workload analytics.
- fuddle 2y agoYou just wrote a blog post about whats needed for a top HN post. You should have known it would do well :) https://fly.io/blog/a-blog-if-kept/ https://fly.io/blog/a-blog-if-kept/
- nicoburns 2y agoFWIW, I'd consider this good publicity. I have no use for GPUs in the cloud (at least not at the prices that they're available at (in general, not just on fly)). So if fly is moving there effort towards things I actually need (like managed databases) then that's going to increase my confidence in them as a platform quite a bit.
- narag 2y agoEmpathy. Such a seemingly good idea. Maybe just a little ahead of its time.
- Philpax 2y agoI noticed quite a few spelling and grammar mistakes - could do with a bit of an edit pass?
- ksec 2y agoIn the current days of AI I think spelling and grammar mistakes is perhaps a great way to tell it is still written by human...... ( Until AI copy this )
- glouwbug 2y agoI'm sure AI is already capable of linking anonymous forum handles by writing style. Best to reinvent ourselves every month like one would with monthly password resets
- mrkurt 2y agoI do, in fact, instruct LLMs to make spelling and grammar mistakes when I have them reply to cold emails.
- ignoramous 2y agoDo you run those LLMs on Fly? ;)
- dangus 2y agoNah, fly.io has a company culture that is all about having lots of bugs and issues, and that includes blog posts. The idea that a cloud compute provider can’t make GPU compute into an profitable business is pretty laughable.
- _zoltan_ 2y agoI have to agree with this. Look at GPU utilization at AWS, Azure, .. they are running close to 100%. for our p5 quota I had to talk to our TAM team on AWS, while most of our quota requests are instant usually.
- akoculu 2y agoI spent a month setting up serverless endpoint for a custom model last year with Runpod. It was expensive and unreliable, in addition to long cold boot times. The product was unusable even as a prototype, to cover the costs, I'd have to raise money first. In a different product, I was given some Google Cloud credits, which unlocked me to put the product in front of customer. This one also needed GPU but not as expensive as the previous. It works reliably and it's fast. Personally, I had two use cases for GPU providers in past 3 months. I think there's definitely demand for reliability and better pricing. Not sure Fly will be able to touch that market though as it's not known for both (stability & developer friendly pricing). P.S If anyone is working on a serverless provider and want me to test their product, reach me out :)
- zackangelo 2y agowould love for you to test a serverless llm product i'm working on, zack [at] mixlayer.com
- dnavani 2y agoGive https://modal.com https://modal.com a spin -- email me at deven [at] modal.com and happy to help you get set up
- BoorishBears 2y agoFwiw Runpod also has a startup program. Ironically GCP and AWS GPUs are so overpriced that getting even half the number of credits from Runpod is like a 4x increase in "GPU runway", especially with .44/hr A40s.
- chr15m 2y agoSide note: "we were wrong" - are there any more noble and beautiful words in the English language?
- brunoqc 2y agoIt's great when people admit they were wrong but I can't help to find those headlines clickbaity. A bit like "stop doing this..." and we think: omg, am I doing the same deadly mistake?
- tptacek 2y agoI love the idea that Kurt needed to better couch a post saying he was wrong about something.
- deleted 2y ago[deleted]
- sonofhans 2y ago“I was wrong.” Closely followed by, “I was right.” :)
- fragmede 2y agoI don't know.
- anotherhue 2y agoNext week: fly introduces game streaming technology for indie game devs.
- iFire 2y agoWhat GPUS services will you keep?
- jeffybefffy519 2y agoI feel like these guys are missing a pretty important point in their own analysis. I tried setting up a ollama LLM on a fly.io GPU machine and it was near impossible because of fly.io limitations such as: 1. Their infrastructure doesnt support streaming responses well at all (which is important part of the LLM experience in my view) 2. The LLM itself is massive, and cant be part of the docker image I was building and uploading. Fly doesnt have a nice way around this, so I had to setup a whole heap of code to pull it in on the fly machines first invocation, which doesnt work well if you start to run multiple machines. It was messy and ended up with a long support ticket with them that didnt get it working any better so I gave up.
- tptacek 2y agoI mean, yes? Managing giant model weight files is a big problem with getting people on-demand access to Docker-based micro-VMs. I don't think we missed that point so much as that we acknowledged it, and found some clarity in the idea that we weren't going to break up our existing DX just to fix it. If there were lots and lots and lots of people trying to self-host LLMs running into this problem, it would have been a harder call.
- akoculu 2y agoDid you consider other use cases in which people need custom models and inference other than just open source LLMs ?
- tptacek 2y agoYes. Click through to the L40S post the article links to (the L40S's aren't going anywhere). There are people doing GPU-enabled inference stuff on Fly.io. That particular slice of the market seems fine?
- johntash 2y agoWhat kind of issues did you have with streaming? I also set up ollama on fly.io, and had no issues getting streaming to work. For the LLM itself, I just used a custom startup script that downloaded the model once ollama was up. It's the same thing I'd do on a local cluster though. I'm not sure how fly could make it better unless they offered direct integration with ollama or some other inference server?
- onli 2y agoNot sure about this: > like with our portfolio of IPv4 addresses, I’m even more comfortable making bets backed by tradable assets with durable value. Is that referencing the gpus, the hardware? If yes, why should they have a durable value? Historically hardware like that deprecated fast and reaches a value of 0, energy efficiency alone kills e.g. old server hardware. Something different here?
- foota 2y agoHistorically GPUs didn't cost 30,000 a pop. Also, the end of Moore's law etc.,.
- thundergolfer 2y agoI think it's referencing only the IPv4 block, but it is a bit confusing. It doesn't make sense to be ref'ing the GPUs because their value is definitely not durable.
- tptacek 2y agoI don't know about "durable", but they're not written off. There is absolutely a market for all this hardware.
- Aeolun 2y agoThere’s certainly more retained value in the physical stuff than in developer time.
- thundergolfer 2y agoYeah fair enough. I think it's just the subjectivity of "durable" at play here. The value of the GPUs may halve in a single year (e.g. H100s), but they'll never* drop to zero in a month. That's at least some kind of durability, because you can get a transaction done in a month. * never say never
- jonathanlei 2y agoAbsolutely - GPUs are definitely not a very liquid asset. As someone who works at a GPU neocloud provider (Voltage Park), server assets at scale definitely face a huge slippage, you can buy for $1 and get quotes for $1.50 but only be able to sell for $0.60
- serjester 2y agoI respect them for being public about this. With that said, this seems quite obvious - the type of customer that chooses Fly, seems like the last person to be spinning up dedicated GPU servers for extended periods of time. Seems much more likely they'll use something serverless which requires a ton of DX work to get right (personally I think Modal is killing it here). To compete, they would have needed to bet the company on it. It's way too competitive otherwise.
- BoorishBears 2y agoAs someone who deploys a lot of models on rented GPU hardware, their pricing is not realistic for continous usage. They're charging hyperscaler rates, and anyone willing to pay that much won't go with Fly. For serverless usage they're only mildly overpriced compared to say Runpod, but I don't think of serverless as anything more than an onramp to renting dedicated machine, so it's not surprising to hear it's not taking off. GPU workloads tend to have terrible cold-start performance by their nature, and without a lot of application specific optimizations it rarely ends up making financial sense to not take a cheaper continous option if you have an even mildly consistent workload. (and if you don't then you're not generating that much money for them)
- tptacek 2y agoMy thing here is just: people self-hosting LLMs think about performance in tokens/sec, and we think about performance in terms of ms/rtt; they're just completely different scales. We don't really have a comparative advantage for developers who are comfortable with multisecond response times. And that's fine!
- cmdtab 2y agoThat reminds me when cloudflare launched their workers gpu product, it was specifically aimed at running models and the pricing was abstracted and based on model output. Did you look what they were doing when building gpu machines? https://blog.cloudflare.com/workers-ai/ https://blog.cloudflare.com/workers-ai/
- 2y ago
- johntash 2y agoI really liked playing around with fly gpus, but it's just too expensive for hobby-use. Same goes for the rest of fly.io honestly. The DX is great and I wish I could move all of my homelab stuff and public websites to it, but it'd be way too expensive :(
- mrkurt 2y agoThis is near and dear to me, because I want people to run stuff like homelabs and side projects. What part of the cost gets out of hand? Having to have a Machine for every process? Do you remember what napkin math pricing you were working with?
- johntash 2y agoHmm, having a machine for every process is part of it but I actually like that kind of isolation. Storage and bandwidth also add up fast. For example, I could get a digitalocean vm with 2gb ram, 1vcpu, 50gb storage, 2tb bandwidth for $12/mo. For the same specs at fly.io, it'd be ~$22/mo not including any bandwidth. It could be less if it scales to zero/auto stops. I recently tried experimenting with two different projects at fly. One was an attic server to cache packages for NixOS. Only used by me and my own vms. Even with auto scaling to zero, I think it was still around $15-20/mo. The other was a fly gpu machine with Ollama on it. The cold start time + downloading a model each time was kind of painful, so I opted for just adding a 100gb volume. I don't actually remember what I was paying for that, but probably another 20/mo? I used it heavily for a few days to play around and then not so much later. I do remember doing the math and thinking it wouldn't be sustainable if I wanted to use it for stuff like home-assistant voice assistant or going through pdfs/etc with paperless. On their own, neither of these are super expensive. But if I want to run multiple home services, the cost is just going to skyrocket with every new app I run. If I can rent a decent dedicated server for $100-$200/mo, then I at least don't have to worry about the cost increasing on me if a machine never scales to zero due to a healthcheck I forgot about or something like that. Sorry if it's a bit rambly, happy to answer questions!
- mrkurt 2y ago
- kirillzubovsky 2y ago[flagged]
- a-r-t 2y agoOff topic, but the font in the article is hard on the eyes.
- hoppp 2y agoGpus don't fit the usual "start with free tier then upgrade when monetizing" approach most devs have with these kinds of platforms. For simple inference, its too expensive for a project that makes no money. Which is most projects.
- mrcwinn 2y agoHas service reliability improved at all? I tried Fly at two different points in time and I’ve never had a worse experience with a service.
- loloquwowndueo 2y agoYou didn’t say at which points in time so it’s kind of hard to say yes but I will say “yes, reliability has improved”.
- mrcwinn 2y agoOkay, I'll try a different question. How's reliability these days, lolo?
- cyberax 2y agoHah. We're doing AI, but we're doing vision-based stuff and not LLMs. For us, the problem has been deploying models. Google and AWS helpfully offered their managed LLM AI services, but they don't really have anything terribly more useful than just machines with GPUs. Which are expensive. I'm going to check fly.io...
- reilly3000 2y agoI shelled out for a 4090 when they came out thinking it would be the key factor for running local llms. It turns out that anything worth running takes way more than 24GB VRAM. I would have been better off with 2+ 3090s and a custom power supply. It’s a pity because I thought it would be a great solution for coding and a home assistant, but performance and quality isn’t there yet for small models (afaik). Perhaps DIGITS will scratch the itch for local LLM developers, but performant models really want big metal for now, not something I can afford to own or rent at my scale.
- UncleOxidant 2y agoThere was a post on r/localLlama the other day about a presentation by the company building Digits hardware for Nvidia. The gist was that Digits is going to be aimed at academic AI research folks and as such don't expect them to be available in large numbers (at least not for this first version). It was disappointing. Now I'm awaiting the AMD Strix Halo based systems.
- wmf 2y agoLaptops with the same chip however...
- unethical_ban 2y agoI haven't tested programming tasks with a local LLM vs. say, Claude 3.5. But it is nice to be able to run 14-32B LLMs locally and get an instant response. I have a single 3090.
- prettyblocks 2y agoSame here. I just built a pc with a 3090 for local llm and stable diffusion and have zero regrets.
- 01HNNWZ0MV43FF 2y agoGosh. Good thing I haven't bought a GPU in almost a decade. With a little luck I'll catch this wave on the back end. I haven't had to learn web or mobile development thoroughly either
- silisili 2y agoI'm admittedly a complete LLM noob, so my question might not even make sense. Or it might exist and I haven't found it quite yet. But have they considered pivoting some of said compute to some 'private, secure LLM in a box' solution? I've lately been toying with the idea of training from extensive docs and code, some open, some not, for both code generation and insights. I went down the RAG rabbit hole, and frankly, the amount of competing ideas of 'this is how you should do it', from personal blogs to PaaS companies, overwhelmed me. Vector dbs, ollama, models, langchain, and various one off tools linking to git repos. I feel there has to be substantial market for whoever can completely simplify that flow for dummies like me, and not charge a fortune for the privilege.
- nemothekid 2y agoThe problem is currently all the "competing" ideas have a ton of tradeoffs and is rarely one size fits all. Furthermore, it's not clear if the idea you choose will be become obsolete by the underlying model architecture getting better. On top of all that, you are essentially competing with Anthropic/OpenAI/Google, where your only advantage is "privacy". Anyone who deeply cares about privacy, and is willing to pay for it, may just likely do it on their own (especially if your CTO is pouting money into "investing" in AI). Anyone who doesn't will likely not want to 6-7 months behind what you can get at OpenAI or Google.
- VectorLock 2y agoOut of curiosity, how much runway does fly.io have (without raising new funding?)
- jeremyjh 2y agoI'd guess significantly less after this debacle.
- inetknght 2y agoKudos to owning up to your failed bet on GPUs even if you are putting a lot of blame on Nvidia for it. And to be fair, you're not wrong. Nvidia's artificial market segmentation is terrible and their drivers aren't that great either. The real problem is the lack of security-isolated slicing one or more GPUs for virtual machines. I want my consumer-grade GPU to be split up into the host machine and also into virtual machines, without worrying about resident neighbor cross-talk! Gosh that sounds like why I moved out of my apartment complex, actually. The idea of having to assign a whole GPU via PCI passthrough is just asinine. I don't need to do that for my CPU, RAM, network, or storage. Why should I need to do it for my GPU?
- yieldcrv 2y agoYes devs want LLMs, but also the price of inference compute plummeted 90% over the last 18 months, which is primarily in gpus So it’s not just that openai and anthropic apis are good enough, they are also cheap enough, and still overpriced compared to the industry Your GPU investment wont do as well as you thought, but also you are wasting time on security. If the end user and market doesnt care then you can consider not caring as well. Worst case you can pay for any settlement with …. more gpu credits.
- mmastrac 2y agoIt's really a shame GPU slices aren't a thing -- a monthly cost of $1k for "a GPU" is just so far outside of what I could justify. I guess it's not terrible if I can batch-schedule a mega-gpu for an hour a day to catch up on tasks, but then I'm basically still looking at nearly $50/month. I don't know exactly what type of cloud offering would satisfy my needs, but what's funny is that attaching an AMD consumer GPU to a Raspberry Pi is probably the most economical approach for a lot of problems. Maybe something like a system where I could hotplug a full GPU into a system for a reservation of a few minutes at a time and then unplug it and let it go back into a pool? FWIW it's that there's a large number of ML-based workflows that I'd like to plug into progscrape.com, but it's been very difficult to find a model that works without breaking the hobby-project bank.
- beebaween 2y agoThis is what services like Vast.ai are for - super cheap GPUs you just use as long as you need etc etc.
- mmastrac 2y agoHmm, that looks interesting -- I might have to explore a bit. The low-end GPU pricing is pretty competitive.
- montecarl 2y agoDo you think that you can use those machines for confidential workflows for enterprise use? I'm currently struggling to balance running inference workloads on expensive AWS instances where I can trust that data remains private vs using more inexpensive platforms.
- mmastrac 2y agoI read through the FAQ and the answer is "no", but they say it basically as "nobody really cares what your data is". I wouldn't put anything confidential through it.
- gopher_space 2y ago
- sergiotapia 2y agoyou guys have all this juicy GPU and infrastructure. why not offer models as apis? i would pay to have apis for: sam2, florence, blip, flux 1.1, etc. whatever use case I would have reached Fly for on GPU, i can't justify _not_ using Replicate. maybe Fly can do better offer premium queues for that with their juicy infra? you're right! as a software dev I see dockerization and foisting these models as a burden, not a necessity.
- latchkey 2y agoThere is no market for MIG in the cloud. People talk about it a lot, but in reality, nobody wants a partial GPU (at least not paying for it). One interesting thing about all this is that 1 GPU / 1 VM doesn't work today with AMD GPUs like MI300x. You can't do pcie passthrough, but AMD is working on adding it to ROCm. We plan to be one of the first to offer this.
- dathinab 2y ago> A whole enterprise A100 is a compromise position for them; they want an SXM cluster of H100s. For a lot of use-cases you need at lest two A100s with a very fast interconnect, potentially many more. This isn't even about scaling with requests but about running one single LLM instance. Sure you will find all of ways how people managed to runt his or that on smaller platforms, problem is that quite often doesn't scale to what is needed in production for a lot of subtle and less subtle reasons.
- freedomben 2y ago> The biggest problem: developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. System engineers may have smart, fussy opinions on how to get their models loaded with CUDA, and what the best GPU is. But software developers don’t care about any of that. When a software developer shipping an app comes looking for a way for their app to deliver prompts to an LLM, you can’t just give them a GPU. I'm increasingly coming to the view that there is a big split among "software developers" and AI is exacerbating it. There's an (increasingly small) group of software developers who don't like "magic" and want to understand where their code is running and what it's doing. These developers gravitate toward open source solutions like Kubernetes, and often just want to rent a VPS or at most a managed K8s solution. The other group (increasingly large) just wants to `git push` and be done with it, and they're willing to spend a lot of (usually their employer's) money to have that experience. They don't want to have to understand DNS, linux, or anything else beyond whatever framework they are using. A company like fly.io absolutely appeals to the latter. GPU instances at this point are very much appealing to the former. I think you have to treat these two markets very differently from a marketing and product perspective. Even though they both write code, they are otherwise radically different. You can sell the latter group a lot of abstractions and automations without them needing to know any details, but the former group will care very much about the details.
- varenc 2y agoAren’t we just continually moving up layers of abstractions? Most of the increasingly small group doesn’t concern itself with voltages, manually setting jumpers, hand-rolling assembly for performance-critical code, cache line alignment, raw disk sector manipulation, etc. I agree it’s worthwhile to understand things more deeply but developers slowly moving up layers of abstractions seems like it’s been a long term trend.
- Moru 2y agoWe certainly need abstractions for the first layer of the hardware. An abstraction of the abstraction can be useful if the first abstraction is very bad or very crude. But we are now at an abstraction of an abstraction x 8 or so. It's starting to get a bit over the top.
- aqueueaqueue 2y ago> developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. Is there not a market for the kind data science stuff where GPUs help but you are not using an LLM. Like statistical models on large amounts of data and so on. Maybe fly.io customer base isn't that sort of user. But I was pushing a previous company to get AWS GPUs because it would save us money vs CPU for the workload.
- thundergolfer 2y agoThere is a market but it most likely requires a thick software layer to enter the competitive space. Modal Labs, Anyscale, and Outerbounds are examples of companies competing for "data science stuff" and have thick software layers over the VMs.
- dijksterhuis 2y agonot from fly.io, but my experience is that most data scientists will just prefer to lump it with the tools they know (pandas / R) on CPUs, rather than delving into things like rapids https://rapids.ai https://rapids.ai -- even if it makes things faster/cheaper. I might have had a bad sample set so far. But the "doing statistics" bit seems to be the interesting thing for them. the tooling doesn't really factor into solutions/plans that often. and learning something new because "engineer say it shinier" doesn't really seem to motivate them much :/
- aqueueaqueue 2y agoDo many DS use Google Colab and click the GPU option? That made me think GPUs would be more popular (due to speed). Also GPUs may be used when productionizing work done by DS but maybe I am in a tiny niche here of (Data Science) intersection (Scale up) minus (Deep learning LLM etc.)
- jonathanyc 2y ago> The biggest problem: developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. I considered using a Fly GPU instance for a project and went with Hetzner instead. Fly.io’s GPU offering was just way too expensive to use for inference.
- Aeolun 2y agoHetzner is more expensive by default though? It starts at $200/month. Which is fine if you are running for 720 hours every month, but you can run more cheaply on fly if it doesn’t get used more than 150ish hours in a month.
- lifeisstillgood 2y agoI have a timeline that I am still trying to work through but it goes like this : 2012 - moores law basically ends - nand gates do t get smaller just more cleverly wrapped. Single threaded execution more or less stops at 2 GHz and has remained there. 2012-2022 - no one notices single threaded is stalled because everything moves to VMs in the cloud - the excess parallel compute from each generation is just shared out in data centres 2022 - data centres realise there is no point buying the next generation of super chips with even more cores because you make massive capital investments but cannot shovel 10x or 100x processes in because Amdahls law means standard computing is not 100% parallel 2022 - but look, LLMs are 100% parallel hence we can invest capital once again 2024 - this is the bit that makes my noodle - wafer scale silicon. 900,000 cores with GBs SRAM - these monsters run Llama models 10x faster than A100s We broke moores law and hardware just kept giving more parallel cores because that’s all they can do. And now software needs to find how to use that power - because dammit, someone can run their code 1 million times faster than a competitor - god knows what that means but it’s got to mean something - but AI surely cannot be the only way to use 1M cores?
- Yoric 2y agoPlenty of non-IT applications use lots of cores, e.g. physics simulations, constraint solving, network simulation used to plan roads or electrical distribution, etc.
- lifeisstillgood 2y agoYes - but the amount of code that loves Amdahls law, is tiny compared to amount code churned out each day that can never run parallel over 1M cores - no matter how clever a compiler gets. I cannot work out if we pack enough parallel problems in the world or just lack a programming language to describe them
- sroussey 2y agoComputers are still Von-Neumann machines, and other architectures lost out due to the great returns on investment for that architecture. However, in the AI world, this might not be the case. For instance, neuromorphic computing is one example, and there are others. Or back to analog again! Superposition is instant—no slow adders with carry bits to propagate! Who knows. Fun times!
- martinsnow 2y ago[flagged]
- keyle 2y agoIt feels to me that instead of quitting on it, you should double down. The reason we don't want GPU, it's that renting is not priced well enough and the technology isn't quite there yet either for us to make consistently good use of it. Removing the offer is just exacerbating the current situation. It feels both curves are about to meet. In either case you'll have the experience to bring back the offer if you feel it's needed.
- jameslk 2y ago> The biggest problem: developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. Fly.io seems to attract similar developers as Cloudflare’s Workers platform. Mostly developers who want a PaaS like solution with good dev UX. If that’s the case, this conclusion seems obvious in hindsight (hindsight is a bitch). Developers who are used to having infra managed for them so they can build applications don’t want to start building on raw infra. They want the dev velocity promise of a PaaS environment. Cloudflare made a similar bet with GPUs I think but instead stayed consistent with the PaaS approach by building Workers AI, which gives you a lot of open LLMs and other models out of box that you can use on demand. It seems like Fly.io would be in a good position to do something similar with those GPUs.
- andrewstuart 2y agoNvidia deliberately makes this hard. Opportunity for Intel and AMD.
- wmf 2y agoMy takeaway from the article is that there's no real market for GPU VMs so it's pointless for Intel and AMD to make them work.
- nitwit005 2y ago> The biggest problem: developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. My current company has some finance products. There was machine learning used for things like fraud and risk before the recent AI excitement. Our executives are extremely enthused with AI, and seemingly utterly uncaring that we were already using it. From what I can tell, they genuinely just want to see ChatGPT everywhere. The fraud team announce they have a "new" AI based solution. I assume they just added a call to OpenAI somewhere.
- devmor 2y agoThe emotion in this article hits home to me. There have been several points in my career where I worked hard for long hours to develop a strong, clever solution to a problem that ultimately was solved by something cheap because the ultimate consumer didn’t care about what we expected them to care about. It sucks from a business perspective of course, but it also sucks from the perspective of someone who takes pride in their work! I like to call it “artisan’s regret”.
- khana 2y ago[dead]
- apineda 2y agoMy issue is that I may or may not understand what's going on, but I simply, for the most part, do not want to spend time maintaining any more than I have to.
- ryuuseijin 2y agoMy heart stopped for a moment when reading the title. I'm glad they haven't decided to axe GPUs, because fly GPU machines are FANTASTIC! Extremely fast to start on-demand, reliable and although a little bit pricy but not unreasonably so considering the alternatives. And the DX is amazing! it's just like any other fly machine, no new set of commands to learn. Deploy, logs, metrics, everything just works out of the box. Regarding the price: we've tried a well known cheaper alternative and every once in a while on restart inference performance was reduced by 90%. We never figured out why, but we never had any such problems on fly. If I'm using a cheaper "Marketplace" to run our AI workloads, I'm also not really clear on who has access to our customer's data. No such issues with fly GPUs. All that to say, fly GPUs are a game changer for us. I could wish only for lower prices and more regions, otherwise the product is already perfect.
- raylad 2y agoI just looked at their pricing and they don't list any GPUs at all that I could find.
- ryuuseijin 2y agoSearch for A100 on this page: https://fly.io/docs/about/pricing/ https://fly.io/docs/about/pricing/
- bottega_boy 2y agoI used the fly.io GPUs as development machines. For that, I generally launch a machine when I need it and scale it to 0 when I am finished. And this is what's really fantastic about fly.io - setting this up takes an hour... and the Dockerfile created in the process can also be used on any other machine. Here's a project where I used this setup: https://github.com/li-il-li/rl-enzyme-engineering https://github.com/li-il-li/rl-enzyme-engineering This is in stark contrast to all other options I tried (AWS, GCP, LambdaLabs). The fly.io config really felt like something worth being in every project of mine and I had a few occasions where I was able to tell people to sign up at fly.io and just run it right there (Btw. signing up for GPUs always included writing an email to them, which I think was a bit momentum-killing for some people). In my experience, the only real minor flaw was the already mentioned embedding of the whole CUDA stack into your container, which creates containers that approach 8GB easily. This then lets you hit some fly.io limits as well as creating slow build times.
- imcritic 2y agoWhat a good and open and honest blog post. And I liked a lot the way it is interlinked with other interesting posts from that blog. I wish I'll have some time to read more articles from that blog.
- abraxas 2y agoIf low cost GPUs are not what they are offering then what are they offering anymore that I wouldn't get a big cloud vender. This looks like self inflicted mortal wound.
- tptacek 2y ago"Mortal wound" lol.
- jonathanlei 2y agoIt's as difficult as a serverless provider to grow as it was for CPUs before GPUs came along. Many companies overinvest in fully-owned hardware, rather than renting from clouds. Owning hardware means you underwrite unrented inventory costs and prevents you from scaling. H100 pricing is now lower than any self-hosted option, even without factoring the TCO & headcount. (Disclaimer: I work at a GPU cloud Voltage Park -- with 24k H100s as low as $2.25/hr [0] -- but Fly.io is not the only one I've noticed purchase hardware when renting might have saved some $$$) [0] https://dashboard.voltagepark.com/ https://dashboard.voltagepark.com/
- hansvm 2y ago> The biggest problem: developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. I don't want GPUs, but that's not quite the reason: - The SOTA for most use cases for most classes of models with smallish inputs is fast enough and more cost efficient on a CPU. - With medium inputs, the GPU often wins out, but costs are high enough that a 10x markup isn't worth it, especially since the costs are still often low compared to networking and whatnot. Factor in engineer hours and these higher-priced machines, and the total cost of a CPU solution is often still lower (always more debuggable). - For large inputs/models, the GPU definitely wins, but now the costs are at a scale that a 10x markup is untenable. It's cheaper to build your own cluster or pay engineers to hack around the deficits of a larger, hosted LLM. - For xlarge models™ (fuzzily defined to be anything substantially bigger than the current SOTA), GPUs are fundamentally the wrong abstraction. We _can_ keep pushing in the current directions (transformers requiring O(params * seq^2) work, pseudo-transformers requiring O(params * seq) work but with a hidden, always-activated state space buried in that `params` term which has to increase nearly linearly in size to attain the same accuracy with longer sequences, ...), but the cost of doing so is exorbitant. If you look at what's provably required to do those sorts of computations, the "chuck it in a big slice of vRAM and do everything in parallel" strategy gets more expensive compared to theoretical optimality as model size increases. I've rented a lot of GPUs. I'll probably continue to do so in the future. It's a small fraction of my overall spending though. There aren't many products I can envision which could be built on rented GPUs more efficiently than rented CPUs or in-house GPUs.
- npn 2y ago> The biggest problem: developers don’t want GPUs. They don’t even want AI/ML models. They want LLMs. No, I want GPU. BERT models are still useful. The point is your service is too expensive that only one or two months of renting is enough to build a PC from scratch and place it somewhere in your workplace to run 24/7. For applications that need GPU power, usually downtime or latency does not really matter. And you always add an extra server to ensure.
- burnto 2y agoI think they’re too early for their core market. It’s taking indie and 0-1 devs awhile to dig into ML because it’s a huge complex space. But some of us are starting to put together interesting little pipelines with real, solid applications.
- the_king 2y agoThis is well written. I appreciated the line, "startups are a race to learn stuff."
- kristopolous 2y agoThey were double wrong. I work at a GPU cloud provider and we can't bring on the machines fast enough. Demand has been overwhelming. People aren't going to fly.io to rent GPUs. That's the actual reality here. They thought they could sidecar it to their existing product offering for a decent revenue boost but they didn't win over the prospect's mind. Fly has compelling product offerings and boring shovels don't belong in their catalog
- tptacek 2y agoSure. If it sounds like we're saying "cloud GPUs are not a product anybody wants", absolutely not. They're just not a knockout hit for us.
- tempaccount420 2y agoBut why not add an option to rent them out without too many abstractions?
- tptacek 2y agoBecause that's not what we're in business to do.
- hamandcheese 2y ago> We were wrong about Javascript edge functions, and I think we were wrong about GPUs. Actually, you're still wrong about JavaScript edge functions. CF Workers slap.
- tptacek 2y agoThey were wrong for us. Cloudflare is in a much different position than we were in 2019 trying to get people to write new Javascript. Clearly, for us, running people's existing applications natively was the better call. We're not dunking on Cloudflare's model.
- taeric 2y agoI was at another team making a similar bet. Felt off to me at the time, but I assumed I just didn't understand the market. I also think the call that people want LLMs is slightly off. More correct to say people want a black box that gives answers. LLMs have the advantage that nobody really knows anything about tuning them. So, it is largely a raw power race. Taking it back to ML, folks would love a high level interface that "just worked." Dealing with the GPUs is not that, though.
- zacksiri 2y agoMost developers avoid GPU because of pricing. It’s simply too expensive to run 24/7 and there is the overhead of managing / bootstrapping instances loading large models to do intermittent instances. That’s the gist of it I think. Unless you have constant load that justify 24/7 deployments most devs will just use an API. Or find solutions that doesn’t need you to pay > $1 / hour.
- PeterStuer 2y agoIt also seems they got caught in the middle of the system integrator vs product company dilemma. To me fly's offering reads like a system integrator"s solution. They assemble components produced mainly by 3rd parties into an offered solution. The business model of a system integrator thrives on doing the least innovation/custom work possible for providing the offering. You posotion yourself to take maximal advantage of investments and innovations driven by your 3rd party suppliers. You want to be squarely on their happy path. Instead this artcle reads like fly, with good intention, was trying to divert their tech suppliers offer stream into niche edge cases outside of maistream support. This can be a valid strategy for products very late into their maturity lifecycle where core innovation is stagnant, but for the current state of AI with extremely rapid innovation waves coarsing through the market, that strategy is doomed to fail.
- deleted 2y ago[deleted]
- amelius 2y agoIn most cases developers don't want GPUs, they just want a way to express a computation graph, and let the system perform the computation.
- scosman 2y ago> But inference latency just doesn’t seem to matter yet, so the market doesn’t care. This is a very strange statement to make. They are acting like inference today happens with freshly spun up VMs and model access over remote networks (and their local switching could save the day). It’s actually hitting clusters of hot machines with the model of choice already loaded into VRAM. In real deployments, latency can be small (if implemented well), and speed is comes down to the right GPU config for the model (why fly doesn’t offer). People have built better shared resource usage inference systems for Loras (openAI, Fireworks, Lorax) - but it’s not VMs. It’s model aware, the right hardware for the base model, and optimizing caching/swapping the Loras. I’m not sure the Fly/VM way will ever be the path for ML. Their VM cold start time doesn’t matter if the app startup requires loading 20GB+ of weights. Companies like Fireworks are working fast Lora inference cold starts. Companies like Modal are working on fast serverless VM cold starts with a range of GPU configs (2xH100, A100, etc). These seem more like the two cloud primitives for AI.
- ec109685 2y agoI think what they mean about latency not mattering is that latency to the LLM provider doesn’t matter. So why run it yourself when there are API’s you can hit that provide a better overall experience (and seems to be dropping in cost 90% year over year).
- scosman 2y agoOh that makes more sense. My bad.
- sgt 2y agoCurrently getting a 502 error when trying to access fly.io
- hankchinaski 2y agoThey should invest and focus on making their platform more reliable. Without that they will continue to be just a hobby toy to play with and nothing more
- cschmatzler 2y agoIt’s nice seeing a major outage a day after this.
- sylware 2y agoGPU is all about performance. Nearly all the time, very high level languages have nothing to do there. The CPU part of high level user applications will probably be written in very high level languages/runtimes with, sometimes, some other parts being bare metal accelerated (GPU or CPU). Devs wanting hardcore performance should write their stuff directly in GPU assembly (I think you can do that only with AMD) or at best with a SPIR-V assembler. Not to mention doing complex stuff around the linux closed source nvidia driver is just asking for trouble. Namely, either you deploy hardware/software nvidia did validade, or just prepare to suffer... it means 'middle-men' deploying nvidia validaded solutions have near 0 added value.
- pier25 2y ago"We started this company building a Javascript runtime for edge computing." Wait... what? I've been a Fly customer for years and it's the first time I hear about this.
- Kwpolska 2y ago> Instead, we burned months trying (and ultimately failing) to get Nvidia’s host drivers working to map virtualized GPUs into Intel Cloud Hypervisor. At one point, we hex-edited the closed-source drivers to trick them into thinking our hypervisor was QEMU. What do Nvidia’s lawyers think of this? There are some things that best not mentioned in a blog post, and this is one of them.
- djhworld 2y agoI get the impression that running LLMs is a pain in general, always seems to need the right incantation of nvidia drivers, linux kernel and a boat load of VRAM, along with making sure the Python ecosystem or whatever you are running for inference has the right set of libaries and then if you want multi-tenant processing across VMs - forget it or pay $$$ to nvidia. The whole cloud computing world was built on hypervisors and CPU virtualization, I wonder if we'll see a similar set of innovations for GPUs at commodity level pricing. Maybe a completely different hardware platform will emerge to replace the GPU for these inference workloads. I remember reading about Google's TPU hardware and was thinking that would be the thing - but I've never seen anyone other than Google talk about it.
- flockonus 2y agoIt feels like giving up on this a bit too soon? I mean, they realized the problem quite right.. their offering doesn't entirely makes sense for their audience when it comes to GPU. _But_ the demand of open source models is just beginning. If they really have a big inventory of GPUs under-utilized and users want particular solutions on demand.... give it to them??? Like TTS STT video creation, real time illustration enhancement, deepseek and many others. You guys are great at devops, make useful offerings on demand, similar to what HuggingFace offers, no???
- siliconc0w 2y agoThey might just be early. The smaller models are getting more and more capable, for high-frequency use-cases it'll probably be worth using local quantized models vs paying for API inference.
- cytocync 2y ago[flagged]
- hinkley 2y agoI feel like one of the mistakes being made again and again in the virtualization space is not realizing there's a difference between a competitor possibly running containers on the same machine with your proprietary data, and Dave over in Customer Relations running a container on your same machine. If Dave does something malicious, we know where Dave lives, and we can threaten his livelihood. If your competitor does it you have to prove it, and they are protected from snooping at least as much as you are so how are you going to do that? I insist that the mutually assured destruction of coworkers substantially changes the equation. In a Kubernetes world you should be able to saturate a machine with pods from the same organization even across teams by default, and if you're worried that the NY office is fucking with the SF office in order to win a competition, well then there should be some non-default flags that change that but maybe cost you a bit more due to underprovisioning. You got a machine where one pod needs all of the GPUs and 8 cores, great. We'll load up some 8 core low memory pods onto there until the machine is full.
- bleemworks 2y agoArticle about GPUs, comments arguing over the definition of complexity in Kubernetes. This is what you call “learned helplessness.”
- KETpXDDzR 2y ago> They want LLMs. That's why NVIDIA has NIMs [0]. A super easy way to use various LLMs. [0] https://developer.nvidia.com/nim https://developer.nvidia.com/nim
- olibaw 2y agoWhat a great blog post. Hope you figure it out, Fly.io. "If you will it, it is no dream."
- east4ming 2y ago[dead]