5 ms·
Fiber networks were using less than 0.002% of available capacity, with potential for 60,000x speed increases. It was just too early. I doubt we wil
by mg 1y ago
Fiber networks were using less
than 0.002% of available capacity,
with potential for 60,000x speed
increases. It was just too early.
I doubt we will see unused GPU capacity. As soon as we can prompt "Think about the codebase over night. Try different ways to refactor it. Tomorrow, show me your best solution." we will want as much GPU time at the current rate as possible.
If a minute of GPU usage is currently $0.10, a night of GPU usage is 8 * 60 * 0.1 = $48. Which might very well be worth it for an improved codebase. Or a better design of a car. Or a better book cover. Or a better business plan.
- cantor_S_drug 1y agoWith improvements on the algorithm side and new techniques, even older hardware will become useful.
- Zigurd 1y agoI get what you're saying and the reasoning behind it, but older hardware has never been useful where power consumption is part of determining usefulness.
- chatmasta 1y agoThis is the biggest threat to the GPU economy – software breakthroughs that enable inference on commodity CPU hardware or specialized ASIC boards that hyperscalers can fabricate themselves. Google has a stockpile of TPUs that seem fairly effective, although it’s hard to tell for certain because they don’t make it easy to rent them.
- xadhominemx 1y agoMore efficient inference = more reasoning token. Hyperscaler ASICs are closing the gap at the hardware/system level, yes.
- Zigurd 1y agoI don't think we will need to wait for anything as unpredictable as a breakthrough. Optimizing inference for the most clearly defined tasks, which are also the tasks where value is most readily quantified, like coding, is underway now.
- Mo3 1y ago> I doubt we will see unused GPU capacity I'd argue we very certainly will. Companies are gobbling up GPUs like there's no tomorrow, assuming demand will remain stable and continue growing indefinitely. Meanwhile LLM fatigue has started to set in, models are getting smaller and smaller and consumer hardware is getting better and better. There's no way this won't end up with a lot of idle GPUs.
- xadhominemx 1y agoTest time compute has made consumption highly elastic. More compute = better results. Marginal cost of running these GPUs when they would otherwise be idle is relatively very low. It will be utilized.
- delusional 1y ago> There's no way this won't end up with a lot of idle GPUs. Nvidia is betting the farm on reinventing GPU compute every 2 years. The GPUs wont end up idle, because they will end up in landfills. Do I believe that's likely, no, but it is what I believe Nvidia is aiming for.
- Workaccount2 1y ago>Meanwhile LLM fatigue has started to set in Has it? I think there is this compulsion to think that LLMs are made for senior devs, and if devs are getting wary of LLMs, the experiment is over. I'm not a programmer, my day job isn't tech, and the only people I know who express discontent with LLMs are a few of programmer friends I have. Which I get, but everyone else is using them gleefully for all manner of stuff. And now I am seeing the very first inklings of completely non-technical people making bespoke applets for themselves. From OpenAI, programming is ~4% of chatGPTs usage. That's 96% being used for other stuff. I don't see any realistic or grounded forecast that includes a diminishing demand for compute. We're still at the tip of adoption...
- Mistletoe 1y agoYou should get on Reddit, people hate AI with a passion there. People I meet in real life hate it also. I think the public actually hates AI more than it should now.
- jdlshore 1y ago> As soon as we can prompt… This is the fundamental error I see people making. LLMs can’t operate independently today, not on substantive problems. A lot of people are assuming that they will some day be able to, but the fact is that, today, they cannot. The AI bubble has been driven by people seeing the beginning of an S-curve and combining it with their science-fiction fantasies about what AI is capable of. Maybe they’re right, but I’m skeptical, and I think the capabilities we see today are close to as good as LLMs are going to get. And today, it’s not good enough.
- Workaccount2 1y agoGetting gold in the math Olympiad is a pretty strong indicator of operating independently on substantive problems. A year ago they need an extensive harness to get silver, and two years ago they could hardly multiply 1000x10000. Terence Tao tweeted yesterday about using GPT5 to help quickly solve a problem he was working on.
- saberience 1y agoYes but why did ChatGPT work on math Olympiad problems? Because it got a prompt giving it the instruction and context etc. Why did GPT5 help Terence Tao solve a math problem, because he gave it a prompt and the context etc. None of these models are useful without a human prompting them and giving it tasks, goals, context etc, they don't operate independently, they don't get ideas of work to be done, they don't operate over long time horizons, they can't accept long term goals and sub-divide those goals into sub goals, and sub tasks etc. They are useless without humans telling them what to do.
- Workaccount2 1y agoYou should see what happens when you let them talk to each other
- WA 1y agoErrors compound? Context drift?
- skrebbel 1y ago> improved codebase I've seen lots of claims about AI coding skill, but that one might be able to improve (and not merely passably extend) a codebase is a new one. I'd want to see it before I believe it.
- Leynos 1y agoIt depends what you're fitting to. At the simplest, you can ask for a reduction in cyclomatic/cognitive complexity measured using a linter, extraction of methods (where a paragraph of code serves no purpose other than to populate a variable) or complex conditionals, move from an imperative to a declarative approach, etc. These are all things that can be caught through pattern matching and measured using a linter or code review tool (CodeRabbit, Sourcery or Codescene). Other things might need to be done in two stages. You might ask the agent to first identify where code violates CQRS, then for each instance, explain the problem, and spawn a sub-agent to address that problem. Other things the agent might identify this way: multiple implications, use of conflicted APIs, poor separation of concerns at a module or class level. I don't typically let the agent do any of this end to end, but I would typically manually review findings before spawning subagents with those findings.
- fragmede 1y agoClaude will refactor but more than that, it can add documentation. And it can be asked about a codebase too. "Where does FOO happen?" "How does BAR work?".
- thenaturalist 1y agoThis is such a short sighted take glaringly ommitting a crucial ingredient in learning or improvement - both for humans or machines alike: feedback loops. And you can't really hack / outsmart feedback loops. Just because something is conceptually possible, interaction with the real rest of the world separates a possible from an optimal solution. The low hanging fruits/ obvious incremental improvements might be quickly implemented by LLMs based on established patterns in their training data. That doesn't get you from 0 to 1 dollar, though and that's what it's all about.
- bwfan123 1y agothis. Was highlighted by Sutton in a recent podcast rather starkly. LLMs are a great tool. But, the real world is far too nuanced to be captured in text and tokens. So, LLMs will be a great productivity boosting tool like a calculator or a spreadsheet. Expecting it to do more is science fiction.
- yubblegum 1y agoI just had to double check (have not been paying attention for a couple of years) but indeed it seems GPU underutilization remains a fact and the numbers are pretty significant. Main issues are being memory bound so the compute sits idle.
- davedx 1y agoTasks being memory bound is not the same thing as GPU's being idle for economic reasons though.
- ninkendo 1y agoThe actual computation speed isn't as important nowadays but it doesn't really change the conclusion with respect to whether they're underutilized. Because the main reason for the price premiums in AI-class GPUs are the gobs of insanely fast memory, and that is very much not underutilized. AI companies grab GPUs with as much memory (at the fastest memory bandwidth) as possible and underclock the GPU to save on power. Linus Tech Tips had a great video about the H200 that touched on this this week: https://www.youtube.com/watch?v=lNumJwHpXIA https://www.youtube.com/watch?v=lNumJwHpXIA
- credit_guy 1y ago> As soon as we can prompt "Think about the codebase over night. Try different ways to refactor it. Tomorrow, show me your best solution." we will want as much GPU time at the current rate as possible. That is nothing. Coding is done via text. Very soon people will use generative AI for high resolution movies. Maybe even HDR and high FPS (120 maybe?). Such videos will very likely cost in the range of $100-$1000 per minute. And will require lots and lots of GPUs. The US military (and I bet others as well) are already envisioning generative AI use for creating a picture of the battlespace. This type of generation will be even more intensive than high resolution videos.
- bigbadfeline 1y ago> "Try different ways to refactor it. Tomorrow, show me your best solution." The cost/benefit analysis doesn't add up for two reasons: First, a refactored codebase works almost the same as non-refactored one, that is, the tangible benefit is small. Second, how many times are you going to refactor the codebase? Once and... that's it. There's simply no need for that much compute for lack of sufficient beneficial work. That is, the present investments are going to waste unless we automate and robotize everything, I'm OK with that but it's not where the industry is going.
- ccorcos 1y agoI’ve never understood why time is the metric people are using here. If LLMs get so much better we can “run them overnight”, what makes you think that they won’t also get faster and so they accomplish exactly what you’re talking about in 5 minutes?