3 ms·
What are companies needing all of these hard drives for? I understand their need for memory, and boot. But storing text training data and text conversations isn
by 55555 8mo ago
What are companies needing all of these hard drives for? I understand their need for memory, and boot. But storing text training data and text conversations isn't that space intensive. There's a few companies doing video models, so I can see how that takes a tremendous amount of space. Is it just that?
- Ekaros 8mo agoHearing about their scrapping practises it might be that they are storing same data over and over and over again. And then yes, audio and video is likely something they are planning for or already gathering. And if they produce lot of video, they might keep copies around.
- red75prime 8mo agoAll the latest general purpose models are multimodal (except DeepSeek I think). Transfer learning allows to improve results even after they exhausted all the text in the internet.
- jmclnx 8mo agoI am surprised by that too. I thought everyone moved to SDDs or NVMe ? I was toying with getting a 2T HDD for a BSD system I have, I guess not now :)
- dominicrose 8mo agoEveryone moved to SDDs or NVMe. If you're right, that includes manufacturers. HDDs still have advantages over SSDs for specific needs, like more reliable long-term unelectrified storage. It's also possible that the high price of SSDs made HDDs an option again.
- pixl97 8mo agoReally if you're writing large solid files hard drives aren't that bad. If you can have the system split out one file per drive at a time then you'll avoid a lot of the fragments
- pixelesque 8mo agoStoring training data: for example, Anthropic bought millions of second hand books and scanned them: https://www.washingtonpost.com/technology/2026/01/27/anthropic-ai-scan-destroy-books/ https://www.washingtonpost.com/technology/2026/01/27/anthrop...
- TiredOfLife 8mo agoAll of Annas archive can be put on 40 drives
- numpad0 8mo agoI think the somewhat hallucinatory canned response is that they distribute data across drives for a massive throughput. Though idk if that even technically makes sense...
- danny_codes 8mo agoSpeaking from personal experience.. we treat cloud storage like an infinitely deep bucket. At rest data efficiency is not really a consideration because compute costs are so absurd. Why worry about a $2M year storage bill when your compute bill is $500M? It’s not worth the engineering time to optimize