3 ms·
It will be interesting to see what the government can do here. Can they use their powers to get their hands on the most data? im still skeptical because new te
by Footnote7341 3y ago
It will be interesting to see what the government can do here. Can they use their powers to get their hands on the most data?
im still skeptical because new techniques are going to give an order of magnitude efficiency boost to transformer models, so 'just waiting' seems like the best approach for now. I dont think they will be able to just skip to the finish line by having the most money.
- tyingq 3y agoIf not "the most" data, they may have the most access to data that's exclusively available to them.
- phkahler 3y agoThat seems like a good reason for them to do this. I wonder how much non-public stuff they have, or it's just meant to incorporate a specific kind of information.
- raccoonDivider 3y agoI just realized that the NSA has probably been able to train GPT-4 equivalents on _all the data_ for a while now. We'll probably never learn about it but that's maybe scarier than just the Snowden collection story because LLMs are so good at retrieval.
- dwaltrip 3y agoHoly shit, you are right. They probably have 10-100x the data used to train gpt-4. Decades of every text message, phone call transcript, and so on. I can’t believe I haven’t seen anyone mention that yet. People keep saying we don’t have enough data. I think there is a lot more data than we realize, even ignoring things like NSA.
- dwaltrip 3y agoApparently there are roughly 2 trillion text messages sent per year in the US [1]. I did a sanity check, that’s like 40 or so a day per person, so sounds reasonable. I couldn’t find the average message length, but I would guess it’s fairly short (with a fat tail of longer messages). To make the math easy, let’s say the average length is ~10 tokens. I’d be surprised if that isn’t correct within a factor of 2 or so. So we have 20 trillion tokens per year from text messages in the US alone. And this is high-quality conversational data. The annual numbers were fairly constant in recent years (and then it drops off), so the past decade of US text messages is about 200 trillion tokens! That’s a metric fuck ton… Much larger than any dataset existing models have been trained on, I believe. I would guess phone transcripts would be an order of magnitude larger at least. Talking is a lot easier than typing on a phone. You could train an absolutely insane model with that amount of data… Damn. [1] https://www.statista.com/statistics/185879/number-of-text-messages-in-the-united-states-since-2005/ https://www.statista.com/statistics/185879/number-of-text-me...
- rrrrrrrrrrrryan 3y agoInterestingly, the reason Google initially created it's Google Voice service back in the day was to gather voicemail audio to train its speech to text engines. It's mind-blowing to me that with all of Google's data, Google isn't the far and away leader in this new space. I have to believe they're paralyzed by the fear of legal repercussions.
- rvba 3y agoWhen chatGPT came out I asked it "which living politician is an unknown spy" and the question got blocked
- deleted 3y ago[deleted]