4 ms·
> the thing that scares me about the existence of this data is that it seems well within the capabilities of current technology to train a model that can replic
by Chance-Device 2y ago
> the thing that scares me about the existence of this data is that it seems well within the capabilities of current technology to train a model that can replicate me, in some sense
There’s already all of your posts on social media accounts, all your emails on various servers, all of your text messages, all the notes you’ve written anywhere in any form that might end up in some database in the future.
It does make me wonder how much of a person could be inferred by an LLM or future AI from that data. It would never be enough though, I think, to do it properly. There are too many experiences and knowledge you have that might influence what you write without being directly expressed.
Will all of our content end up in some database in the future, and someone decides to make agents based on what they can link to specific identities? Interesting thought.
- brookst 2y agoSeems likely, given the number of models already fine tuned on notable historical figures. Not sure if it makes it better or worse that most of us are probably mostly useful as virtual focus groups / crowds rather than any particular interest in you or me as individuals.
- darknavi 2y agoThat is interesting. I can imagine replicating my speaking/typing mannerisms quite well if I think about stages in my life. Maybe a yearly snapshot, so I could talk to my self as a teen, college student, early professional, etc.
- javajosh 2y agoThat only captures your output, not your input. The best people to simulate in this world would be so-called terminally online people virtually all of whom's input is itself online. So for those who've read a lot of paper books or done a lot of traveling or had a lot of offline conversations or relationships, I think it would be difficult to truly simulate someone.
- visarga 2y agoI think aggregate information across billions of humans can compensate. It would be like a human personality model, that can impersonate anyone. How do you train such a model? Simple - Collect texts with known author and date. They can be books, articles, papers, forum and social network comments, emails, open source PRs, etc. Then assign each author a random ID, and train the model with "[Author-ID, Date] Text", and also "Text [Author-ID, Date]". This means you have a model that can predict authors and impersonate them. You can simulate someone by filling in the missing pieces of knowledge from the personality model. Currently LLMs don't learn to assign attribution or condition on author. A whole layer of insight is lost, how people compare against each other, how they evolve over time. It would allow more precise conditioning by personality profile.
- gavmor 2y agoWhile I agree somewhat with my sibling comment's assertion that "aggregate information across billions of humans can compensate", somewhat, I'd like to offer that a lot of important output is non-digital, as well! For example, lately I've spent a lot of time with resin printers, laser cutters, vacuum chambers, and the meaningful positioning of physical models on large sheets of paper. It'll be a while yet before my haphazard, freewheeling R&D methods are replicable by robots. (Although it's tough to measure the economic value of these labors.)
- ericjmorey 2y agoThe vast majority of my communication is not in text. Most of what I have written or typed is not in anyone's database. I'm not sure how that compares to others.
- deadbabe 2y agoA fatal assumption people make about a person’s online corpus is that the person is actually expressing their true thoughts and personality instead of a LARPed version of themselves that deliberately acts more inflammatory to get engagement. If the person is not being genuine, you will not simulate their true personality and interests, you will be simulating their character. Most people are probably not genuine, except in their one on one conversations with people they know in real life.
- Chance-Device 2y agoIf they’re consistently “not genuine” in their interactions with others, then what’s the difference? I agree that people will have different presentation depending on context; I don’t speak at home exactly the way that I post on here for example. But you are certainly capturing the same aspect or persona that other people see in that context. Like with most things AI it’s about data, you need copious amounts of data in varying contexts. To really capture a person, you’d probably have to get them to write out their thoughts as well. Including all the ones they would never say.
- crackalamoo 2y agoI made a custom GPT of myself using my blog. It understood who I am, but wasn't able to replicate me very well, and mostly sounded like generic ChatGPT with some added interest in my interests. I would imagine fine tuning with enough data would be different though
- WA 2y agoYour output is the map, the map of your experiences. If you make a map of the map (by training a model on your output / the map), this is two abstractions away from the human being experiencing the world with all the errors and uncertainties encoded in both maps.
- SJC_Hacker 2y agoProlific authors, such as Hitchens (passed away in 2011) have been convincingly by duplicated by AI. https://www.youtube.com/watch?v=0qIdEteK0VE https://www.youtube.com/watch?v=0qIdEteK0VE
- phito 2y agoHow could you tell without being very close to him personally?
- HPsquared 2y agoThe model would also need all your personal inputs: everything you've ever seen or heard etc.
- Chance-Device 2y agoNo it wouldn’t, it’s not like ChatGPT needed to be trained on everything all the Kenyan RLHF’ers ever saw or heard.
- ForTheKidz 2y agoOur social media personas are also tiny subsets of our actual personalities. I think most people don't reveal their full character in any one medium. now phone conversations—that would be a goldmine and a nightmare.
- Chance-Device 2y agoHmm. Good thought. Perhaps we could, for strictly educational reasons, set up some agencies who could collect these phone conversations. Maybe the Chronicle & Information Agency? Or the National Scholarship Agency?
- bflesch 2y agoone should operate under the assumption that all phone conversations are being recorded and will be stored eternally. with today's technology they are automatically converted to text, and some palantir-rebranded chatGPT model ranks it in different categories such as as counter-intelligence, organized crime, or terrorism. This is state of the art and certainly done on a national scale by someone (with or without approval of your own government).
- ForTheKidz 2y agoAgreed. I didn't really have the nightmare of "what if the government is impersonating me with hundreds or thousands of hours of phone conversation" until this thread, though.
- TechDebtDevin 2y agoI can clone your voice ezpz, without much expertise with a 10 min phone recording. Some banks still use "my voice is my password" for authentication. Crazy.
- djfivyvusn 2y ago[dead]
- jocoda 2y agoI'm comforted by the thought that like me, most people with nothing to say are determined to let us know that. So I was never really bothered by the belief that everything online is being stored somewhere, because I was certain that there was so much crap to wade through that no one could make any use of it. Not so sure about that any more...
- sharpshadow 2y agoI’ll hope so. Much worse would be if one’s public content gets erased from history. I fear the loss of original sources when LLMs get placed in between. We already have the unresolved issue that the training data is partly illegal and can’t be published. Accessing information through LLMs is much more efficient and is great progress but it’s build in to censor parts of the source information and likely the censored information is lost in transition. Somehow there should be a global data vault initiative, where at least the most important information about our human endeavour is stored. It gives me a chill down my spine when I hear that content from the internet archive is being deleted on request erased from history..
- 34679 2y agoThe first time I changed a system prompt, I changed it to "You are George Carlin." So, I think we're already doing that, in a way.
- physicsNolike 2y ago[dead]