3 ms·
I've been wondering about this problem from a non-expert position for awhile. How can we capture, say the entire hard-information - or semantic latent space; of
by kfrzcode 4y ago
I've been wondering about this problem from a non-expert position for awhile. How can we capture, say the entire hard-information - or semantic latent space; of a given body of work. Pick, say, the Simpsons universe. Could we have an AI consume all of the screenplays, scripts, fandom wikis, etc, and create a pre-trained model of every considerable aspect of the fictional universe? That pre-trained model would just be a distillation of the informational content and built to interact in both plain-language, ChatGPT style interface but also more technically and output JSON for queries. Say I want to list every episode of the show that includes a theme of love always ending in tragedy, except of course for Marge and Homer. Or I want to define a histogram of which illustrators worked when which voice actors did a specific set of characters.
Obviously a more useful example would be domain-specific knowledge in industry or medicine, but a generalized approach to ontological encoding from a given dataset would probably require a lot of interesting techniques and math that's way beyond my head.
But that's probably something easily done with current technology, what I'm interested in is learning how to talk about/learn about this concept. Distilling "knowledge," or "concepts" into "parameters", like, defining a DNA-like code for a given corpus of data... sorry I'm rambling but hopefully someone can relate
- reidjs 4y agoSemi-related: One of the first things I tried with 'GPT' was training it on every Simpson's screenplay and then write its own. Example: https://github.com/reidjs/simpsons-gpt/blob/master/output/output_t_0.8_k_30.txt https://github.com/reidjs/simpsons-gpt/blob/master/output/ou... The results were... not great. Maybe on the new GPT iteration it will be more successful though.
- deleted 4y ago[deleted]