10 ms·
Show HN: Goopt – Search Engine for a Procedural Simulation of the Web with GPT-3
- zuzun 5y agoSo basically modern Google without ads.
- kelseyfrog 5y agoI love this. It's the reification of the Dead-Internet Theory - a tangible artifact embodying the feeling that the internet was replaced by its own simulacrum powered by AI.[1] The existence of Goopt is the culmination of DIT as self-fulfilling prophecy. We can almost see the beginnings of an outline begin to form around an Internet Turing test. How well can we discern the real internet from the fake one. Consequently, what happens when the line becomes so blurred that we lose the ability to perceive the difference? 1. https://www.theatlantic.com/technology/archive/2021/08/dead-internet-theory-wrong-but-feels-true/619937/ https://www.theatlantic.com/technology/archive/2021/08/dead-...
- joken0x 5y agoI think there will come a point when the traditional internet will also become diluted in artificiality, content from the procedural web will start to creep into the traditional web, and we won't be able to distinguish. It's interesting to think of it in terms of Baudrillard's Simulacra and Simulation.
- kelseyfrog 5y ago> I think there will come a point when the traditional internet will also become diluted in artificiality There is already a concern in corpora creation in ML/AI projects. Researchers would like very much to have human-only generated content when training models on internet-sourced text. People posting GPT-created output has the potential to taint these corpora and create all sorts of strange loopy feedback.
- rm_-rf_slash 5y agoKnowing people it wouldn’t take long for incoherent generative syntax to become a meme and reinforce the human corpora with new syntactic slang.
- joken0x 5y agoThat's right, the human creation is and will become more and more a very precious treasure.
- jerf 5y agoWe're passed that point. Probably by a couple of years at a minimum. No sarcasm. Content farms are definitely using techs like this. GPT-3 is really good at generating text but still has some characteristic failures, and I encounter content farm web pages (despite my best efforts) that have clearly used it or something like it as a tool. Even just in the last couple of weeks I've been seeing some new, innovative content farms managing to pollute my search results that I've not seen before.
- zitterbewegung 5y agoIf you put your OpenAI key and start running this then they will ban your account because it will be against their TOS. With some minor modifications you could port it to goose.ai and it isn’t against their TOS. EDIT: Forking it here https://github.com/zitterbewegung/Goopt https://github.com/zitterbewegung/Goopt to add the functionality above.
- dschnurr 5y agoOpenAI engineer here. Cloning this project and running it locally with your own API key does not violate any policies. However the way this project is configured, publishing it to the web would expose your API key in the client-side source code, which violates our policies since it would allow your account to be compromised.
- zitterbewegung 5y agoSorry I should have read the TOS
- FrenchDevRemote 5y agohey, you beat me to it :D did you manage to get it to work? I tried on my own but got error messages
- yoland68 5y agoYou might have to enable billing, that was the issue for me
- zitterbewegung 5y agoNo it’s a config error.
- zitterbewegung 5y agoWhat did you change ?
- fudged71 5y agoThis is an incredible idea. There are so many unexplored possibilities when you re-write and re-format the web. Can you mix procedural and static content? How can you verify accuracy of information? What if you could refine a web page’s content just-in-time? Modifying the query and context etc. Through a lens of Roam/Notion: what if everything were a block that could be individually linked? what if every block could be edited by anybody? what if anyone could add links and annotations across pages? a blend of web and wiki?
- joken0x 5y ago* Can you mix procedural and static content? Just the idea of the wiki is interesting here. Perhaps there could be a wiki that stores content in a static way, that is edited by users putting the best content they find on the procedural web. It would be a valuable place to find good ideas or ideas that we might not have thought of but someone else did. This could also serve as feedback for AI models. Although it is also true that we would not be able to distinguish if non-human opinions start to creep in and end up contaminating the site. * How can you verify accuracy of information? I think this is one of the main difficulties, as the AI would have to understand contexts and have a notion of truth, I think this would already start to touch the capacity of "consciousness". * What if you could refine a web page’s content just-in-time? You will be able to do this for every part that you don't like enough and want something better, or just to see something different.
- marmarama 5y agoIf GPT-3 can produce procedurally generated web content this convincing, search engines are screwed, right? We won't be able to find anything useful on any current search engine because there's no straightforward algorithmic way to tell useful content from endless link farms full of utterly convincing but totally useless content.
- moffkalast 5y agoAt least we can still use Google to search Reddit.
- jay00 5y agoUntil all reddit posts will be gpt-3 generated.
- kelseyfrog 5y agoIt should be possible to train an upvote prediction model conditioned on submission title. This could then be used to optimize GPT-3-family models to produce text which had the highest predicted upvote response. It's a couple-weekend project and I'd be surprised if an AI-hobbiest hadn't done it already.
- jazzyjackson 5y agoIn the trivial case, karma farming bots just keep a database of all Reddit history (it is a public dataset, few hundred gigabytes) and repost the top comments (top threads even) whenever they detect a reposted link (extra points for similarity / reverse image searching) It’s a project I have on the back burner to analyze Reddit history to check what ratio of comments are actually original, and I’d like to build a link aggregator that sorts by novelty.
- kingcharles 5y agoI've thought about this too, and the fact that I've not seen such a bot so far is pretty unbelievable. It's not a huge amount of work to code it. Working across the whole of Reddit (or HN for that matter), it would gather an ungodly amount of karma (and awards) in a small amount of time.
- Geee 5y agoThis is clearly the future. All information will be generated on the fly and tailored for you. AI can match your level of knowledge, your language, your preferred style etc. AI can simplify / extend topics on demand, and also generate illustrations and videos to help explain topics. I think most of the current form of pregenerated web with search actually becomes completely unnecessary, and it'll basically stop existing.
- joken0x 5y agoYou get it, man, that is exactly the question. It is time to think about the many possibilities, problems, dilemmas, paradoxes, etc. It is very interesting and disturbing at the same time.
- Geee 5y agoIt changes everything. Thanks for coming up with this. I have thought about AI generated content before but not in this way. I just realized that we don't need the content web; we need just raw data sources and AI that generates content on the fly, for the user. The AI works for and is directed by the user; that's why it actually can reduce gibberish and make information more accessible and useful. This sets it apart from the current crop of content generation bots.
- debdut 5y agoIt's so cool someone made this, but > The procedural web will be the future of the web. It will offer us infinite content Yup it'll be infinite "garbage"
- ushakov 5y agothe current web isn't to far off with SEO articles and ad-video autoplay
- joken0x 5y agoExactly, it is inevitable, the traditional web will be diluted with the same garbage.
- robbedpeter 5y agoNot necessarily - federated media, webs of trust, and diligent curation across many smaller communities could allow for something that replaces Twitter, reddit, and centralized media hubs. Search within that context is easier - p2p/torrent streaming with crypto incentivized seeding can scale distribution. The current state of adtech and near total surveillance isn't sustainable as more people wake up to the downsides, and as fake crap begins to accumulate. Decentralization of social media, advertising, e-commerce and other web 2.0 staples will be a natural evolution of technology. The story goes "under Google's model of the walled garden web, SEO, spam, and bots achieved parity in all content metrics except actual meaningfulness to the user." Despite having all the compute and talent you could possibly bring together, Google is failing to uphold its core technology. They incentivized bad faith behavior, and are reaping the consequences of that. The acceleration of seo hacking and artificial worthless content is asymmetrical to the acceleration of the capabilities and market model Google has created. A search engine can navigate self selected communities, human curated lists, and creatively bundle lists of lists to achieve high quality results based on actual humans self selecting and acting in their own interests. You can do things with higher quality classification and even provide regex over crawled data without huge technical barriers. Search agents will come about, whether locally or cloud hosted, and will eventually replace centralized engines like Google. There are non doomed visions of the future. Maybe we won't suffer a digital trashocalypse.
- chrisgp 5y agoDoes the full version of this require Strong AI to truly replace the internet? What level of AI is necessary to convincingly replicate human understanding and explanation of information?
- joken0x 5y agoIt is something that is still not clear to me, I see the difficulty of the task but also the rapid evolution of AI models. Maybe it will surprise us in not too long.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- mattnewton 5y ago> The procedural web will be the future of the web. Isn't the "procedural web" built of mountains of (hopefully) human written content? How will the system get content about new subjects without the humans writing it? Isn't a system like GPT-3 currently limited to reflecting the ground truth data it has seen?
- lumost 5y agoFor how long? Think of the marketing and censorship opportunities when you can directly tune not just the content that gets seen but also the content itself! Content is still at least somewhat robust to censorship as it's sometimes difficult to remove all references to a banned book. Imagine if banning content also automatically rewrote all references such that they no longer made reference to the content? Or if one could simply pay and have all reviews of a mediocre book changed to make it the greatest book ever? Note the above is a statement on some of the risks to a procedural web. Not a real market opportunity.
- visarga 5y ago> Isn't a system like GPT-3 currently limited to reflecting the ground truth data it has seen? This limitation went away recently. A variant called RETRO (Retrieval-Enhanced Transformer) can use a search engine to take in the exact information up to date [1], assuming you can curate your own text corpus. It's also 25x smaller. [1] https://deepmind.com/research/publications/2021/improving-language-models-by-retrieving-from-trillions-of-tokens https://deepmind.com/research/publications/2021/improving-la...
- ameminator 5y agoUnfortunately, does not come with a Gwenyth-Paltrow-scented candle.
- sqs 5y agoHaha, this is an amazing concept. It feels like a satire or an art piece. I love it, but it kind of gives me the "is the world real?" feeling.
- joken0x 5y agoSimulations on simulations. How to distinguish the real?
- randomstring 5y agoCuil was way ahead of its time. https://news.ycombinator.com/item?id=1255122 https://news.ycombinator.com/item?id=1255122 http://cuiltheory.wikidot.com/ http://cuiltheory.wikidot.com/ Cuil Theory
- TehCorwiz 5y agoI haven't thought of that in years. I gave up my crusade to revive use of the interrobang(‽) in writing a while ago. After reviewing the materials I see that Cuil Theory has come a bit further since I last read. I believe that Goopt would be somewhere around -2‽ from Cuil theory itself, negative because it's literal reality, but distant because it's an abstract embodiment. Slightly off-topic. During my cursory reading I see that imaginary Cuil got fleshed out. I'd like a second opinion. The way it reads to me is that 'i‽' is almost the literal definition of solipsism.
- aghilmort 5y agovery cool / congrats! was recently tweeting about same -- the potential to help humans search the existing web better beyond just keyword search, i.e., query rewriting, summarizing and extending existing content, etc. there's also the flip side around SEO spam, which is partly why founded Breeze, a newish topic search engine that leverages curation to hedge against the dark side of human / bot spam, etc. bottom line, love this, having worked with GPT-3 in past and the direct impact on day job, all things search
- joken0x 5y agoThanks man! It's a good idea, maybe a similar filter or curator will be needed for the procedural web, but for the dark part of the AI; disinformation, meaningless content, etc.
- skybrian 5y agoIn the event you think you're looking at a simulation of the Internet, maybe start out by checking if news, maps, and weather are realistic, to see how good their world simulation is. Live news video should be interesting too. But if your browser is compromised so that encryption doesn't work, I think you have bigger problems.
- joken0x 5y agoThe concept of simulation is very interesting, Baudrillard's Simulacra and Simulation and other related theories will gain more strength and meaning.
- amznbyebyebye 5y agoIs there any use of ML to distinguish the AI generated dead internet from the real one?
- deleted 5y ago[deleted]
- visarga 5y agoFor generated headlines humans are a coin toss, can't tell them apart. But transformers can reach 85% accuracy. https://aclanthology.org/2021.nlp4if-1.1.pdf https://aclanthology.org/2021.nlp4if-1.1.pdf
- gernb 5y agoIs there any way to get access to a GPT-3 like API that can be run locally (color me ignorant, I know generating net for GPT-3 is huge but I have no idea how small the usable result can be stored so that usage can happen locally instead of to some cloud server
- sbierwagen 5y agoPublic models like GPT-NeoX-20B need a minimum of 45GB of VRAM. That's two 3090s, (Maybe four, five grand, depending on how much effort you spend on bid sniping ebay auctions) or a single A100 80GB. ($20,000+) Also note that NeoX-20B is pretty good, but it's not GPT-3 quality.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- joken0x 5y agoGoopt is an experiment in what the "procedural web" could be. This new web will use procedural content generation to create varied content, completely synthetic, since these are generated by algorithms and artificial intelligence. For now, the content is only text that is being generated automatically with GPT-3, the recent OpenAI model. Goopt works as if it were a search engine, allowing us to search for any term, obtain related results and access their content. Simulating in this way the experience of browsing the web. I started working on this because I think the idea of the procedural web is interesting, I think this technology will come in the future maybe not too far away, so we have to start thinking about the possibilities, problems, dilemmas, paradoxes. Well, I think this is a big change, which has a direct impact on the information we consume and how we do it, the online experience, it could even shape our behavior as it is our information, reference, entertainment system, etc. We must be ready for a change of this size, so I invite you to give life to this topic and put it up for discussion. What is Procedural Web? This is a term that has not yet been treated as such, so I have tried to give my vision on the matter. The procedural web will be the future of the web. It will offer us infinite content, since it will not be necessary that someone has written or created it before. All the content will be synthetic and generated at the moment, with infinite possibilities. From informative text, articles, images, videos, to games, applications and services with interfaces and functionality that are automatically generated. All this adapted to our queries and needs, and increasingly personalized to our preferences. Web 4.0 could be the propitious evolution for the procedural web. The automation that this new version of the web poses about applications, services, interfaces, APIs, devices and others, could be exploited by connecting them with the procedural web. These reality data interfaces would help empower and enhance their generative capabilities, as well as connect their functionality to the real world, having services and devices on which to execute actions. This makes the procedural web even more interesting, because it endows it with cybernetic capabilities. The procedural web is based on natural language processing (NLP) and procedural content generation (PCG). Advances in these fields, as well as in the field of computational creativity, will allow us to generate increasingly better synthetic multimedia, thus nurturing the procedural web with more and better content. It will be interesting to see if this content comes to satisfy us more than the traditional web and human creation. Maybe one day we won't ever be able to distinguish. Demo video and usage guide in the GitHub repository: https://github.com/jokenox/Goopt https://github.com/jokenox/Goopt
- npunt 5y agoProcedural generation is going to be fantastic for personalizing language and explaining ideas; it's pretty obviously the future for anything written that is learning & information oriented. However I'm concerned these personalization systems will be too accommodating and just tell people what they want to hear. One way people use search is for motivated reasoning: I believe something, I search for it, I find confirming evidence. A procedurally generated system I imagine is especially prone to this kind of massaging of queries to get desired outputs - tweak a few input parameters in the form of a query, and out pops the answer you seek. It's a hard problem to test for. Funny thing is when people search to find what they want to hear, they're often driven to content farms, and those are increasingly ML-generated. Seems this is just cutting out the middlemen!