3 ms·
This is the dictionary definition of an unfunded mandate. I'll just say that I'm not seeking a glorped together description of the games in question, but a sem
by textfiles 4y ago
This is the dictionary definition of an unfunded mandate.
I'll just say that I'm not seeking a glorped together description of the games in question, but a semantic listing of the contents of original articles, ads, discussions and mentions related to the titles. LLMs do not do this, and if you're training the LLMs (at cost) to do this, you're already having to do the very same searching out of materials within the corpus related to what you want. It'd be like gathering every TV guide in a stack to then ask it to describe things that are like All in the Family, instead of just going through them as you gather them to indicate what each issue has information about.
- TedDoesntTalk 4y ago[flagged]
- textfiles 4y agoYou've modified your original comment (which I didn't know Hackernews allowed people to do). So now I look completely weird because it "seems" you only typed one word and I wrote something. You wrote, essentially: Why not train a LLM AI on the magazines, newsletters, documentation, etc, and have it generate the descriptions based on prompts? ...and I call this an "unfunded mandate", which is where someone comes along to a project, and tells them a bunch of things they should take on and do, instead of either offering to do something themselves, or pointing to a resource or project that fulfils the need that is being asked for. The article is about me finding out that certain kinds of activities are jammed in a Catch-22 of necessity but lack of potential funding (time or money). Your idea shifts no needle and makes things more complicated, to no greater chance of success, but plenty of chance (contemporarily) of hallucination and misleading results. Feel free to delete, I guess, even more.
- gwern 4y ago> LLMs do not do this, and if you're training the LLMs (at cost) to do this, you're already having to do the very same searching out of materials within the corpus related to what you want. A more reasonable suggestion would be not training a LLM (which one doesn't want to do anyway) but treating it as a retrieval+summarization task: search the corpus for mentions and similar-by-embedding documents, and summarize. LLMs are good at abstractive summarization with minimal hallucination or error. This can serve as an 'annotated bibliography', a first pass for a human writing it themselves, or the collective summaries be fed into the LLM for a summary. The main problem here is I guess that most of the relevant texts have poor or no OCR, so one can't do that in the first place. But there's a good chance that that will mostly stop being an issue in a few years as 'text' LLMs move to images (see eg PIXEL https://arxiv.org/abs/2207.06991 https://arxiv.org/abs/2207.06991 or Kosmos https://arxiv.org/abs/2302.14045 https://arxiv.org/abs/2302.14045 or https://arxiv.org/abs/2010.10648#google https://arxiv.org/abs/2010.10648#google https://arxiv.org/abs/2012.14271 https://arxiv.org/abs/2012.14271 https://arxiv.org/abs/2209.14156 https://arxiv.org/abs/2209.14156 ) and they will either OCR, embed, or just process images of complex text directly. So, something to keep an eye on, perhaps: there's never going to be enough humans to do all this archiving properly, but perhaps there may eventually be enough GPUs to do it...