4 ms·
Don't even remember when I opened Stack Overflow, won't miss that condescending place.
by sekai 3y ago
Don't even remember when I opened Stack Overflow, won't miss that condescending place.
- the_duke 3y agoJust wait for people to stop using SO, at which point the LLMs won't have a high quality training set for new questions, so you won't get good answers from the LLMs anymore...
- littlestymaar 3y agoDepends on the language, but many things happen on Discord now (which is very annoying since it's not indexable by search engine and you need to ask the question to get the answer…)
- sumitkumar 3y agoThe LLMs are generating training data at a faster rate than SO. All the prompts and the responses will eventually be 99.99% of the training data.
- vorticalbox 3y agodoes this not create a feed back loop, if you're training data based on things the LLM said?
- tinco 3y agoThey're probably generating based on GitHub code. If I were training a code model I'd take a snippet of code, have the existing LLM explain it. Then use the explanation and the snippet for the test data.
- DSingularity 3y agoSurely you are joking. You want us to rely on models that are overfit to hallucinated LLM interactions.
- bongobingo1 3y agoJust open enough issues on the parent libraries that they give up and conform to the hallucinations.
- clbrmbr 3y agoI’ve been doing this in my private codebase. When copilot hallucinates a function, I just go and write the thing. It’s usually a good idea, and it will re-hallucinate the same function independently in another file.
- the_duke 3y agoThe only way this is useful in the context of code is if: * The LLMs have a sufficient "understanding" of the request and of how to write code to fulfill the request * Have a way to validate the suggestion by actually executing the code (at least during training) and inspecting the output From what I've seen we are still far away from that, Copilot and GPT-4 seem heavily reliant on very well-commented code and on sources like Stackoverflow
- hobabaObama 3y agoLLMs also train on official documentations which is where 90% of problems get solved.
- m_fayer 3y agoWhat will happen to official docs when it becomes clear that the only thing that reads them are llm-training runs?
- terhechte 3y agoThe LLMs will read the actual source code which is way better than the documentation (as any iOS engineer will tell you). For private codebases the companies can provide custom-trained LLMs. Techniques like "Representation Engineering" will at some point also prevent against accidental leakage of private codebase source code.
- tiborsaas 3y agoCall it a win?
- m_fayer 3y agoWon't you think of all the technical writers?!
- RamblingCTO 3y agoIn what world are you living in? That's maybe true in noob land. Literally all the problems I have are being solved in github issues, if at all. When has documentation been 90% sufficient for anything? In the 80s? /e: sorry, sounds a bit stand off-ish. Let me give an example: I was trying to find a way to clone a gorm query to keep the code clean. The documentation doesn't have anything (no, .Session isn't a solution) and the only place I had was issues discussing that. Apparently you can't. So I'll be ditching gorm and move to pgx in the near future. That's how it happens for me all the time. The documentation is lacking the hard part, always.
- hackerlight 3y agoWe will figure out synthetic code data by then.
- dcow 3y agoSO: the community that optimized for moderator satisfaction over enduser utility.