3 ms·
>was reviewed by independent researchers That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing https://andrewwu.substa
by jrflowers 20d ago
>was reviewed by independent researchers
That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing
https://andrewwu.substack.com/p/the-slop-vestigation-and-ethics-washing https://andrewwu.substack.com/p/the-slop-vestigation-and-eth...
Edit: Does everybody else get no results when searching for ‘slopvestigation’ on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago
- Ylpertnodi 20d ago'Slopping': when you have to buy something you know is poor quality, but if it works...
- derpyzza 20d agodoesn't show up for me either
- talon8635 19d agoIsn’t the use of LLMs to unwind the events evidence of the scope/breadth, and a testament to the complexity and uniqueness of what happened? Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?
- jrflowers 19d ago> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis? Do you think the only thing a person can do on the computer is use a chat bot?
- talon8635 19d agoWell as a programmer who doesn’t really use them, no.
- sensanaty 18d ago> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis? When we have an error or issue in the $WORK codebase on LIVE/PROD, that's precisely what we do. We sit down, analyze the logs for our services over the relevant date ranges and try to piece together exactly what happened and why. We have a huge number of logs too, but thanks to the magic of proper SWE (which you'd think OAI would have with their magic AIs) we've managed to partition our observability tooling so that you can digest only what you need. That's basically how any serious organization does things, instead of just throwing a non-deterministic black box at the problem. Especially because logs are by their very nature noisy, and they will saturate any model's context window very quickly leading to massive hallucinations and what ultimately amounts to making shit up that isn't anywhere in the logs (ask me how I know)