4 ms·
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence. You don't just get to subpoena your neighbor
by fc417fc802 29d ago
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.
You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.
- Arodex 29d agoBut I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career. Any other argument, fc417fc802?
- tristanj 29d agoYou're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it. See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot https://en.wikipedia.org/wiki/Russell%27s_teapot
- magicalist 29d ago> You're making a classic a burden-of-proof fallacy This is incorrect, and you invoke Russell's teapot incorrectly too. It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation. But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false. Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible. This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.
- tristanj 28d agoNo, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set. First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster. Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized. Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one. Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list. You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from. This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
- Arodex 28d agoThen OpenAI should acknowledge that they can't prove they solved the problem independently, and credit the external researchers. It cuts both ways: if OpenAI really needs to access user data, even anonymised, to improve its models, they have to waive any pretention to solve "independently" any problem other people worked on with its tools. Otherwise they (OpenAI) have to firewall/cleanroom themselves.
- fc417fc802 28d ago
- Arodex 28d agoThe accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.