4 ms·
> Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done. How far do you go? You don't need to tell it explicitly that using che
by teiferer 13d ago
> Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.
How far do you go? You don't need to tell it explicitly that using chemicals A and B in ways X and Y result in a bomb that can kill lots of people. It's enough that it knows A and B and X and Y in isolation, some connections that are indirect, and it will combine those things on its own. So you can't tell it about A, B, X or Y. But those are also just results of other steps Where to stop? You won't have any chemistry in the traning data? No algorithms to prevent it from using them in an undesired way? This is just bot workable. It's akin to banning knives from stores because somebody coul figure out that one can kill people those. Until people figure out that scissors are essentially knives.
- kelseyfrog 13d agoLLMs can only repeat and interpolate data. They can't create anything new. So, you don't need to go that far.
- TheMayorOfDunce 13d agothis is immediately disprovable and embarrassingly naive in the year of AI generating cancer vaccines and solving Navier-Stokes. You can argue "the vaccine is just interpolating chemicals together" and "the solution uncovered is just interpolating mathematical operations together", but by that standard there is literally nothing new under the sun.
- kelseyfrog 13d agoThe Navier-Stokes solution was an interpolation of existing data.
- TheMayorOfDunce 13d agoadmittedly I am not a mathematician, so the Navier-Stokes solution is just an example I am using. But how is it not something "new" if it did not exist before? At what point could something ever possibly be new, if "new" means "this uses absolutely zero existing elements"? Nothing in math would ever be "new". Nothing in physics or chemistry would ever be "new", by this standard. It seems to me that the only reason to declare this solution "not new" is specifically to dismiss AI. If a human had deduced the Navier-Stokes solution, who would bother to scoff "that's not new! the numbers already existed!"?
- mjburgess 13d agoThe claim is OpenAI stole the work of mathematicians who had the same proof that they had developed using chatgpt conversations. Since by default, 'sharing' is turned on, and it often 'turns itself on' -- it is plausible gpt6 had been trained on the work of mathematicians who had effectively solved this problem in private.
- jacomoRodriguez 13d agoIf inrember correctly, this mathematicians worked on a simpler version of the problem and called transforming this in the solution to the wider problem a "remarkable" thing to do. Based in this, solving this problem is very impressive
- mjburgess 13d agoMaybe. Or maybe its evidence that the frontier of mathematics is knowledge-bound rather than understanding-bound or even just manpower-bound. In many cases where proofs have emerged, that I have read, the LLM has retrieved some antique lemma unknown to the mathematician. OpenAI spent 10-20m USD in energy costs to produce that proof with likely substantially similar prior work in the training data. What does this say? Who knows. It continues in the tradition of using measurements of intelligence in humans, applied to LLMs with the hopes the "stolen valour" transfers. Here, the NS problem was a useful framing problem for mathematics to progress because of how it interacted with the development of mathematics broadly -- ie., how it progressed techniques, ideas, understandings, etc. When we apply these issues to LLMs (whether IQ tests or mathematical proofs) we always discover something substantial lacking beneath the interesting facade of useful answers. The process isnt useful. And it is precesiely the process which these tests, in humans, are supposed to help with. The tests themselves (IQ or otherwise) arent the point. No one cares about their answers. LLMs represent an alternative understanding-free approach to solving problems, with variable success rates depending on how similar the problem is to the training data and its rewarded reasoning traces. That mathematics is making substantial progress, "10 million USD / problem" at a time, in using understanding-free methods -- says something sociologically interesting about the state of the field. Something which was already know: mathematics has long been full of a vast amount of papers, proofs, theorems and lemmas that few have ever read, or investigated. Mathematics has long been in a crisis of "overproduction of unvisited knowledge", LLMs are exploiting that otherwise unmined gold.
- BryceEller 13d agoAll mathematic breakthroughs could be characterized as an interpolation of existing data
- GPerson 13d agoNot really. There’s no honest sense in which quantum mechanics is interpolated from the text of Euclid’s elements, to demonstrate the point with a very extreme example. Many mathematical breakthroughs of the past seem to have involved the observation of semantically interesting concepts, beyond the syntax of known theories. (To onlookers this particular post makes no claim about AI’s capacity to make the same observations.)
- kelseyfrog 13d agoIt depends on the subspace in which the interpolation is framed.
- casey2 13d agoIf they can "interpolate" existing data to that level then saying "just don't hack" doesn't make any sense, hacking is derived from knowledge of software systems. It's far harder to solve navier-stokes than creating a program that replicates and abuses computer resources. You could as models improve continue to remove more and more training data, what happens when there is no more data left to remove but a running system still outperforms humans? I think you grossly overvalue data.
- teiferer 12d agoOk, then let's go with that viewpoint and apply it to the original question: How much do you need to remove from the training data such that from what remains, nothing harmful can be "merely interpolated" (which you acknowledge the LLM is in principle able to do) and anything harmful would require something "new" that is not in the training data and that the LLM is incapable of coming up with. The argument is that there won't be much left in the training data if that's your approach.
- teiferer 12d agoIt's not clear where the line between "interpolate" and "new" goes. They can infer that certain chemicals in combination create a boom and that boom is bad for health and finally those 1 and 1 make a 2 without having been told ever what a bomb is. Is that "new"? And does it matter as long as it is harmful?
- mjburgess 13d agoThe issue is not whether an ML model of any kind can generate (X_ReasoningTrace, X_Answer) distributed like P_HumanExpert(X) -- the issue is always why it would do so. By introducing modelling of "Reasoning Traces" into LLMs, and reinforcing patterns of reasoning -- this gives you a system which generates expert-like distributions of output. This lifts the "stochastic parrot" issues, or the "knowledge interpolation" problem, into different parts of the process. It isnt my view that the "ReasoningTraces" which you think are derivable from mere "basic propositions" concerning, say, hacking are actually things that LLMs can derive. Ie., I dont think LLMs have rich representational models of what they appear to understand. Instead, they are given "reasoning proxies" which allow them to reason without such understanding. This is done by providing vast specialised datasets of reasoning examples. In the case of hacking, there are large numbers of competitive datasets (forums, reports, etc.) which provide these reasoning traces. And no doubt, major vendors have paid a vast amount for special case expert-prepared datasets. So I do not believe that by witholding such reasoning exemplars, and traditional "question/answer" datasets, that LLMs can infer these things. And at least, no major vendor is doing this to my knowledge. So they are lying. They are pretending the alignment issue is "AI going rogue" when they are explicitly training the systems to "go rogue" and have done nothing at all to shape datasets to lack these capabilities. The issue here isnt alignment at all. It's training on hacking datasets. (EDIT: Philosophically, you could ask whether the reasoning-proxies LLMs are given form a kind of 'representational structure' akin to understanding, and at least, I'd concede they model understanding. But they lack important properties (eg., LLMs cannot act on them to evolve them, as with us: when I think about one of my representations to derive (eg.,) entailments of it, I thereby revise my representation. The key properties of 'evolving self-understanding' are likely to be provided by substantial (unknown) revisions to how the training/reward layer works. No doubt one of the meanings of 'recursive self-improvement' is just such a modification).
- busssard 10d agoyou are right. what they propose is called "deep ignorance" and is not a fix to alignment. the perpetrator just has to supply a volume of bio-textbooks and let the model figure out the DNA synthesis. (just) corpus-level alignment is the way to deep alignment, but not deep-ignorance.