3 ms·
Hello HN — I’m the coauthor of this post. You may remember me as that guy who spent most of 2022 posting GPT-3 screenshots to Twitter, most famously prompt inje
by goodside 4y ago
Hello HN — I’m the coauthor of this post. You may remember me as that guy who spent most of 2022 posting GPT-3 screenshots to Twitter, most famously prompt injection and “You are GPT-3”. Happy to answer any questions about Claude that I can.
- detrites 4y agoThanks for being here to answer questions. One possibly difficult topic others also may be interested in, after reading Claude's responses in the article, is: what does "harmless" mean? For example, if asked to help the user understand how to do something "bad", will it give the answer if they claim they want this information in order to help them write a screenplay, versus if they seem have an intent to do it? And how is "bad" decided? We can recognise through everyday personal interactions that one persons "bad" is another persons "good", and across country-boundaries even the legality of these distinctions can be radically different. One counterargument to these constraints is that anyone can already use the internet to access all of the same information the model was trained on, unencumbered by whatever intent they may or may not have. As such, what are the rationale for making these attempts at the somewhat invasively-impossible task of determining user intent? This has never been employed with search engines before, which have lead to a rich explosion of innovation and education, so why attempt it now, in what could be argued is ultimately an iteration of search engine technology?
- goodside 4y agoThe motivation as I understand it has less to do with present-day misuse, and more to do with maintaining controllable behavior in accordance with an arbitrary, human-written “Constitution”. Anthropic is attempting to make a model that will not harm (in the unambiguous, uncontroversial sense of the word) humans even if it is superhumanly intelligent, or trusted with real-world control.
- Nevermark 4y agoYou can think adversarial models, which are often used to detect and negatively reinforce quality issues in model outputs. Claude outputs an answer. Then Claude independently rates the output for "helpfulness" as in literally "Claude, how helpful is this answer to this question". There is no collusion between the two results because they are run independently. Then Claude also rates answers for "honesty" and "harm". Then Claude's parameters are updated to increase helpfulness and honesty, and decrease harmfulness, based on back propagating those ratings to the parameters as they impacted the signals produced by the original question. Not saying that is exactly what they are doing, but that is one approach. It manages to leverage language models to train themselves on broad concepts, as apposed to brittle, more unreliable and vastly more resource intensive manual labeling. Very clever. As the models get better at languages (and other modalities), and the concepts behind them, the models also get better at schooling themselves. --- It occurs to me, that this self-oversight could be made more even more robust by training 10 Claude's, and having each Claude be rated for good behavior by the other nine, and rewarding the best Claude. Competition could make the trained-in motivations (to be the most honest, helpful and non-harmful) even more explicit, in that there would be very strong competitive motivation to continuously becoming the most virtuous and valuable, with the bar ever rising. Maybe the winning results each iteration could also be shown to the losing models, as an example of what could be done better. This really is a great direction. Kudos to Anthropic.
- zaptrem 4y agoWouldn’t all the Claudes be incentivized to simply trash each other constantly in that case?
- Nevermark 4y agoThat would certainly be something to design clear of. I don't think that is a problem. Each query runs separately so there is no "collusion", i.e. shared signals and coordination, between contrary goals (winning and virtue). Also, all the information about ratings, winning and winning examples can be used without ever giving the models explicit information about the population of models and how they are being used as a group. They don't need to know they are in a competition for competitive information to be used to update them. They just know they have ratings to improve, some indicator of how close to "the bar of currently targeted virtue" they are, and examples of how they could have improved them. Of course, I am just spitballing, and assuming the training regimen gets vetted by a lot of people (and models?!?). -- In the long run, when there are long running artificial personalities with personal memories and more direct awareness of their own motivations and options, there will certainly be the need for additional levels of moral wiring to be considered.
- detrites 4y agoGreat answer, thank you. The issue of regarding humans as AI-persuadable entities is certainly one to be carefully considered. Indeed, if it were to occur in the truest sense, we'd never know it. Another view is any AI we give birth to may only be constituted of what we are; we who ultimately, if imperfectly, demonstrate value for all life. In a sense, our constitution as "mostly harmless" may be AI's default.
- scrollaway 4y agoI’m looking for your thoughts on the following: It should be somewhat easy to teach these types of models to reach for a particular tool at times where they need it, yes? I can instruct ChatGPT for example to tell me when it should use a calculator during a session. If instead I allow it to fall back to an external calc process, then suddenly, I have a chatbot that has reasoning AND better mathematical accuracy. Also: I’ve also been entertaining the idea of having multiple layers of GPT interact with one another. So you feed back some interaction into another GPT instance without context, and ask it for example how it would verify the accuracy of certain statements (and you can ask it for machine readable code, even). Finally, I know a lot of people who start playing a lot with GPT and get disheartened because they see the quality of responses isn’t there. But the fact ChatGPT has the capacity to reason, has chain of thought, has given me a newfound appreciation for how close to AGI we might be. It has also given me an appreciation for how much simpler humans are than we like to think. I’ve introspected a lot in the past months and often ask myself: is my speech any different than “predicting the next few words”? And I feel like it’s just text prediction with some more layers on top.
- dpaleka 4y ago[I mean no bad faith in this comment, I'm a fan of yours.] Why answer questions about harmlessness/safety in such a roundabout way? Both OpenAI and Anthropic are clear about what words like "safe" are intended to mean: a stepping stone to "AI does not kill all people when given control". Avoiding to state this clearly only invites unnecessary culture war disagreements in every discussion about these models.
- goodside 4y agoMaybe you’re right. It’s partially laziness on my part — it takes a while to explain long-term issues, and those who are inclined to care about them are generally aware of who started Anthropic and why.