5 ms·
The vast majority of human data can be refined into something approximating high quality textbooks, which is what happened here too.
by airgapstopgap 3y ago
The vast majority of human data can be refined into something approximating high quality textbooks, which is what happened here too.
- skepticATX 3y agoOnly 1/7 of the data they are using is synthetic though. The other 6/7 was just filtered from a larger dataset, not refined in any way.
- WJW 3y agoI wonder if that is true. Intuitively it seems to me that there are probably many areas of knowledge that cannot be reliably summarized into a textbook. Most subjective experiences, for example. Qualia have been notoriously difficult to describe too, even if the might be absolute. Even for fields that lend themselves well to being converted to textbook format there is often a tradeoff between accuracy and conciseness. The more you refine, the more nuance you throw away. It seems like a Very Hard Problem (tm) to know which data are superfluous and which are not, especially at scale.
- cjbprime 3y agoWell, used to be a Very Hard Problem. Now you can just test it: train a model on the textbook and ask it to solve many problem challenges that it hasn't seen before. After you get a score for the model trained on the whole textbook, try removing each sentence in the book in turn. If removing that particular sentence decreased the test scores then keep it in, else throw it away.
- WJW 3y agoI don't think that is sufficient: - Perhaps the model can be improved by adding sentences, but which ones? There is a potentially infinite amount of sentences to add and no good way to select them - Your proposed method handwaves the method of acquiring "many problem challenges that it hasn't seen before", which just moves the problem. Constructing a problem set containing the full range of potential problems to solve is again a Very Hard Problem. The second problem is not just theoretical either: see https://sitn.hms.harvard.edu/flash/2020/racial-discrimination-in-face-recognition-technology/ https://sitn.hms.harvard.edu/flash/2020/racial-discriminatio... for example, where a facial recognition algorithm performed much worse on non-white women because the training set didn't contain enough pictures of them. AI software is notorious for finding this kind of loophole and for overfitting itself for the training set rather than for the real world.
- mejutoco 3y agoVery interesting point. I wonder: are subjective experiences knowledge? Not scientific knowledge. One one side they are part of who we are. On the other side, same as an airplane does not copy a bird 100%, it makes sense that to make a machine "think" we would feed it rational content. That is, content that follows the scientific method.
- WJW 3y agoI suppose that depends on what type of AI you are trying to build. One that is used to help design airplanes can mostly get by with pure scientific knowledge, although it still needs to know about softer phenomena like claustrophobia and personal space to understand that packing in humans as close as physically possible is not desired. On the other end of the AI spectrum, an AI being trained to be a childrens toy, old people companion, nursing robot or even psychologist would be incredibly deficient if it didn't understand emotions. I don't think using purely content that follows the scientific method would be sufficient for such an AI.
- mejutoco 3y agoGood points. I especially like the claustrophobia one in planes. On the other side, there are textbooks about "emotions". I own a Cognitive Behaviour Therapy manual (seemed interesting) that goes step by step through how to conduct a session (it is aimed at therapists or at the reader being their own) as it progresses. As a textbook, similar professional references could be more useful than say reddit or some blogs. Those other sources on the internet will be mostly derived from the reference materials in the field, or from personal anecdotes. I can imagine textbooks exists on how to deal with people on the spectrum, or for them to recognize social cues, or how symptoms of depression look like. I am still skeptical. You did not say anything of the sort, but the scientific method is not the opposite of emotions. It shows us what we know of emotions so far. Psychology has the reproducibility scandal but it is our best try. Even in the case of planes, I would not be surprised if an engineering manual mentions a minimum space for passengers to be comfortable, as part of regulations, or any other data that might as a side-effect solve the claustrophobia issue. TL;DR. Scientific knowledge is not antithetic to "human" knowledge. Scientific knowledge is what we actually know. Otherwise we have mysticism, or faith.
- bitshiftfaced 3y agoIf there's something that cannot be reliably summarized into a textbook, then can we expect to see an agent reliably output it as text in the first place?
- WJW 3y agoDepends on the sophistication of the agent. I have agency myself and can express my current mental state quite readily, or can express personal opinions about a wide variety of subjects. I'm not sure how you would write a useful textbook about it though.
- bitshiftfaced 3y agoWe already summarize people's mental states and personal opinions in textbooks. No problem there.
- dist-epoch 3y agoIf it can't be reliably summarized, how can a human judge the output of the agent then? Compare it against what? It's a bit circular.
- bitshiftfaced 3y agoThe point is that if something is difficult to express or encode in to words for a textbook, then there's an underlying reason that would apply to not just textbooks but also to others who try to write about the same thing.
- AlotOfReading 3y agoYes. Translation of classical languages is a good example. There are lots of subtle nuances that aren't captured by any pedagogical text and we accept that humans require require apprenticeships/graduate study to become competent (self study is not enough), but the output is obviously text.
- bitshiftfaced 3y ago
- deleted 3y ago[deleted]