9 ms·
Is anyone playing with the combination of generative AI and OpenCyc?
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- cjbprime 2y agoHaven't LLMs simply obsoleted OpenCyc? What could introducing OpenCyc add to LLMs, and why wouldn't allowing the LLM to look up Wikipedia articles accomplish the same thing?
- drdaeman 2y agoI’m not familiar with Cyc/OpenCyc, but it seems that it’s not just a knowledge base, but also does inference and reasoning - while LLMs don’t reason and will happily produce completely illogical statements.
- p1esk 2y agoCan you please give an example of a “completely illogical statement” produced by o1 model? I suspect it would be easier to get an average human to produce an illogical statement.
- n2d4 2y agoGive it anything that sounds like a riddle, but isn't. Just one example: > H: The surgeon, who is the boy's father, says "I can't operate on this boy, he's my son!" Who is the surgeon of the boy? > O1: The surgeon is the boy’s mother. Also, just because humans don't always think rationally doesn't mean ChatGPT does.
- pezezin 2y agoHaha, you are right, I just asked Copilot, and it replied this: > This is a classic riddle! The surgeon is actually the boy's mother. The riddle plays on the assumption that a surgeon is typically male, but in this case, the surgeon is the boy's mother. > Did you enjoy this riddle? Do you have any more you'd like to share or solve?
- throw310822 2y agoHa, good one! Claude gets it wrong too, except for apologizing and correcting itself when questioned: "I was trying to find a clever twist that isn't actually there. The riddle appears to just be a straightforward statement - a father who is a surgeon saying he can't operate on his son" More than being illogical, it seems that LLMs can be too hasty and too easily attracted by known patterns. People do the same.
- varjag 2y agoIt's amazing how great these canned apologies work at anthropomorphising LLMs. It wasn't really in haste, it simply failed because the nuance fell below noise in its training set data but you rectified it with your follow-up correction.
- throw310822 2y agoWell, first of all it failed twice: first it spat out the canned riddle answer, then once I asked it to "double check" it said "sorry, I was wrong: the surgeon IS the boy's father, so there must be a second surgeon..." Then the follow up correction did have the effect of making it look harder at the question. It actually wrote: "Let me look at EXACTLY what's given" (with the all caps). It's not very different from a person that decides to focus harder on a problem once it was fooled by it a couple of times already because it is trickier than it seems. So yes, surprisingly human, with all its flaws.
- varjag 2y agoBut thing is it wasn't trickier than it seemed. It was simply an outlier entry, like the flipped tortoise question that tripped the android in the Bladerunner interrogation scene. It was not able to think harder without your input.
- wanderer2323 2y agoEasy, from my recent chat with o1: (Asked about left null space) ‘’’ these are the vectors that when viewed as linear functionals, annihilate every column of A . <…> Another way to view it: these are the vectors orthogonal to the row space. ‘’’ It’s quite obvious that vectors that “annihilate the columns” would be orthogonal to the column space not the row space. I don’t know if you think o1 is magic. It still hallucinates, just less often and less obvious.
- brokensegue 2y agoaverage humans don't know what "column spaces" are or what "orthogonal" means
- Sabinus 2y agoAverage humans don't (usually) confidently give you answers to questions they do now know the meaning of. Nor would you ask them.
- throw310822 2y agoAh hum. The discriminant is whether they know that they don't know. If they don't, they will happily spit out whatever comes to their mind.
- leklund 2y agoSure average humans don’t do that, but this is hackernews where it’s completely normal for commenters to confidently answer questions and opine on topics they know absolutely nothing about.
- mdp2021 2y agoAnd why would the "average human" count?! "Support, the calculator gave a bad result for 345987*14569" // "Yes, well, also your average human would" ...That why we do not ask "average humans"!
- 2y ago
- rcxdude 2y agoSuch systems tend to be equally good at producing nonsense: mainly because it's really hard to make a consistent set of 'facts', and once you have inconsistencies, creative enough logic can produce nonsense in any part of the system.
- 2ro 2y agoCyc apparently addresses this issue with what are termed "microtheories" - in one theory something can be so, and in a different theory it can be not so: https://cyc.com/archives/glossary/microtheory/ https://cyc.com/archives/glossary/microtheory/
- cookiengineer 2y agoLLMs have just ignored the fundamental problem of reasoning: symbolic inference. They haven't "solved" it, they just don't give a damn about logical correctness.
- wordpad25 2y agological correctness as in formal logic is a huge step down LLMs understand context and meaning and genuine intention
- 2ro 2y agoCyc apparently addresses this issue with what are termed "microtheories" - in one theory something can be so, and in a different theory it can be not so: https://cyc.com/archives/glossary/microtheory/ https://cyc.com/archives/glossary/microtheory/
- optimalsolver 2y ago>A microtheory (Mt), also referred to as a context, is a Cyc constant denoting assertions which are grouped together because they share a set of assumptions This sounds like a world model with extra-steps, and a rather brittle one at that. How do you choose between two conflicting "microtheories"?
- rurban 2y agoNo, they are completely orthogonal. LLM are likelyhood completers, and classifiers. OpenCYC brings some logic and rationale into the classifiers. Without rationale LLM will continue hallucinating, spitting out nonsense.
- tpm 2y agoLLMs don't know what is true (they have no way of knowing that), but they can babble about any topic. OpenCyc contains 'truth'. If they can be meaningfully combined, it could be good. It's the same as using LLM for programming, when you have a way to evaluate the output, then it's fine, if not, you can't trust the output as it could be completely hallucinated.
- brokensegue 2y agosometimes i think projects like cyc are like 3n+1 problem for AI. it's so alluring.
- K0balt 2y agoHmm… maybe we could train /tune a model on symbolic logic similar to or even using CycL instead of python, and then when we have it “write code” it would be solving the problem we want it to think about, using symbolic logic? You might be on to something here. The problem being there isn’t billions of tokens worth of CycL out there to train on, or is there?
- JimDabell 2y agoI think it would be interesting to use an LLM to distill Wikipedia into a set of assertions, then iterate through combinations of those assertions using OpenCyc. You could look for contradictions between pages on the same subject in different languages, or different pages on related subjects. You could synthesise new assertions based on what the current assertions imply, then render it to a sentence and fact-check it. You could use verified assertions to benchmark language parsing and comprehension for new models. Basically unit test NLP. You could produce a list of new assertions and implications introduced when a new edit to a page is made.
- mindcrime 2y agoAlong with that, a portion of the content of Wikipedia is already available in structured assertion form, thanks to DBPedia[1] and Wikidata[2]. I don't know the exact percentage, but it's a starting point at the very least. [1]: https://www.dbpedia.org/ https://www.dbpedia.org/ [2]: https://www.wikidata.org/wiki/Wikidata:Main_Page https://www.wikidata.org/wiki/Wikidata:Main_Page
- throw310822 2y agoI've tried to use ChatGPT to produce wikidata queries- it sounds like a great combination. Unfortunately it's pretty hard to make it produce valid queries, and to find the wikidata documentation to teach it.
- mindcrime 2y agoMy interest is a little more like flipping that around the other way. Take output from the LLM, get it into structured assertion form, and then use those assertions as part of a query / inference process which pulls in other "known good" (or "ground truth" if you will) assertions from sources like DBPedia or Wikidata. The idea being to either verify the LLM output, or possibly to extend it using inferred conclusions. I think the way I think about it is somewhat akin to what AWS are doing here, where they talk about using automated reasoning to reduce hallucinations from LLM's: https://aws.amazon.com/blogs/machine-learning/reducing-hallucinations-in-large-language-models-with-custom-intervention-using-amazon-bedrock-agents/ https://aws.amazon.com/blogs/machine-learning/reducing-hallu...
- mindcrime 2y ago> Is anyone playing with the combination of generative AI and OpenCyc? OpenCyc specifically? No. But related semantic KB's using the SemanticWeb / RDF stack? Yes. That's something I'm spending a fair bit of time on right now. And given that there is an RDF encoding of a portion fof the OpenCyc data out there[1], and that may make it into some of my experiments eventually, I guess the answer is more like "Not exactly, but sort of, maybe'ish" or something. [1]: https://sourceforge.net/projects/texai/files/open-cyc-rdf/ https://sourceforge.net/projects/texai/files/open-cyc-rdf/
- jillesvangurp 2y agoMicrosoft published quite a bit on the notion of graph RAG which is about enhancing regular RAG (retrieval augmented search) with graph databases. The idea is that instead of just pulling semantically related information to the query, it also pulls in information about things connected to those things. This gives the LLM more contextual information to work with. Sounds like it would probably have. If you combine that with entity recognition and automated graph construction from unstructured data, you might get something that is vaguely useful.
- 2ro 2y agoif anyone is interested: https://2ro.co/post/768337188815536128 https://2ro.co/post/768337188815536128 (EZ - a language for constraint logic programming)
- pcblues 2y agoI'll be the antiquated person here. Writing/speaking well has brought people to knowledge because the authors are uniquely positioned to use genius, humour, gentleness and generosity to bring an inquisitive but ignorant person into a new area of knowledge. When the value of that can be quantified, then we can compare AI "generation" of "efficiently written, useful knowledge" with what we had/have. Same for fiction, visual art, raising children, caring for old people, and so on, and so on.
- willvarfar 2y agoWith big enough cohorts we can AB test by quantifying outcomes? Treatment group A has AI-generated texts and Treatment group B has the originals. We can have a questionnaire for how the group felt about the material etc, but we can also perhaps measure life outcomes over a bigger period e.g. performance at school or in the market place? As I write this it feels naive and reminds me of a thousand ill-thought-out AB website tests etc. But still :D
- bryanrasmussen 2y agoexactly how long are you planning on running these tests for - sounds like minimum 20 years?
- willvarfar 2y agoYeap :) There are plenty of very long-term studies going on where groups are identified and then followed over years or decades. Another approach is today you might go looking for people with different levels of airborne lead exposure as a child, and then compare the average income today to see if their is correlation etc. In these kinds of studies the treatment groups weren't wittingly part of the experiment, but you can look backwards. Either way, in 20 years we'll probably be able to identify schools where AI texts predominated earlier than other schools and, adjusting for other factors, try and tease out the actual impact of that choice.
- 2y ago
- ornornor 2y agoAt the risk of sounding dumb: I don’t understand what are the practical applications of OpenCyc. I get LLMs, you can ask questions and they’ll answer, they can write an article, they can summarize documents… what are the practical applications of OpenCyc?
- PeterStuer 2y agoOpenCyc is the remaining evolution of Cyc, which was based on the idea that symbolic AI (knowledge graphs/Semanic Networks) would lead to AGI through scaling the knowledge base. You formalize the world so you can use logic to reason about it. Later approaches of the same idea coning out of the academic Databases community reinvented this particular wheel with far better PR and branded it the Semantic Web with 'Ontologies' and RDF, OWL and its ilk. While reasonable (pun intended) for small vertical domains, the approach has never made inroads in more broad general intelligence as IMHO it does not deal well with ambiguities, contradictions, multi level or perspective modeling and circular referential meaning in 'real world' reasoning and also tends to ignore agentive, transformative and temporal situatedness. Its ideal seems to be a single thruth never changing model of the universe that is simply accepted by all.
- 2ro 2y agoCyc apparently addresses this issue with what are termed "microtheories" - in one theory something can be so, and in a different theory it can be not so: https://cyc.com/archives/glossary/microtheory/ https://cyc.com/archives/glossary/microtheory/
- mindcrime 2y agoAll true in general, but allow me to add that there has been work on incorporating probabilistic / "soft" reasoning into the Semantic Web / RDF / OWL world. For some time now there has been PR-OWL (Probabilistic OWL)[1], and the recent work on RDF* (RDF-STAR)[2][3] emphasizes its application in terms of being able to - among other things - do stuff like adding weights (confidence scores, fuzzy probabilities, what-have-you) to RDF assertions. So the need to pursue these paths is understood, although I suppose one can argue that progress has been slow and painstaking. [1]: https://www.pr-owl.org/ https://www.pr-owl.org/ [2]: https://www.w3.org/2021/12/rdf-star.html https://www.w3.org/2021/12/rdf-star.html [3]: https://w3c.github.io/rdf-star/UCR/rdf-star-ucr.html https://w3c.github.io/rdf-star/UCR/rdf-star-ucr.html
- thom 2y agoProbably, but the bitter lesson still applies.
- dannyobrien 2y agoUnsure specifically, but there's a long-standing movement to combine GOFAI symbolic approaches, and modern neurally-influence systems. https://en.wikipedia.org/wiki/Neuro-symbolic_AI https://en.wikipedia.org/wiki/Neuro-symbolic_AI
- amelius 2y agoHow would you phrase that question in opencyc?
- transfire 2y agoIsn’t Cycorp?
- deleted 2y ago[deleted]