10 ms·
CauseNet: Towards a causality graph extracted from the web
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- pavlov 1y agoThe sample set contains: { "causal_relation": { "cause": { "concept": "boom" }, "effect": { "concept": "bust" } } } It's practically a hedge-fund-in-a-box.
- bbor 1y ago> CauseNet aims at creating a causal knowledge base that comprises all human causal knowledge and to separate it from mere causal beliefs Pretty bold to use a picture of philosophers as your splash page and then make a casual claim like this. To say the least, this is an impossible task! The tech looks cool and I'm excited to see how I might be able to work it into my stuff and/or contribute. But I'd encourage the authors to reign in the rhetoric...
- johnecheck 1y agoIndeed. I can't take an epistemology project seriously if it has no humility. Building a perfectly accurate model of the world isn't possible. We need to create tools that make it easier for regular people to build more accurate models, not delude ourselves with dreams of perfection.
- ProofHouse 1y agoWell of course because no such model of the world can or does exist
- maweki 1y agoIt's nice to see more semantic web experiments. I always wanted to do more reasoning with ontologies, etc., and it's such an amazing idea, to reference objects/persons/locations/concepts from the real world with uris and just add labeled arrows between them. This is such a cool schemaless approach and has so much potential for open data linking, classical reasoning, LLM reasoning. But open data (together with RSS) has been dead for a while as all big companies have become just data hoarders. And frankly, while the concept and the possibilities are so cool, the graph databases are just not that fast and also not fun to program.
- thedudeabides5 1y agosemantic web/OWL was always way too heavy to imagine humans using, you could imagine AI doing the heavy lifting here though..
- thicknavyrain 1y agoI know it's a reductive take to point to a single mistake and act like the whole project might be a bit futile (maybe it's a rarity) but this example in their sample is really quite awful if the idea is to give AI better epistemics: { "causal_relation": { "cause": { "concept": "vaccines" }, "effect": { "concept": "autism" } } }, ... seriously? Then again, they do say these are just "causal beliefs" expressed on the internet, but seems like some stronger filtering of which beliefs to adopt ought to be exercised for an downstream usecase.
- kolektiv 1y agoOh, ouch, yeah. We already know that misinformation tends to get amplified, the last thing we need is a starting point full of harmful misinformation. There are lots of "causal beliefs" on the internet that should have no place in any kind of general dataset.
- Amadiro 1y agoIt's even worse than that, because the way they extract the causal link is just a regex, so "vaccines > autism" because "Even though the article was fraudulent and was retracted, 1 in 4 parents still believe vaccines can cause autism." I think this could be solved much better by using even a modestly powerful LLM to do the causal extraction... The website claims "an estimated extraction precision of 83% " but I doubt this is an even remotely sensible estimate.
- kykat 1y agoIn the precision dataset, there are the sentences that led to this, some are: >> "Even though the article was fraudulent and was retracted, 1 in 4 parents still believe vaccines can cause autism." >> On 28 February 1998 Horton published a controversial paper by Dr. Andrew Wakefield and 12 co-authors with the title "Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children" suggesting that vaccines could cause autism. >> He was opposed by vaccine critics, many of whom believe vaccines cause autism, a belief that has been rejected by major medical journals and professional societies. All that I've seen don't actually say that vaccines cause autism
- refactor_master 1y agoMight as well go ahead and add https://tylervigen.com/spurious-correlations?page=135 https://tylervigen.com/spurious-correlations?page=135 from the looks of it.
- jack_riminton 1y agoReminds me of the early attempts at hand categorising knowledge for AI
- rhizome 1y ago"The map is not the territory" ensures that bias and mistakes are inextricable from the entire AI project. I don't want to get all Jaron Lanier about it, but they're fundamental terms in the vocabulary of simulated intelligence.
- tgv 1y agoThis makes little sense to me. Ontologies and all that have been tried and have always been found to be too brittle. Take the examples from the front page (which I expect to be among the best in their set): human_activity => climate_change. Those are such a broad concepts that it's practically useless. Or disease => death. There's no nuance at all. There isn't even a definition of what "disease" is, let alone a way to express that myxomatosis is lethal for only European rabbits, not humans, nor gold fish.
- koliber 1y agoExactly. In some cases disease causes death. In others it causes immunity which in turn causes “good health” and postpones death.
- Nevermark 1y agoContradictory cause-effect examples, each backed up with data, are a reliable indicator of a class of situations that need a higher chain-effect resolution. Which is directly usable knowledge if you are building out a causal graph. In the meantime, a cause and effect representation isn't limited to only listing one possible effect. A list of alternate disjoint effects, linked to a cause, is also directly usable. Just as an effect may be linked to different causes. Which if you only know the effect, in a given situation, and are trying to identify cause, is the same problem in reverse time.
- koliber 1y agoIt is my opinion that if we examine any factor closely, it will have multiple disjoint effects. As in nothing is absolutely unilateral in its effects. Some of those effects will depend on certain conditions. If it is possible to specify condition, annotations, and other nuances such as levels of confidence or source of the opinion, such a database might be pretty useful.
- deleted 1y ago[deleted]
- 1y ago
- TofuLover 1y agoThis reminds me of an article I read that was posted on HN only a few days ago: Uncertain<T>[1]. I think that a causality graph like this necessarily needs a concept of uncertainty to preserve nuance. I don't know whether this would be practical in terms of compute, but I'd think combining traditional NLP techniques with LLM analysis may make it so? [1] https://github.com/mattt/Uncertain https://github.com/mattt/Uncertain
- 9dev 1y agoRight. The first example on the site shows disease as a cause, and death as an effect. This is wrong on several levels: There is no such thing as healthy or sick. You’re always fighting off something, it just becomes obvious sometimes. Also, a disease doesn’t necessarily lead to death, obviously.
- notrealyme123 1y agoI get some vibes of fuzzy logic from this project. Currently a lot of people research goes in the direction that there is "data uncertainty" and "measurement uncertainty", or "aleatoric/epistemic" uncertainty. I foumd this tutorial (but for computer vision ) to be very intuitive and gives a good understanding how to use those concepts in other fields: https://arxiv.org/abs/1703.04977 https://arxiv.org/abs/1703.04977
- koliber 1y agoI wonder how they will quantize causality. Sometimes a particular cause has different, and even opposite, effects. Alcohol causes anxiety. At the same time it causes relaxation. These effects depend on time frame, and many individual circumstances. This is a single example but the world is full of them. Codifying causality will involve a certain amount of bias and belief. That does not lead to a better world.
- lwansbrough 1y agoI was hoping this would be actual normalized time series data and correlation ratios. Such a dataset would be interesting for forecasting.
- ivape 1y agoI don’t know if it’s inadvertent, but it’s headed toward just becoming an engine for over fitted generalizations. Each casual pair will just emerge based on frequency, which will reinforce itself in preemptively and prematurely classifying all future information. Unfortunately, frequency is the primary way AI works, but it will never be accurate for causality because causality always has the dynamic that things can happen just “because”. It’s hacked into LLMs via deliberate randomness in next-token prediction.
- huragok 1y agothe cyc of this current ai winter
- daloodewi 1y agothis will be super cool if it can be done!
- ProofHouse 1y agoI think this is many years old
- deleted 1y ago[deleted]
- rwmj 1y agoIsn't this like Cyc? There have been a couple of interesting articles about that on HN: https://news.ycombinator.com/item?id=43625474 https://news.ycombinator.com/item?id=43625474 "Obituary for Cyc" https://news.ycombinator.com/item?id=40069298 https://news.ycombinator.com/item?id=40069298 "Cyc: History's Forgotten AI Project"
- HarHarVeryFunny 1y agoSeems like a subset of CYC - attempting to gather causal data rather than declarative data in general. It's a bit odd that their paper doesn't even mention CYC once.
- 2OEH8eoCRo0 1y agoEverything old is new again
- TomasBM 1y agoCyc is hardly [1] mentioned in modern work under the knowledge representation and reasoning umbrella, because most [2] of it was/is unavailable or unknown to most researchers. It's hard to build on something that's primarily marketing material. [1] I could be wrong, but even those that mention Cyc use it only as a historical example of early work in KRR / symbolic AI. [2] OpenCyc being the small subset which is available, tho I haven't met anyone who worked with it.
- athrowaway3z 1y agoA cool idea, in desperate need of an example use case.
- AlienRobot 1y agoI wonder what is this for.
- larodi 1y agoWhy not use PROLOG then, is the essence of cause and effect in programming. And also can expound syllogisms.
- orobus 1y agoThe conditional relation represented in prolog, and in any deductive system, is material implication (~PvQ), not causation. You can encode causal relationships with material implication but you’re still going to need to discover those causal relationships in the world somehow.
- cubefox 1y agoConditional statements don't really work because "if A, then B" means that A is sufficient for B, but "A causes B" doesn't imply that A is sufficient for B. E.g. in "Smoking causes cancer", where smoking is a partial cause for cancer, or cancer partially an effect of smoking. "A causes B" usually implies that A and B are positively correlated, i.e. P(A and B) > P(A)×P(B), but even that isn't always the case, namely when there is some common cause which counteracts this correlation. Thinking about this, it seems that if A causes B, the correlation between A and B is at least stronger than it would have been otherwise. This counterfactual difference in correlation strength is plausibly the "causal strength" between A and B. Though it doesn't indicate the causal direction, as correlation is symmetric.
- larodi 1y agoI didn't say one does not to discover the causal relationships, but once discovered, such relationships can be explored and followed and _inferred_ on in a very syllogistic manner. My comment was really about the proposal in the article. On the other hand, what we seem to have with LLM models, and the transformer approach in particular, is a sort of probable statistical correlation, calculated by brute-forcing and approximation (the gradient descent). So this is not true causation also, it becomes one only after a human observes it and agrees it follows certain causality. /Not sure whether I can state that it is also material but in another non-logical sense, perhaps would sound nonsensical, but the apparent logical structure in the LLM production rather emerges from training patterns, not from explicit logical operations./ There's nothing wrong having a graphical structure which models causality, and of course - this needs to be discovered first. But then we have LZW/Sequitur using very brute-force way in order to find the minimal grammar for compressing certain data lossless-ly, thus discovering some logical structure (and correlation), but this is not yet causation. Indeed finding patterns != finding causal relationships. My gut feeling is we want something that would result in a correct PROLOG-like set of inference rules, but based on actual causality, not conflating correlation. And then this - for a larger corpus - world's knowledge, but we don't have the means (yet) to figure out the correlation, even though approaches exist for smaller corpus. It is perhaps the gradient descent and the fact that this composition of tensor algebra is differentiable that is the ingenious thing about the ML we deal with now, but everyone is dreaming of some magic algo which would allow finding the causation so that it results in non-probabilistic graphical model, or at least a model that we can follow the stochastic branching on in a observable manner. It is indeed ingenious to fold multi-dimensional spaces, multiple times, in order to disambiguate the curvature of bunny's ear from the one of bear's ear. But it just does not feel right to do logic and causation by means of differential calculus and stochastic structures.
- Unirely01 1y ago[dead]
- bbstats 1y agoCausality is literally impossible to deduce...
- amelius 1y agoCan't an LLM extract this type of information with reasonably high accuracy?
- kruffalon 1y agoI read it as "casual" rather than "causal", got very dissapointed while reading the article! An inventory of casual knowledge would be really fun, although it's hard to think what it would consist of now that I think about it... There is this concept of "hidden knowledge" about all the things you know at work that no one really thinks about is knowledge so it's hard to let newcomers know about it. But that does sound different than "casual knowledge", and so does "trivia". Oh well!
- aleatorianator 1y agoit's as simple as precisely describing "common sense"
- kruffalon 1y agoYes, that very simple task :D But I also wonder if that is actually an equivalent. I think of "common sense" as think you should know to function well in your current society. I really don't know what "casual knowledge" would mean. In my head it's some kind of low stakes knowledge for everyday life (but more 'useful' than trivia). Maybe the order is "common sense", "casual knowledge" and trivia?
- eloeffler 1y agoI did, too! And it reminded me of a project idea I had a while ago: A time traveller's wiki that collects casual knowledge for different times (and different places). Such as: "Buying a train ticket in Paris in 1972". But it was a shower thought and it's pretty hard to imagine how this knowledge should be collected and especially presented. In a way, wikipedia is already doing this by keeping records of articles as they change over the years :) The article about train tickets wasn't so good as an example but "computer monitor" from 2004 is kind of fun to read :) Unfortunately, "casual knowledge" is often omitted when writing informative articles. In this example, there is no mention that power buttons are often located somewhere in the back of the monitor, which was good to know in 2004. Also, some monitors are drawing power from the computer, thus they won't power up before the computer will. And speaking of that: You may want to turn of your computer after shutdown! Edit: This would probably be useful for novelists and filmmakers (in addition to the casual time traveller)
- RianAtheer 1y ago[dead]
- mark_l_watson 1y agoThis might be of at least some value to augment training LLMs? I spent a lot of time in the 1980s and early 1990s using symbolic AI techniques: conceptual dependency, NLP, expert systems, etc. While two large and well funded expert system projects I worked on (paid for by DARPA and PacBell) worked well, mostly symbolic AI was brittle and required what seemed like an i finite amount of human labor. LLMs are such a huge improvement that the only real use I see in projects like Cause et, the defunct OpenCyc project, etc. the only possible practical use might be as a little extra training data.
- fohara 1y agoThe associated paper references Judea Pearl's theories on causality, but curiously doesn't mention the DoWhy implementation [0], which seems to have some recognition in the causal inference space. [0] https://github.com/py-why/dowhy https://github.com/py-why/dowhy
- growingkittens 1y agoOrganizing all knowledge requires a flexible system of organization (starting with how the categories are organized and accessed, not the data). Random thoughts about organizing knowledge: - Categories need fractal structures. - Categories need to be available as subcategories to other categories as a pattern. - Words need to be broken down into base concepts and used as patterns. - Social information and context alter the meaning of words in many cases, so any semantic web without a control system has limited use as an organization tool.
- MangoToupe 1y agoWittgenstein is calling
- pfdietz 1y agoThis is difficult, but then I just had someone earnestly inform me that the covid virus doesn't cause covid, so I think there's a need here, if only to have an automated way of identifying idiots.
- kgrizzle 1y agoReminds me of the cyc project. https://en.wikipedia.org/wiki/Cyc https://en.wikipedia.org/wiki/Cyc
- sinuhe69 1y agoI find the simple expression of a causes b as in this database without qualification not very helpful. At least, we need causal graphs/causal digram loops to describe these causal relationships better. [0] https://en.wikipedia.org/wiki/Causal_graph https://en.wikipedia.org/wiki/Causal_graph Harvard has a free course about it: https://www.edx.org/learn/data-analysis/harvard-university-causal-diagrams-draw-your-assumptions-before-your-conclusions https://www.edx.org/learn/data-analysis/harvard-university-c...
- quirk 1y agoThe fact that they are using Wikipedia for a primary data source exempts them from any further serious consideration.
- circlemaker 1y agoThis made me think of a much more interesting project. A compendium of information automatically extracted from research articles. Essentially one totalizing meta analysis. E.g. If it reads an article about the relationship between height and various life outcomes in Indonesian men, then first, it would store the average height of Indonesian men, the relationship between the average height of Indonesian men and each life outcome in Indonesian men, the type of relationship (e.g. Pearson's correlation), the relationship values (r value), etc. It would store the entity, the relationship, the relationship values, and the doi source. Something like a quantitative Wikipedia.