9 ms·
I don’t use Semantic Web technologies anymore, though they still influence me
- tylerjwilk00 7y agoOr responsive web design techniques apparently.
- StuffedParrot 7y agoIt reads fine for me on mobile.
- AndrewStephens 7y agoI have some low-level hate for the Semantic Web. I run a small personal blog that I maintain using a relatively simple static site generator that I created that turns markdown files into clean(ish) html. A couple of months ago I got interested in adding semantic information to my posts so I modified the generator to add some of the common semantic tags. It was an annoying job, since the semantic information pollutes the structure of the html. Can anyone tell me what the semantic web does for me as a small-time publisher? Is it for search engines? Does it really matter that a book review (for instance, I have a few) is tagged properly?
- sjg007 7y agoEmbedding semantic information would allow Google to further refine search traffic to your web page. I assume it may also make you more authoritative wrt to the content you publish.
- lazyjones 7y ago> Can anyone tell me what the semantic web does for me as a small-time publisher? Is it for search engines? Yes, in practice it is mostly for bigger fish in the pond to easily identify and steal your content as needed. For example, Google was using reviews from small competitors' sites in Google Shopping.
- abathur 7y agoI think this is one of the big issues. The semantic information does make it easier for end users to find what they're looking for, but it also made denial of traffic possible. In a lot of cases, the information was there to get eyeballs--so this is undesirable. I guess if you don't really care about the eyeballs it can be "useful" for the big fish to pay most of the cost of serving the fraction of your server response that the end user was looking for...
- TeMPOraL 7y agoSo the root problem is actually that people care about the eyeballs. Nothing good comes from such incentive.
- abathur 7y agoMaybe. Not sure what I think about that framing. FWIW, I was picking "eyeballs" as something wider than just ad revenue. I think ads are the big share, but I'm sure there are people/orgs who want eyeballs for other reasons like ego/status, promote their company/brand/service/products, etc. In some sense I think your framing is accurate, but I don't know about whether we'd be better off (have an informational ecosystem that is more net-positive?) without status chasers. Some share of them are inevitably gaming the system and diluting the ecosystem; others probably add net value in pursuit of eyeballs?
- TeMPOraL 7y agoIn context of semantic web, pursuit of the eyeballs is a problem because it makes the people owning/creating the data also want to be delivering that data directly to the users, and be the only ones allowing to do so. Semantic web works for the opposite goal - to allow the data to be automatically transmitted, processed and understood by software, and only perhaps eventually delivered in some form to end user. As for building more net-positive information ecosystems, going for the eyeballs instead of actually caring to deliver good information isn't necessarily bad per se, just suboptimal. It's better for an eyeball-chasing site to publish some information, if otherwise that information wouldn't be published at all. But it's the eyeballs being your primary revenue source that will make you work hard to make the data as useless as possible outside your own publication - which leads to a very unhealthy information ecosystem.
- decebalus1 7y ago> It was an annoying job, since the semantic information pollutes the structure of the html. In what way? Both the html and the metadata is intended to make your website machine-friendly. You may find the html structure polluted, but crawlers would find it more informative.
- onli 7y agoMore a side note, but if you run a blog you might know that the trackback url can be specified via a RDF tag. That's a kind of semantic information, one example for one type of usage: Given other clients (here: other blogs) additional information (here: where to send the Trackback POST). The markup you added - it depends on what exactly you did. Did you add the markup for schema.org? That's in practice solely for Google. The SEO promise there is that Google will make use of the information provided and format some information nicely, which can lead to more clicks. https://moz.com/learn/seo/serp-features https://moz.com/learn/seo/serp-features explains that not badly. For things like reviews I can imagine it to be quite useful.
- zozbot234 7y ago> Does it really matter that a book review (for instance, I have a few) is tagged properly? If the semantic web was better supported, you could have a semantic annotation precisely identifying the books you are reviewing (whether by ISBN edition or otherwise), and reusers of your content (users, search engines or others) would be able to programmatically associate your review with similar content.
- coddle-hark 7y agoThat seems like it would be abused to the point of the semantic information being completely useless.
- abathur 7y agoI guess by abuse you mean ~black-hat SEO? It seems likely (and perhaps obvious) that: - people will try to abuse it - abuse will keep it from supporting naive trust of semantic information published by untrusted third parties But we're also already roughly in this scenario, and it seems like it might be easier to model and spot/discard abuse of semantic information.
- have_faith 7y agoI can't imagine what semantic tags would pollute a blog's markup as most of the semantic tags were designed to structure simple text content like a blog post. Do you have any examples? > Is it for search engines? Yes. And Accessibility.
- Vinnl 7y agoI think you might be confusing semantic HTML with the semantic web. (Which is understandable given the mention of semantic tags.) Using semantic HTML means using <article> rather than yet another <div>. What GP is referring to, however, is adding extra information to your HTML detailing what kind of data is in your tags, e.g.: <p vocab="http://schema.org/" typeof="Person"> <span property="name">Christopher Froome</span> was sponsored by <span property="sponsor" typeof="http://schema.org/Organization"> <a property="url" href="http://www.skysports.com/">Sky</a></span> in the Tour de France. </p> Here, the vocab, typeof and property attributes are used to add semantic information to the HTML. It might also give you an idea of why one might consider that a chore, especially if it doesn't appear to provide any benefit, like making your site accessible to users of screen readers.
- have_faith 7y agoYou're right, I was conflating the two overlapping concepts.
- zmix 7y ago"Semantic Web" is a wide area. What technologies did you use? Care to post a little example, as to what and how it pollutes the HTML structure?
- sawaruna 7y agoShoutouts to the 11 other people on HN still working with rdf and similar in 2020.
- zcw100 7y agoI could write a book on what's wrong with the semantic web. One of the worst isn't even technical, it's the community. There are some great people in the community but there are also a large number of extremely toxic people that drive people away. If the technology ever takes off it's going to be because some outside community cherry-picks the good parts and tells those people to f-off. That's already starting to happen and you'll hear no end of bitching from people in the semantic web community about how they're reinventing what they've already done years ago. Guess what? You're right. You're so toxic that it's worth redoing everything if it means they don't have to deal with the toxic attitudes.
- markhollis 7y agoI'm curious what those toxic attitudes are. Surely the "we already invented it and you're reinventing it" can't be the only case. I'm also curious if it's in an academia or in industry.
- wrnr 7y agoMy take on the attitude in academia: Here we describe a set of algorithms that can solve a class of problems that previous algorithms can't. In the 60' someone published a solution to a problem we have improved upon with the novel innovation of called "hyperlinks". The technical, social and economical shortcomings of our solution are invalid because it is decentralised and therefor morally superior to the current offerings, used the world over, of industry practitioners who are only doing it for the money. More funding is needed for further research.
- nl 7y agoIn general the decentralised fetishism isn't something that is big in academia (as in the academia that publishes paper). There's lots of issues in academia and even more with the semantic web, but fetishism of decentralisation isn't it.
- tasogare 7y agoI shared the experience described by the grand parent. In particular I remember I had some argument on HN with a few people and the sheer amount of bad faith and technical inaccuracy thrown at me was jaw dropping. At this point I consider SW more a cult than a technology. On the research side there are two kinds of research papers: the one that proposes an ontology for a domain, and the one that describes the conversion of an existing resource to RDF. I've never seen a paper where SW was used for something new and interesting and that would have been impossible without SW. That being said, they are also both technical and conceptual pain points that are plaguing RDF. Basically the tech is trying to address too many things: both metadata and data, and every kind of data. "IRIs that can be URLs than can be sometimes dereferenced and sometimes not, but it's better if they are and then it's Linked Data" kind of thing makes it hard to assume (and thus build) anything. So, RDF have been success in a few domains (biology) but in most case it doesn't offer a real competitive advantage over simpler and more expressive technologies such as graph databases. PS: @zcw100 if you where to really write a book about semantic web, drop me a line please.
- liminal 7y agoI really want to like semantic web technologies, but every time I try to get into them I'm stymied: * A zillion standards that all reference each other * Two zillion incomplete and incompatible implementations of those specifications * No sense of direction within it all (what's the easy path?) * Multiple rebrandings of the same ideas (Semantic Web, Linked Data, Solid...)
- hos234 7y agoI am still a fan of Googles OpenRefine tool. It's reconciliation feature that helps disambiguate Named Entities etc based on wikidata is really powerful - https://github.com/OpenRefine/OpenRefine/wiki/Reconciliation https://github.com/OpenRefine/OpenRefine/wiki/Reconciliation You can hook in your own reconciliation end point which we do at work to expand internal knowledge graphs.
- jyrkesh 7y agoThis is awesome, thanks so much for sharing. I'm really surprised I've never come across it because I've thought of building something like this before. I really want to look into how this could ingest my own post-GDPR data exports, as well as data sanitization for ML projects.
- nl 7y agoNote that OpenRefine isn't really kept up to date. The basic capabilities work ok, but lots of the additional capabilities have atrophied away.
- contravariant 7y agoWith the recent widespread interest in Category theory I still think it's a damn shame that RDF wasn't designed to treat relationships as stand-alone entities. Perhaps property graphs work better in that regard, although it's a bit weird how properties aren't themselves relationships, but perhaps that's a necessary concession to keep things efficient.
- at_a_remove 7y agoAt an old job, I knew some very idealistic folks who kept pushing semantic web business. "Let's do that everywhere!" As an exercise, I would have them open a browser, visit various sites, and then look at the source. "Go on, check to see if it validates," I would say with an anticipatory grin. Whether hand-crafted HTML or generated by any number of frameworks, many sites can barely manage to close their tags, asking for semantic references is a "just won't happen in practice" thing. I have also seen a great deal of consultant money, programmer time, sys-admin sweat, and the like focused on these toweringly-designed, completely-unused triple stores, layer upon layer of hot technologies (ever-moving, construction on the tower never ceased) fused together to create a resource-intense monstrosity that, at the end of the day, barely got used. But hey, let's look at that jazz semantic web example one more time. The most painful part is that I understand the urge to build a gleaming repository for information, where the cool URIs never change; SPARQLing pinnacles, ready to broadcast the Library of Alexandria, glimmer; and the serene manifold of abstract information lies RESTful ... but I have come to understand that the web of today is an endlessly bulldozed mudscape where Someone Very Important has to have that URL top-level yesterday (never mind that they will forget about it tomorrow), of shoddy materials and wildly varying workmanship, and where nobody is listening to your eager endpoints because the commercials are just too loud. I too once labored for information architecture, to have the correct thing in the obvious place, with accurate links and current knowledge, to provide visitors with the knowledge they desired ... but PR preempted all of it to push yet more nice photographs in yet another place: the Web as a technology for distributing images that would once live on glossy pamphlets. The vision is lovely, but we who have always lived in the castle have walked alone.
- riffraff 7y agoI would argue the problem is not the broken tags, but the business disadvantage to exposing semantic data. Remember when microformats were all the rage, and you could get hReview or hRecipe or XFN data everywhere? Then every host in turn realized that actually, it's _better_ if people can't scrape your site, and it's even better if they can't even see it and it's behind a login wall.
- 7y ago
- mark_l_watson 7y agoThese a fair criticisms of the semantic web. One thing the author misses (does not touch on at all) is domain specific RDF resources for biology, medicine, etc. schema.org and WikiData are great resources and for large companies, using these as a foundation for their own internal Knowledge Graphs can make sense. This expense is (maybe?) too large for small and medium size companies, they would not get enough benefit for the cost. I worked with Google’s Knowledge Graph as a contractor, and I am still a believer in the technology but I also respect other people’s well founded scepticism.
- ansible 7y agoI had a lot of interest in the semantic web when I first started learning about it. However, the efforts I've seen seem to be missing some critical factors for longer-term success. I think we've got a lot of work to do with regards to knowledge representation in general. One of the big things for me is that the context for any fact is critical for it to be true or not. You can have a fact like "Tim Cook is the CEO of Apple", represented in a graph like you would expect. However, that is only true today. Ten years ago it was Steve Jobs. Without explicit context encoded in the information graph, this web of data isn't as useful as it could be. Context is important for reasoning in all kinds of situations. "What if Steve Ballmer was CEO of Apple?", is a hypothetical context, where it may be useful to do reasoning about. The context of "Who is the most distinguished captain of the Enterprise?" could be about the real world US Navy, or a fictional Star Trek universe (of which there are multiple).
- JimmyRuska 7y agoCheckout datomic, you can query for facts given a certain point in time. This is common among customer preference storage too, like, "what's your favorite song or movie can change easily over time." Being able to query the state at a certain date can be helpful. Some predictive analytics can also be done like people with these preferences, how did they change once they had a family, or moved to another "life stage". Though you probably don't need datomic, it would be not too complicated to model this in neo4j or some other RDF graph that supports arbitrary sized tuples. Datomic just supports this feature as a first class value offering.
- mdlincoln 7y agocontext-dependent, or "reified" assertions are a pain point for sure. I come from the perspective of cultural heritage data, where context is king. Which expert made this attribution for this painting? Who owned it _when_? According to which archival document? etc. Almost all the engineering problems cited in the original post are still basically there, but graphical models are still the least painful way of doing this, particularly when trying to share data between institutions. Example: https://linked.art/model/assertion/ https://linked.art/model/assertion/
- abathur 7y agoI'm really interested in semantic authoring (not really structuring data with semantics--but marking semantics within running text), though I guess I'm disinterested in the semantic web. I agree with a lot of the problems noted in other posts, and would add two other problems from the authoring side: 1. Identifying and employing sound semantics requires a level of thought and clarity that I don't think most people are habituated to working at. It raises the bar somewhat on who can be contributing (either they have to understand and take care with the semantics, or you need a separate person to handle them?) 2. I may be missing some good tools, but I haven't been able to find a good low-friction semantic authoring experience. Even if you are mentally prepared to write with explicit semantics, it still adds a lot of friction to the writing process (or requires subsequent semantic-edit passes).
- austincheney 7y agoWhen writing data structures that are not for describing or defining services I still can't help but think in triples. I also can't help but think of each data facet as though it were something described with meta-data would provide sufficient context that it would make sense if it were read out loud to a stranger.
- lyrachord 7y agoHow to describe semantics? Seems nobody is intensely aware of that at all. WTF? the language, of course. U must know how many forms of languages in semantics domain. XML will die, JSON after. huh? json5, jsn, ...... shows that, and proves that none is the best. BTW it is just like the papermaking technology, albeit Earth rotates without paper.
- buboard 7y agomodern NLP makes the semantic web completely obsolete. if anything, you need less markup because it's confusing and more often than not, just wrong.
- drongoking 7y agoThis is too extreme. If, like Google, you have a flock of Ph.D.s who you can put onto an NLP problem to extract semantics from text, then semantic markup becomes less valuable. Not all of us are in that situation. And I don't think parsing text is the only application of the semantic web. Having hugs databases full of knowledge is interesting in itself. As for semantic markup being confusing and usually wrong, I don't know where you get that.
- buboard 7y agoYeah, but i think there is a difference between standardized markup data formats describing e.g. proteins, and generic text with annotations. The latter are redundant
- sproketboy 7y agoIf you want Semantic web all you need to do is add rel=Data to the link tag. The browser would use that to fetch the json data for the page just as it uses link to fetch styling. Now that I've solved this for you, anything else?
- JimmyRuska 7y agoSemantic web tech solves a common problem. You have a database where you want to have some shared schema among many groups, and you want a way to infer facts based on first order logic. You want to be able to query multiple sources and reason about facts when taking into account multiple sources. Whether you use semantic web tech or not that's still a common problem that doesn't always have a good plug and play solution. There's still a lot of places using jsonld format for metadata and cataloging information. You can google cooking recipes and get ratings, cook time; search for movies and see how high rated the movie is and who made it with a synopsis of the plot, all of these are product metadata powered by rdfs or jsonld metadata, a relic of the semantic web. It would be incorrect to say semantic web is dead. Any AI that can effectively use wikidata as a fact table would be jeopardy grade. There's still new tools coming out like RDFox that apply first order logic at multicore speed across huge datasets for reasoning. There is work being done to make it horizontally scalable. I think people will just go on an endless loop of getting the same pain points and creating new tools using the trending tech of the day, but even in this day and age, sometimes something like prolog or picat is what you need.
- zozbot234 7y ago> you want a way to infer facts based on first order logic Isn't that computationally infeasible? Semantic web standards are based on description logics, i.e. multi-modal logics chosen specifically for computational expediency. Also, I wouldn't describe JSON-LD as a "relic" of anything. It's a fairly recent standard in the grand scheme of things, and many interesting projects these days implicitly rely on it.
- aggerdom 7y agoSo not an expert in this area, would love if someone corrects me. My understanding is generally FOL is infeasible. Propositional logic even can be computationaly difficult [1]. My understanding is that most of the semantic web stuff is done using a description logic of some flavor. These will be named based on the properties of the logic. The important thing is that they are generally decidable, and you can use something like MALET or some other solver to infer things from your database or ontology.(You give up some expressivity for decidability) Not sure how much stuff is going on with that these days. Played with a petrology ontology in protegé some back in college, but haven't followed the space. I remember OWL being important, but can't remember why at the moment. [1] For example if you try to figure out if a formula is satisfiable. You can for sure do this using truth tables. The catch is that you're looking at 2^n complexity where n is the number of propositions in your formula.
- tannhaeuser 7y agoI wouldn't call semweb dead; it just has found its niche(s) and is even stabilizing and gaining in those areas. I actually landed a gig for graph DBs, SPARQL, etc. in lab informatics for bio/chem. Earlier this year I attended a keynote held by Wikimedia Deutschland's Franziska Heine pushing for large publicly available RDF data sets, etc.