41 ms·
Make the “semantic web” web 3.0 again – with the help of SQLite
- _ea1k 5y agoIs there really that much web safely exposable data in sqlite for this to make sense? I'm not really seeing how this is obviously better than the metadata ideas that preceded it.
- rossdavidh 5y agoSome: weather, ratings, topography, dictionaries and encyclopedias, sports scores, market prices, some other stuff. All public knowledge, but not necessarily publicly available (easily) in raw form.
- lostmsu 5y agoBut doesn't it have to be immutable for the proposal to work?
- vorpalhex 5y agoNo, as long as the data model is stable you can add new rows. You might want some kind of versioning for messy column changes, particularly removals.
- lostmsu 5y agoI do not understand how that would work. The clients don't have any way to synchronize with server changes, so they can read data in an inconsistent state.
- regularfry 5y agoI think the way you'd have to do it is to effectively publish new database versions to their own path. Symlinked as much as possible, so they can back onto the same database, but you'd do something like this: http://host/db/1.0.0 http://host/db/1.0.0 is your original database. http://host/db/1.0.1 http://host/db/1.0.1 is what you publish when you fix a bug. http://host/db/1.0 http://host/db/1.0 could be a redirect to /1.0.1. http://host/db/1.1.0 http://host/db/1.1.0 is what you create when you add a new column. It's backwards compatible with /1.0.*, so you can either leave those paths working or you can redirect them to /1.1.0, depending on what guarantees you want to be making to your clients. http://host/db/2.0.0 http://host/db/2.0.0 is the version with an old column deleted, but you'd want to check in your access logs that nobody was still requesting any 1.0 version before publishing it. Either way, when this gets published you want to stop serving the /1.0.* path because 1.0 and 2.0 now can't come from the same backing file. But 1.1 and 2.0 can come from the same file if you've given everyone time to stop using the deleted column. It's not a great scheme, but it does give you a way to get new client connections onto the right version. For clients that have a session which lives across database upgrades, I think what you'd need is a `schema_version` table the client could check however often makes sense, and let themselves get reset if they find there's a new version available.
- lostmsu 5y agoThat can't really work for any kind of data, that is updated often enough.
- regularfry 5y agoIt works for the data itself just fine. Hopefully schema changes aren't terribly frequent.
- lostmsu 5y agoWhat do you mean "works fine"? AFAIK to read consistent data you need to be reading from a database snapshot. So with every data update on the server you need to make a new snapshot (to include new data), and publish it under new version (otherwise clients might be in the middle of some query reading data from the old snapshot, and combine metadata from the old snapshot with data from the new one).
- regularfry 5y agoSQLite already supports multiple-readers/single-writer. You don't need a new snapshot to add data.
- lostmsu 5y agoOne of us does not understand how that works. According to me you need to acquire read/write locks in order to ensure consistency in presence of a writer. You can't do that by reading ranges of a static file with independent read requests.
- _ea1k 5y agoI get that, but his argument is that you could do it with no additional work. "The beauty of this technique is that you are already using SQLite because it's such a powerful database; with no additional work, you can throw it on a static file server and others can easily query it over HTTP." I doubt very many of those already use sqlite. As soon as the additional work is added, its probably easier to just expose it as XML or JSON.
- bokchoi 5y agoI never got on the semantic web train, but a translation layer does allow you to make underlying schema changes. I poked around the ANSIWAVE BBS and it looks fun!
- 0xbadcafebee 5y agoAmong the 30-odd technologies that make up the Semantic Web[1] (it never died, it's just a collection of tech, lots of organizations use it daily) are graph databases[2]. Graph databases are necessary to implement semantic web databases. SQLite is not a graph database. Even if you used SQLite to implement a graph database, it would not solve any significant problems of the semantic web, such as access to data, taxonomies, ontologies, lexicons, tagging, user interfaces to semantic data management, etc. It's a really odd suggestion that you would just copy around a database or leave it on the internet for people to copy from. For the BBS mentioned here, that might actually be illegal, as it might contain PII, and on other sites possibly PHI. Many countries now have laws that require user data to remain in-country. Besides the challenges of just organizing data semantically, there still needs to be work done on data security controls to prevent leaking sensitive information. The funny thing is, that isn't even hard to do with the semantic web. You classify the data that needs protecting and build functions and queries to match. You can tie that data to a unique ID so that people can "own" their data wherever it goes, and sign it with a user's digital certificate which can also expire. But all of that (afaik) doesn't exist yet. Everyone is more concerned with blockchains and SQL, either because the fancy new tech is sexier, or the old boring tech doesn't require any work to implement. The Semantic Web never caught on because it's really fucking hard to get right. No companies are investing in making it easier. Maybe in 20 years somebody will get bored enough over a holiday to make a simple website creation tool that implicitly creates semantic web sites that are easy to reason about. It'll probably be a WordPress plugin. [1] https://en.wikipedia.org/wiki/Semantic_Web https://en.wikipedia.org/wiki/Semantic_Web [2] https://graphdb.ontotext.com/documentation/enterprise/introduction-to-semantic-web.html https://graphdb.ontotext.com/documentation/enterprise/introd...
- sekao 5y ago> Graph databases are necessary to implement semantic web databases. The online docs (and TBL himself) rarely mention of graph databases, but obviously the idea is tied tightly to RDF. Separating it from that implementation detail is part of the point, though. Getting people to represent their data via an additional format was never going to work. > For the BBS mentioned here, that might actually be illegal, as it might contain PII Can't imagine the purpose you had in even making this point. In theory, any arbitrary database exposed publicly could be illegal to replicate due to copyright, PII laws, etc. But that has nothing at all to do with a technical discussion of a technique for exposing data. What a bizarre point to make. As an aside, I'm glad you removed the "Uh........." from the beginning of your post. We're all making an effort to reduce the typical HN snark in the comments, and there's always room for improvement :D
- luhn 5y agoThe author seems to assume that everybody is using SQLite, but SQLite for a production database is an extremely niche choice. Attempting to expose more popular options like PostgreSQL or MySQL as SQLite would be extremely difficult because SQLite only supports a subset of SQL, whereas PostgreSQL and MySQL both implement their unique superset (for the most part) of SQL. But it doesn't matter. The API doesn't matter. Web 3.0 was never about APIs, it was about data. A standardized API is only useful if it outputs standardized data. Having a bunch of bespoke SQLite tables scattered across the web gets us no closer to the ideal of Web 3.0.
- jumpkick 5y agoI thought Web 3.0 was about NFTs.
- dspillett 5y agoPeople are talking that way now. Before what is being called “Web3” ATM Web 3.0 was to be the “semantic web” (where Web 2.0 was the “interactive web” - with many things become read/write instead of read-only and greater interactivity, both in terms of social interactivity and individual interactive apps being web based, being the focus of technological enhancements). This article is talking about that earlier definition, and a way it might once again be the definition, perhaps relegating Web3 to Web 4.0 (or web we-worked-out-it-was-a-ponzi-scheme-so-just-stopped with something else being Web 4.0, if you take the more cynical view).
- tyingq 5y ago"Before the term was hijacked by crypto-grifters (and, admittedly, a few genuinely neat projects), web 3 (point oh) referred to Tim Berners-Lee's project to promote a standard way to expose and parse metadata on the web."
- Closi 5y agoSQLite is the most used database engine in the world, so I wouldn't call it niche. In fact, by some estimates, it is probably used more than all other database engines combined. The only difference is that it is usually run locally (compared to Postgres and your other examples), but something doesn't have to run remotely to be considered running in production :)
- xmly 5y agoI do not understand this conclusion: "Data on the web will only be "semantic" if that is the default, and with this technique it will be." Why would it be semantic?
- SkeuomorphicBee 5y agoBecause the backend data is exposed to the world, with all its original semantic structure (relational model) intact, before it is flattened into a document view.
- visarga 5y agoReal world is messy, some companies have different key-value pairs on the same kind of document (invoice, purchase order, utility bill, etc). I counted 20K different keys, some semantically synonymous, in a few thousand invoices. Even the table part can have different columns. What do you do when schemas don't quite match?
- mk81 5y ago
- alexchamberlain 5y agoGenerally relational models are _not_ semantic models.
- deleted 5y ago[deleted]
- rpwverheij 5y agoI agree. Semantic data would mean that others can easily understand the meaning of the data. In the semantic web this would be by using ontologies, which define the types of things and relations between these types of things. But just having your schema visible doesn't mean anyone understands it straight away, they would still need to make an effort to understand the schema of that specific application. And the schema is probably unique to that application. The end result is pretty much the same as for example exposing your database as a GraphQL endpoint. Take "The Graph" a web3 project exposing data of many different blockchain projects as GraphQL endpoints. It's nice, but I still need to make an effort to understand the meaning of each property in each endpoint. And a "transaction" in one project is not linked to the meaning of a "transaction" in another. A bit off topic, but ironically I therefor don't find the name 'The Graph' to be all that accurate. Point in case: YES to at least remembering that web3 is (also) the Semantic Web. But no, this solution is not semantic data.
- fleddr 5y agoThe semantic web is not a technical problem, it's an incentive problem. RSS can be considered a primitive separation of data and UI, yet was killed everywhere. When you hand over your data to the world, you lose all control of it. Monetization becomes impossible and you leave the door wide open for any competitor to destroy you. That pretty much limits the idea to the "common goods" like Wikipedia and perhaps the academic world. Even something silly as a semantic recipe for cooking is controversial. Somebody built a recipe scraping app and got a massive backlash from food bloggers. Their ad-infested 7000 word lectures intermixed with a recipe is their business model. Unfortunately, we have very little common good data, that is free from personal or commercial interests. You can think of a million formats and databases but it won't take off without the right incentives.
- onion2k 5y agoEven something silly as a semantic recipe for cooking is controversial. Somebody built a recipe scraping app and got a massive backlash from food bloggers. Their ad-infested 7000 word lectures intermixed with a recipe is their business model. Taking someone else's content and republishing it without permission isn't cool, even if you wrap it in a nice machine readable format.
- fleddr 5y agoI fully agree, and that's one of the problems I was describing. There's very little content free of commercial interests. If this is true, it blocks a lot of potential use cases of a semantic web.
- Micoloth 5y agoI'm more and more seeing that this is true. Still, it is sad. The question isn't even, what can one do, because obviously nobody can change how incentives works in a given society. The question is: is there a timeline in which the right incentives (to share data) start being enforced? How would that play out?
- OliverJones 5y ago> The semantic web is not a technical problem, it's an incentive problem. True. Demonstrable in the health-care IT world. Think of electronic health records. My personal portable electronic health record would either be a bunch of images of scrawled notes and maybe some nice medical images ( = nonsemantic web) Or it would be in a highly wrought format, i dunno, XML or something, with carefully worked out schemata for everything from flu shot records to heart transplants (= semantic web). Back in 2007-2010, "electronic health records" EHR were spottily and sloppily implemented by some providers. But, in the US, a federal law pushed more widespread implemetation. Now my online EHR, and yours, is decidedly app-mediated and non-semantic, on a web site portal. Export to JSON? Hah. No. The hospitals and health care systems only did it because of incentives. I happened to work at a B2B SaaS company focused on making connections between hospitals and rehab/skilled nursing providers. A rehab outfit can't decide to accept a patient without seeing her medical records and doctors' orders. So our customers had a real incentive to be able to share records. It worked. But the data we had access to (go read about HL7) was not even close to semantic. And our SQL database schemas were, umm, quinquiremes of Nineveh, really intricate, somewhat brittle. Let's leave privacy issues out of the conversation for a moment. Publishing the schema and accepting random queries would help NOBODY except some partner outfit willing to develop and test useful stuff. Hey, I got an idea! Let's give them an API! Oh, wait, nobody wants to bother with an API? OK, how about a nice web site! And we're back where we started. With a universal semantic web, the same problems would crop up everywhere.
- echelon 5y agoWhat a lot of folks don't realize is that the Semantic Web was poised to be a P2P and distributed web. Your forum post would be marked up in a schema that other client-side "forum software" could import and understand. You could sign your comments, share them, grow your network in a distributed fashion. For all kinds of applications. Save recipes in a catalog, aggregate contacts, you name it. Ontologies were centrally published (and had URLs when not - "URIs/URNs are cool"), so it was easy to understand data models. The entity name was the location was the definition. Ridiculously clever. Furthermore, HTML was headed back to its "markup" / "document" roots. It focused around meaning and information conveyance, where applications could be layered on top. Almost more like JSON, but universally accessible and non-proprietary, and with a built in UI for structured traversal. Remember CSS Zen Garden? That was from a time where documents were treated as information, not thick web applications, and the CSS and Javascript were an ethereal cloak. The Semantic Web folks concurrently worked on making it so that HTML wasn't just "a soup of tags for layout", so that it wasn't just browsers that would understand and present it. RSS was one such first step. People were starting to mark up a lot of other things. Authorship and consumption tools were starting to arise. The reason this grand utopia didn't happen was that this wave of innovation coincided with the rise of VC-fueled tech startups. Google, Facebook. The walled gardens. As more people got on the internet (it was previously just us nerds running Linux, IRC, and Bittorrent), focus shifted and concentrated into the platforms. Due to the ease of Facebook and the fact that your non-tech friends were there, people not only stopped publishing, but they stopped innovating in this space entirely. There are a few holdouts, but it's nothing like it once was. (No claims of "you can still do this" will bring back the palpable energy of that day.) Google later delivered HTML5, which "saved us" from XHTML's strictness. Unfortunately this also strongly deemphasized the semantic layer and made people think of HTML as more of a GUI / Application design language. If we'd exchanged schemas and semantic data instead, we could have written desktop apps and sharable browser extensions to parse the documents. Natively save, bookmark, index, and share. But now we have SPAs and React. It's also worth mentioning that semantic data would have made the search problem easier and more accessible. If you could trust the author (through signing), then you could quickly build a searchable database of facts and articles. There was benefit for Google in having this problem remain hard. Only they had the infrastructure and wherewithal to deal with the unstructured mess and web of spammers. And there's a lot of money in that moat. In abandoning the Semantic Web, we found a local optima. It worked out great for a handful of billionaires and many, many shareholders and early engineers. It was indeed faster and easier to build for the more constrained sandboxiness of platforms, and it probably got more people online faster. But it's a far less robust system that falls well short of the vision we once had.
- firechickenbird 5y agoIsn’t this Web 1.0 instead? You are only reading data, yeah ok with sql, but you still can’t modify it. And also there are already very good standards like Rdf, Owl2, spraql, which are more expressive than sql for consuming the info
- netcan 5y agoWhether or not it has legs, at least this is an interesting idea.
- recursivedoubts 5y agoHumans, as of now (and as far as I'm aware, being outside the AI labs at the big tech companies and DARPA) have agency, and so are in a unique position to take advantage of the uniform interface of REST/the web in a flexible manner. I wrote an article about this on the intercooler.js blog, entitled "HATEOAS is for Humans": https://intercoolerjs.org/2016/05/08/hatoeas-is-for-humans.html https://intercoolerjs.org/2016/05/08/hatoeas-is-for-humans.h... The idea that metadata can be provided and utilized in a similar manner doesn't strike me as realistic. If it is code consuming the metadata, the flexibility of the uniform interface is wasted. If it is a human consuming the metadata, they want something nice like HTML. For code, why not just a structured and standardized JSON API? This appears to be what we have settled on, and I don't see any big advantage extending REST-ful web concepts on top of it. The machines just ignore all that meta-data crap.
- netcan 5y ago>> why not just a structured and standardized JSON API? So in this version of the idea... because structuring data requires work. Unstandardized data exists already. Some of it is already SQLITE. A lot of the rest is in other SQLs, and that might be a smaller bridge. Author claims (if I'm understanding correctly) that a static website could easily query sqlites over HTTP, and bam, web 3.0. Honestly, it's hard for me to think/discuss these ideas without examples, even if contrived. What kind of websites would be built this way? What data will they be querying? A web app that uses photos and address books on the users phone? An alternative UI for news.yc?
- 5y ago
- sharperguy 5y agoI wonder if combining this idea with some kind of microtransactional currency such as the bitcoin Lightning Network or even a simple Chaumian e-cash system (1) would help to get around the issue of requiring clickbait, advertising and SEO with every single piece of data. Would be great if providers could offer data in raw form without the overhead of all the gunk that gets them paid. 1. https://en.wikipedia.org/wiki/Ecash https://en.wikipedia.org/wiki/Ecash
- vmception 5y agoif "exposing and parsing metadata" never took off as the meaning for web 3.0, by the author's own admission, why try to resurrect it with this title? clickbait? we love web3 labeled articles here fortunately the more popular variant of web 3.0 doesn't even need the developer to make a database or anything on the backend. just frontend development, and deploying code once to the nearest node. frontend is optional depending on your userbase.
- fjfaase 5y agoSQL is not better than XML or JSON for representing data. They are all mappings of much richer data structures on a limited data model. But even setting aside these problems, there are some problems with a distrubted semantic web that are barely ever mentioned: the step from going from data to 'semantic' facts, how to deal identifying sources, and versioning/updates. I think it is very important to record who (person or institution) is the source of a certain fact or the 'linking' of facts between multiple sources. Cryptographic keys, just as in blockchains, could help to link data of distruted sources such that it is possible to verify the source of a fact to sources/authorities and correct errors or deal with updates in case they occur.
- OliverJones 5y agoThere's one exception to the equivalence of SQL on the one hand to XML or JSON on the other hand. The point of SQL (and other DBMS paradigms) is to give access to data that's orders of magnitude bigger than the RAM in which the app runs. That has stayed true for at least a quarter century, during which RAM and database sizes both did the Moores-law exponential expansion.
- fjfaase 5y agoA relational database maps everything to unordered relationships. Representing or manipulating a tree like structure is complex. Just representing an ordered list is complex. In XML and JSON everything is ordered and querying it as relational database is cumbersum. Graph databases and OO databases are somewhere in the middle. But as I wanted to point out, which data models are used, is not the major obstacle to the semantic web. It is these other problems that are not addressed.
- KarlKemp 5y agoI feel Wikidata has a generally sane approach to these questions.
- fjfaase 5y agoWikidata is interesting. But it is a centralized approach. Is there an interface which gives the full breakdown of sources and where you can, through a chain of certificates (like those used for ssh and https), verify the sources?
- fiatjaf 5y agoWhat is this data we want to semantically link by the way?
- punnerud 5y agoThis post came 3months before phiresky, and should get credit for being first to making it practical: https://news.ycombinator.com/item?id=25842999 https://news.ycombinator.com/item?id=25842999
- fizx 5y agoEffectively, this is a pretty cool way to get an OSS version of something like https://www.snowflake.com/guides/public-data https://www.snowflake.com/guides/public-data. Also something like a lighter version of https://datasette.io/ https://datasette.io/. The thing it's kinda missing for me is the ability to compose multiple SQLite databases, possibly provided by different domains. It'd be nice to join together different public datasets. In a weird personal example, if Strava exposed SQLite, I'd love to do a join to weather.com and see when the last time I biked in the rain was. It'd be cool if one half of some table was at foo.com and I could add a few rows to it on my bar.com domain, and then the combined dataset was queryable as a single unit.
- simonw 5y agoI've done some thinking about joining datasets together. Datasette grew a `--crossdb` option a while back, which means that if you attach multiple SQLite files to the same Datasette instance you can run joins across them: https://docs.datasette.io/en/stable/sql_queries.html#cross-database-queries https://docs.datasette.io/en/stable/sql_queries.html#cross-d... So one option is to download the database you want to join against and run the joins locally. Datasette encourages making the raw SQLite database file available, so if it's less than about 100MB this may be a good way to do it. If you're willing to do the joins in your client-side code, Datasette's default JSON API can help. You can write an application (including a client-side JavaScript application) which fetches and combines data from multiple different Datasette JSON instances by hitting their APIs. My last idea is the most out-of-left-field: since Datasette lets you define custom SQL functions using Python code, it would be feasible to create a Python function which itself makes a query via the JSON API against another Datasette instance! You could then use that to simulate joins in SQL queries that you run against a single Datasette instance. I've not built a prototype of this yet, and to be honest I think combining data fetched from multiple JSON APIs (which is possible today) will provide just-as-good results, but it's an interesting potential option.
- dgudkov 5y agoI have a similar idea with PDF documents. Instead of having the royal PITA of parsing generated PDFs (e.g. invoices), things would be much simpler if every generated PDF came with a built-in SQLite or JSON that contains the structured data of that PDF. One day I will do it. Speaking more broadly, whether we talk about HTML or PDF it's the same problem: documents should have two representations - human-friendly and machine-friendly until AI gets so good that only having the human-friendly representation is enough.
- ttyprintk 5y agoIf you want to start with the most controllable representation of a piece of paper, consider that OpenOffice (word processing and drawing) embeds its structured file format into PDFs. Maybe those are the PDFs we scrape first, leaving the bag-of-jpegs for later.
- daveydave 5y agoI have an application that converts word documents to RDF conformant with the SPAR ontologies (mainly DoCO http://www.sparontologies.net/ontologies/doco http://www.sparontologies.net/ontologies/doco), so it contains things like headers, numbering, contains/within relationships explicit in the RDF. I've used it successfully with PDFs by converting to DOCX first. Is this the sort of thing you had in mind? Not here to sell it! I think this is a genuinely interesting unexplored area ..
- dgudkov 5y agoThe PDF format supports attachments (embedded files). I'm thinking about a set of libraries and/or a command-line utility that would make it trivially easy to attach a SQLite|JSON file to a PDF or extract one from a PDF. This won't fix existing files, of course, but at least for those apps that generate PDFs it will be easier to embed a SQLite/JSON into a generated PDF.
- oever 5y agoXMP is meant for adding semantic information to (parts of) PDF files. https://en.wikipedia.org/wiki/Extensible_Metadata_Platform https://en.wikipedia.org/wiki/Extensible_Metadata_Platform
- tingletech 5y ago"The semantic web is the future of the web, and always will be." -- Peter Norvig https://www.youtube.com/watch?v=LNjJTgXujno&t=1257s https://www.youtube.com/watch?v=LNjJTgXujno&t=1257s
- deleted 5y ago[deleted]
- z3t4 5y agoOne problem is that it's the one hosting the data that pays the bandwidth. Yes when you download a video from Youtube, Google has to pay your ISP! (Google will strong-arm peering agreements though, but that doesn't take away from my point) Someone have to pay for the infrastructure. Right now the one hosting (not the consumer) pays for the infrastructure. So there are not really any incentive to host data for free - like a kiosk offering free goods and services. The problem is if you would get paid by sending stuff, everyone would be spamming data everywhere. Imagine if you would have to pay 10c every time someone sent you an e-mail. Something that I think would help are micro transactions. And a built into browsers so that you could easily make a micro-transaction. We already have Bitcoin and other crypto currencies, but they are too big to run inside the browser of a mobile phone, if it wasn't for the high transaction costs - the ledger/blockchain would be even bigger... Today publishers - those that publish stuff on the web earn money by showing ads. And ads initially worked very well for a few years around 2000 before people started cheating with bots. But you can still make individual deals with webmasters and choose to trust them. Also a lot of "the web" has moved to videos and Youtube. The average web user choose to watch a video rather then reading a text article covering the topic of interest.
- chupchap 5y ago> Google has to pay your ISP This is the standard ISP arguement that I never understood. Google pays their ISP and any other ISPs they are using for peering, but they don't HAVE to pay your ISP. You pay for your usage.
- lobocinza 5y agoI as a consumer also have to pay for Internet access.
- fizx 5y agoYou can host this on a requestor pays S3 bucket.
- kdunglas 5y agoAPI Platform is a popular and easy to use semantic web framework: 1. You design your data model as a set of PHP classes, or you generate the class from any RDF vocabulary such as Schema.org 2. API Platform uses the classes to expose a JSON-LD API with with all the typical features (sorting, filtering, pagination…) 3. You use the provided "smart clients" to build a dynamic admin interface or to scaffold Next, Nuxt or React Native apps (these tools rely on the Hydra API description vocabulary, and work with any Hydra-enabled API) In addition to RDF/JSON-LD/Hydra, API Platform also supports ActivityPub. https://api-platform.com https://api-platform.com
- no_wizard 5y agoAren't you the same Dunglas that is head of this project? Great idea by the way! had to the pleasure of working with api platform in some Symfony applications in a past gig. I can vouch its easy enough to use, but the GraphQL integration (at least at that time) was really slow. I have not found PHP to be the ideal runtime for GraphQL
- serverholic 5y agoIt's clear that people want web apps, not the semantic web. I really don't see why people care so much about this.
- simonw 5y agoI've been exploring the idea of using SQLite to publish data online via my Datasette project for a few years now: https://datasette.io/ https://datasette.io/ Similar to the OP, one of the things I've realized is that while the dream of getting everyone to use the exact same standards for their data has proved almost impossible to achieve, having a SQL-powered API actually provides a really useful alternative. The great thing about SQL APIs is that you can use them to alter the shape of the data you are querying. Let's say there's a database with power plants in it. You need them as "name, lat, lng" - but the database you are querying has "latitude" and "longitude" columns. If you can query it with a SQL query, you can do this: select name, latitude as lat, longitude as lng from [global-power-plants] Here's a demo using exactly that query: https://global-power-plants.datasettes.com/global-power-plants?sql=select+name%2C+latitude+as+lat%2C+longitude+as+lng+from+%5Bglobal-power-plants%5D+limit+100 https://global-power-plants.datasettes.com/global-power-plan... That URL gives you back an HTML page, but if you change the extension to .json you get back JSON data: https://global-power-plants.datasettes.com/global-power-plants.json?sql=select+name%2C+latitude+as+lat%2C+longitude+as+lng+from+%5Bglobal-power-plants%5D+limit+100&_shape=array https://global-power-plants.datasettes.com/global-power-plan... Or use .csv to get back the data as CSV: https://global-power-plants.datasettes.com/global-power-plants.csv?sql=select+name%2C+latitude+as+lat%2C+longitude+as+lng+from+%5Bglobal-power-plants%5D+limit+100 https://global-power-plants.datasettes.com/global-power-plan... But what if you need some other format, like Atom or ICS or RDF? Datasette supports plugins which let you do that. I'm running the https://datasette.io/plugins/datasette-atom https://datasette.io/plugins/datasette-atom datasette-atom plugin on this other site. That plugin lets you define atom feeds using a SQL query like this one: select issues.updated_at as atom_updated, issues.id as atom_id, issues.title as atom_title, issues.body as atom_content, repos.html_url || '/issues/' || number as atom_link from issues join repos on issues.repo = repos.id order by issues.updated_at desc limit 30 Try that query here: https://github-to-sqlite.dogsheep.net/github?sql=select%0D%0A++issues.updated_at+as+atom_updated%2C%0D%0A++issues.id+as+atom_id%2C%0D%0A++issues.title+as+atom_title%2C%0D%0A++issues.body+as+atom_content%2C%0D%0A++repos.html_url++%7C%7C+%27%2Fissues%2F%27+%7C%7C+number+as+atom_link%0D%0Afrom%0D%0A++issues+join+repos+on+issues.repo+%3D+repos.id%0D%0Aorder+by%0D%0A++issues.updated_at+desc%0D%0Alimit%0D%0A++30 https://github-to-sqlite.dogsheep.net/github?sql=select%0D%0... The plugin notices that columns with those names are returned, and adds a link to the .atom feed. Here's that URL - you can subscribe to that in your feed reader to get a feed of new GitHub issues across all of the projects I'm tracking in that Datasette instance: https://github-to-sqlite.dogsheep.net/github.atom?sql=select%0D%0A++issues.updated_at+as+atom_updated%2C%0D%0A++issues.id+as+atom_id%2C%0D%0A++issues.title+as+atom_title%2C%0D%0A++issues.body+as+atom_content%2C%0D%0A++repos.html_url++%7C%7C+%27%2Fissues%2F%27+%7C%7C+number+as+atom_link%0D%0Afrom%0D%0A++issues+join+repos+on+issues.repo+%3D+repos.id%0D%0Aorder+by%0D%0A++issues.updated_at+desc%0D%0Alimit%0D%0A++30 https://github-to-sqlite.dogsheep.net/github.atom?sql=select... As you can see, there's a LOT of power in being able to use SQL as an API language to reshape data into the format that you need to consume.
- gibsonf1 5y agoWell, then again the original idea is taking off with https://solidproject.org/ https://solidproject.org/ with millions of pods by Tim Berner-Lee's Inrupt to go online starting this Spring.
- gfody 5y agogross sparkle queries
- indigodaddy 5y agoLooks neat thanks for the link. Also appears to be fairly straightforward to self host a pod: https://solidproject.org//self-hosting/css https://solidproject.org//self-hosting/css
- rapnie 5y ago> to go online starting this Spring Any announcement I missed? Solid project exists for a long time and seems that many specs are still very early days.
- gibsonf1 5y agoYes, the entire country of Flanders is getting pods for every citizen in March. Then there are patient pods for UK NHS, and then pods for BBC content users...
- seumars 5y agoThis could be a nifty way of getting RSS back.
- wombatmobile 5y ago> The semantic web will never happen if it requires additional manual labor. Is manual labor the reason things turned out the way they did, with google spending whatever it took to index and monetise the whole web the way it did? Or might money have something to do with it?
- dustractor 5y agoRange requests. Hmm. That would lead to some interesting semantics.
- togaen 5y agoTerrible idea. Why would anyone want to deal with interfacing a bunch of randomly structured databases whose tables can change at any time without warning. Nightmare.
- scotty79 5y agoNot when compared to alternative of accessing random, misdocumented, randomly limited, arbitratily formatted subset of that which slowly bit-rots.
- FridgeSeal 5y agoAs opposed to a bunch of websites serving an archaic, poorly formatted blob of text-the “correct” parsing of which has now become _so complicated_ that it’s basically infeasible for anyone not willing to build a whole web-browser?
- cxr 5y agoParsing is not the hard part of dealing with the Web platform.
- regularfry 5y agoIt's certainly not the easy part.
- deleted 5y ago[deleted]
- luckystarr 5y agoYes, it's still terrible for the consumer of the data. But I like it not because of that. The positive thing I feel when reading about this, is that it dramatically lowers the barrier for the producer of the data to expose it in a meaningful way. While previously it was necessary to think about the format and write code to expose the data while now its possible to just throw the data over a wall. You could use a framework to automate the first thing, but this would be specific to one programming language, while the second approach works with all languages. So it lowers the total effort to get to the goal, effectively side-stepping the "have to implement framework or serialization code" issue. Warning: heavy speculation below So if more people would build sites using this technique, the pressure for better tools (at a higher level than right now) for consumers would increase, so these would be built by someone. As you have a proper standard (there is only one SQLite) you would have a new "ecosystem" growing. This would lower the pain for the consumers of said data. You'd still have to implement it in every programming language that wants to access the data, but this is another problem.
- pietroppeter 5y agoVery nice indeed! I am sorry I did not notice before the discussion about previous blogpost on the subject [0] “Using the SQLite-over-HTTP "hack" to make backend-less, offline-friendly apps” Are there more than 2 blogposts? Cannot find a posts page. [0]: https://news.ycombinator.com/item?id=29758613 https://news.ycombinator.com/item?id=29758613
- SergeAx 5y agoWhat is the difference between this idea and exposing a read-only MongoDB (JSON in, JSON out) via HTTP endpoint? In the end, 50 ranged HTTP requests is not that great for unstable connection, 1 request is way better, and you need a server anyway. Okay, offline first, what does that mean? Should I download the entire 600mb SQLite database? Should I do it every time it changes? Who will pay for the bandwidth? We can not employ standard HTTP proxy and caching mechanism here, it is not for 600mb files.
- sekao 5y agoMaking it available via a static file host dramatically lowers the barrier. If you have an interesting dataset, you can even throw it on github pages and pay nothing; that likely is not true for your mongo db server. A typical request using indexes will be less than 10 separate 1KB GET requests, not 50. But yeah, more work needs to be done on performance. Whether it makes sense to fully download the dataset depends on the project; maybe it does not. But it doesn't have to be a monolithic file. You can use SQLite's multiplex VFS to split the SQLite file into many smaller pieces (and still update the db later!).
- SergeAx 5y agoI think this can be carefully thought through and become something interesting. I don't fancy SQL, to be honest, despite having "structural" in its name it is too chaotic for XXI century. Rows and columns! Schema-enforced document databases, on the other hand, are neat and mostly people- and machine-readable at the same time. Another idea: downloading only indexes may greatly reduce number of requests needed to query the data.
- ricardobeat 5y agoThat’s what those tiny 1KB requests are doing mostly, downloading indexes. With http pipelining taken into account, it’s way faster than trying to preload the full index data. Despite the number of requests seeming excessive to us, the performance of this setup is already in the ballpark of your usual underpowered MySQL going through an app server.
- uoaei 5y agoThis may be more likely to happen if there was a compromise between the two: query the database, but maintain a database that is queryable using SPARQL and can export to TTL files. Then the linked data revolution can continue and we don't have to maintain finnicky webpages but rather a relatively static database.
- __MatrixMan__ 5y agoBut how are we going to make sure that users see ads between each query?
- galaxyLogic 5y agoSQLite is a relational database. Wouldn't it be a better fit to use a graph-database as the backend for anything "web"? The idea is good, that a web-page should be generated from some data somewhere. But "web" is much about not a single document but the links between the documents, which allow you to to represent a "semantic net". The data should be about the links between them. Now where is such a database? And how can it "sharded" into multiple databases running in thousands of locations on the internet?
- mikestaub 5y agoOriginTrail.io is using Arangodb.com
- Groxx 5y agoWhat is a graph database? A miserable little pile of joins. Though to be serious: what do you expect a graph database to provide that sqlite cannot / does not do efficiently?
- patates 5y agoTreating links as an entity. In a RDBMS, it is not possible to map a bidirectional relation semantically. It needs to have a direction (you can define it in both directions but then you have 2 relations). You then also need to duplicate link properties (in the join table) or normalize to yet another join table. That has been my pet peeve for a while, and can make it hard to define navigation. Totally not a deal-breaker though. I'd still use sqlite because I ### love it :)
- eoo 5y agoNeeds cost-based querying paid for with lightning network to be viable at scale.
- regularfry 5y agoWhy is that uniquely the case here?
- visarga 5y ago> Data on the web will only be "semantic" if that is the default, and with this technique it will be. Not going to work unless imposed by some external force. The semantics of the web can more practically be extracted with neural nets, but it's a long tail and there are errors. Lots of good work recently in parsing tables, document layouts and key-value extraction. LayoutLM and its kin comes to mind.[1] [1] https://scholar.google.com/scholar?cites=9435785928704193879&as_sdt=2005&sciodt=0,5&hl=en https://scholar.google.com/scholar?cites=9435785928704193879...
- lmm 5y agoI'm not a fan of SQL, but I do think exposing your original source data in its original form is valuable (though it has little to do with being semantic). I carefully set up my blog to expose the raw markdown that is the source form of my blog posts in the source HTML itself, with the minimum necessary cruft around it to render it as a viewable webpage.
- deleted 5y ago[deleted]
- WolfOliver 5y agoHow does this relates to an headless CMS?
- phh 5y agoI think the author missed the "semantic" part. If you push your own SQLite, then no, I don't have the semantic meaning of the website. Only a standardized semantic file format ala RDF can achieve that.
- sekao 5y agoOut of context, you won't necessarily be able to glean meaning from either an arbitrary SQLite database or arbitrary RDF tuples. Both are equally meaningful or meaningless depending on the observer...at the end of the day, they are just structured data with labels that (hopefully) the observer understands. One doesn't have inherently more semantic meaning than the other.
- tzury 5y agoA more readable version https://outline.com/E5J2Ft https://outline.com/E5J2Ft
- jhoelzel 5y agoWell yes and no. I can see this working in theory, but in reality semantic means standardised as much as it means accessible. In a world where my blogpost objet has the same information as your blogpost object, this works without a problem. In a world where I actually want to up my database to you, we could agree on a format. Both of these cases, from where i stand, seem very unlikely and we have not even talked about the pople that would clone your data 1 to 1 just to host an ad filled alternative of your site in real time.
- shp0ngle 5y ago.... the original SQLite-over-HTTP-ranges was a clever hack to host database-like data on github. But I don't think it should be actually used for anything serious. And I don't really get the connection with "semantic web", which was essentially idealistic vaporware of the 2000s.
- hankman86 5y agoNot going to happen. The reason for the Semantic Web never taking off were never technical. Websites already spend a lot of money on technical SEO and would happily add all sorts of metadata if only it helped them rank better. Of course, many sites’ metadata would blatantly “lie” and hence, the likes of Google would never trust it. Re exposing an entire database of static content: again, reality gets in the way. Websites want to keep control over how they present their data. Not to mention that many news sites segregate their content as public and paywalled. Making raw content available as a structured and query able database may work for the likes of Wikipedia or arxiv.org. But it’ll not likely going to be adopted by commercial sites.
- moigagoo 5y agoFeel kinda disappointed that the blog isn't hosted on Ansiwave :-)
- hankman86 5y agoBtw, it’s funny how the failed “semantic” web is now labelled Web 3.0
- jillesvangurp 5y agoI think both the capital S Semantic Web and the lowercase semantic web (microformats) kind of just fizzled out towards the end of last decade without changing much at all on the actual web. The lower case variety kind of survives as a smart thing to do to "help" search engines a little but otherwise has very little real world relevance. All talk of doing anything with on page information in browsers evaporated a long time ago. E.g. MS had some plans with this with early versions of Edge and there were some nice extensions for Chrome and Firefox as well. Not a thing any more. Most of that got unceremoniously ripped out of browsers a long time ago. At this point it's basically just good SEO practice to use microformats as search engines can use all the help they need to figure out what is what on a page. Other than that, whether you render your data to a canvas, a table, or nice semantic HTML has very little relevance for anyone. It's all just pixels that hit your eyeballs in the end. There's nothing else that looks at that information. With the exception of search engines. And they were part of web 1.0 already. The capital S Semantic Web with ontologies, triple databases, etc. never really got out of the gates and is perpetually stuck in people doing very academic stuff or specialist niche stuff that largely does not matter to anyone else. The exception is graph databases, which are still used in some data/backend teams for some stuff. And of course a few of those also pay lip service to some of the Semantic Web W3C standards from the early 2000s even though that is not the main thing they do anymore. Either way, too much of a specialist thing to call it a semantic web (capital or lower case). Most of the web uses exactly none of this stuff. But nice tools to have if you need them. You could argue a lot of the people involved moved their focus to AI and machine learning, which certainly looks like it is having a very large impact on e.g. search engines. I guess web3 has that in common with web 3.0 (other than the number 3). There are a few people who desperately (and loudly) want the web to go their way and insist it must be the future. But most people couldn't care less. In the end people just vote with their feet and gravitate to technologies that work for them or solve a problem they have and ignore things that don't do anything useful for them. In the case of Semantic Web, there was nothing there that you could coherently explain (i.e. without using all sorts of abstractions, complex stuff, and simplistic hyperbole). There were a few startups and lots of hype. They did a bunch of stuff. Most of those startups no longer exist or have faded into irrelevance. And the few that survived carved out a few interesting niches but did not end up producing any mainstream, must have technology. Certainly no unicorns there. Wolfram Alpha probably is one of the more well-known ones that actually shipped something useful. But it's a destination and not the web. Web3 has the same issues. Most threads on HN on web3 devolve into people talking about what it is, ought to be, or isn't and why that is or isn't important. That seems to be impossible to do without using a lot of hyperbole and BS. Very little substance in terms of widely adopted technology or even in terms of what that technology looks like or should look like. It's Web 1.0 all over again. Step 1 Blockchain, Step 2: ????, Step 3: Profit (or not). Most of the web is just a slightly slicker version of what we had 15 years ago (web 2.0). AJAX definitely became common place. We now have mature versions of HTML, SVG, CSS, etc. that actually work. And with WASM we can finally engineer some proper software without having to worry about polyfills and other crazy hacks to make javascript do stuff it clearly is not very good at. I'm looking forward to the next 15 years. It's going to be interesting and possibly a wild ride.
- mro_name 5y agoI don't get it, why the data and the final visual have to be both present/created ON THE SERVER. There's been a technology around for so long, that it is forgotten meanwhile (like the semweb itself): xslt. A lot can be done by just publishing raw xml data plus a visual representation generated in the browser right before display. I'm doing so with RDF https://demo.mro.name/geohash.cgi/about https://demo.mro.name/geohash.cgi/about, GPX https://demo.mro.name/geohash.cgi/u154c https://demo.mro.name/geohash.cgi/u154c, homegrown xml http://rec.mro.name/stations/b2/2022/01/12/1005 http://rec.mro.name/stations/b2/2022/01/12/1005 or atom feeds https://demo.mro.name/shaarligo https://demo.mro.name/shaarligo and on and on. The server is a source of data, its filesystem the database, and the client has to make sense of it. There is no API but GET requests. Works wonders for all but big data queries, naturally. So you publish raw data (TimBL, you want it that way) plus a recipe for a visual representation and the browser shows a sensible view to begin with.