16 ms·
Extracting data from Wikipedia using curl, grep, cut and other bash commands
- vram22 10y agoThere is also a wikipedia library for Python. An example of its use: Using the wikipedia Python library (to search for oranges :) https://jugad2.blogspot.in/2015/11/using-wikipedia-python-library.html https://jugad2.blogspot.in/2015/11/using-wikipedia-python-li... And there maybe libraries for other languages too, since the above library wraps a Wikipedia API: https://en.wikipedia.org/wiki/Wikipedia:API https://en.wikipedia.org/wiki/Wikipedia:API
- ianseyer 10y agoThis is the kind of query that excites me for WikiData's development. http://wikidata.org http://wikidata.org
- lacksconfidence 10y agosee also the SPARQL search: https://query.wikidata.org/ https://query.wikidata.org/
- minimaxir 10y ago> You will not need to open an editor and write a long script and then to have an interpreter like Node.js to run it, sometime the bash command line is just enough you need! This is a bad attitude to have for working with data processing, where QA is necessary and the accuracy of the output is important. A 50 LOC scraper with comments and explicitly-defined inputs and output from functions is far preferable to a 8 LOC scraper that those without bash knowledge will be unable to parse. And the 8 LOC bash script is not much of a time savings as this post demonstrates; you still have to check each function output manually to find which data to parse / handle edge cases.
- jc4p 10y ago100% this. Coming from someone whose made disgusting chained commands to parse access logs for relevant information, you _need_ better documentation than you can do with a single line of commands. I have so many `zgrep ... | awk ... | sort` scripts that I legitimately couldn't tell you what they do anymore, just what the correct output is. I really like the feeling of proving to myself that I know how to use my command line well, but in most cases I end up spending longer trying to remember the differences between OS X built-in `sed` and the `sed` I know and blah. Lots of wasted time. I've started keeping a "scratch" git repo with scripts for one-offs, so the next time I say "oh I want to go through this CSV and run X on each line with Y" I can look at my old code and replace the necessary parts.
- Symbiote 10y agoYou can put comments in Bash scripts, including in multi-line pipelines echo x |\ # Comment here cat | # Or like this cat An interactive shell with the option interactive_comments set (the default) ignores these comments, which is useful if you wish to copy+paste+execute bits of scripts.
- MichaelBurge 10y agoI've used Makefiles to coordinate a zoo of perl/bash scripts before. It's pretty effective.
- black_knight 10y agoMakefiles are underrated. The concept is so simple – it is basically a dependency graph with attached shell scripts – yet so powerful. Not just for building software, but also for everyday tasks where you need some files to be updated under some condition. I recommend mk (original make replacement for Plan 9, available for different OSes through Plan9Port [0,1]). It has a bit more uniform syntax, and can check that the dependency graph is wellfounded. [0] https://github.com/9fans/plan9port https://github.com/9fans/plan9port [1] https://swtch.com/plan9port/ https://swtch.com/plan9port/
- gcb0 10y agoanother advantage of make is that every host have it. like vim. mk is not there yet.
- viraptor 10y agoDid you mean "vi"? Lots of systems don't install vim by default.
- vacri 10y agoI've been getting into Makefiles more over the past year, and still a noob at it. I see them as a necessary evil - they're baroque and they hurt... and yet they make other things easier. I'm using them less to compile stuff and more to package/upload things, but they certainly have their place.
- Symbiote 10y ago> far preferable Well, that depends. Shell scripts can be commented, and they can be built progressively and interactively by building a pipeline. That's a great choice for a one-off task, and in my experience much faster than many other approaches. It also works well for tasks close to the system. For example, our users are able to download large archives of data, and we keep over 99% of such downloads indefinitely. We delete downloads > 100GB after they're more than 6 months old. With a shell script run by cron that's achieved with find + rm + curl|jq (to tell the API the download is deleted).
- minimaxir 10y agoTo clarify, by data processing I mean extract-transform-load, where the outputs of the extract/transform phase(s) may not be immediately obvious. Bash as an simple automation tool for simple tasks is fine.
- module0000 10y agoWhich company/industry do you work in, where people who would be parsing your code are unable to grok something as basic as bash?
- 0942v8653 10y agoI don't think it's really bash, more the utilities' inconsistent APIs and single-letter flags (which are not descriptive at all). Each command has its own API with different options, so it takes a lot of exposure to really learn them.
- ktRolster 10y agoIf you don't like single letter flags, most have options for longer, more descriptive flags that do the same thing. For example, you can either do: Make -I dir or Make --include-dir=dir Choose whichever one makes most sense.
- gcb0 10y agothis. anyone who uses single letter options on a script should be punished. single letters are for one time typing.
- viraptor 10y agoSingle letters are sometimes more known than their full option equivalents. Quick, without a man page, what's the short alias for `tar --get` ?
- niftich 10y agoI have no idea because I memorized 'tar zxvf' a long time ago [1] and I have to google everything else. [1] https://xkcd.com/1597/ https://xkcd.com/1597/
- 10y ago
- nostrademons 10y agoMost interesting data-analysis problems require multiple iterations. It's not a bad idea to take a couple quick exploratory passes at your data with interactive command-line tools, look at its shape, figure out what areas you need to take a closer look at, and then write real programs to work specifically with those. I've gotten surprisingly far using just curl/cat/cut/sed/awk/wc & friends. When I need to build on top of that, I go write a real program, but the UNIX-fu tells me what I need to build on top of.
- LukeShu 10y agoWhat? For a one-off, one-time "I wonder who has the most medals" curiosity script? > 8 LOC scraper that those without bash knowledge will be unable to parse. And someone without knowledge of Node.js will be unable to parse the 50-line JavaScript monstrosity. > And the 8 LOC bash script is not much of a time savings as this post demonstrates; you still have to check each function output manually to find which data to parse / handle edge cases. It only took that long because he detailed every step as a tutorial. If you have basic literacy of the shell, it's no time at all. To refute you, I decided to solve it myself, using only the "hint" at the beginning of the article to use ?action=raw to work with the source wikitext (I had not yet read the rest of the article; I had not seen his solution). It took me literally 2 minutes: (Setting in url='https://en.wikipedia.org/wiki/List_of_Olympic_medalists_in_judo?action=raw' https://en.wikipedia.org/wiki/List_of_Olympic_medalists_in_j... for readability on HN) [2016-08-15 17:52] curl -s ${url} [2016-08-15 17:52] curl -s ${url}|grep flagIOCmedalist [2016-08-15 17:52] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}' [2016-08-15 17:53] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/[[\]]//g' [2016-08-15 17:53] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2 [2016-08-15 17:54] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/\(\[\|\]\)//g' [2016-08-15 17:54] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/\(\[\|\]\)//g'|sort |uniq -c [2016-08-15 17:54] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/\(\[\|\]\)//g'|sort |uniq -c|sort -n You can see the only place I fumbled a bit was with escaping brackets inside of brackets in sed, which is admittedly a little wonky. Sure, for something that might need to run repeatedly, it probably doesn't handle future edge cases that might arise. But it's not production, it's a one-off exploring the data script. Not every one-off program you write needs to be production quality. And being worried about "those without bash knowledge"... don't be afraid to use your operating system! That said, this line of solutions has a big edge case: it relies on the editors of Wikipedia being consistent and formatting each row as one line in the source. If I were to solve this totally on my own, choosing to ignore the "hint", and work with the rendered HTML, and just do most of it with a nokogiri one-liner. Knowing that there's an ?action=something to get just the page HTML without the navigation and such. I spent about 3 minutes finding the "render" action (documented here: https://www.mediawiki.org/wiki/Manual:Parameters_to_index.php#Actions https://www.mediawiki.org/wiki/Manual:Parameters_to_index.ph... ). Then anther 3-ish minutes poking around the DOM in my browser to get an idea of what I'm working with, and visually inspecting the layout of the article; each cell with a medalist has two links, the first to the medalist, and the second to the country. Then it took me a whopping 5 minutes to hammer out the rest of the one-liner. (Similarly, url='https://en.wikipedia.org/wiki/List_of_Olympic_medalists_in_judo?action=render' https://en.wikipedia.org/wiki/List_of_Olympic_medalists_in_j... for HN readability) [2016-08-15 18:14] curl -s ${url} [2016-08-15 18:15] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}' [2016-08-15 18:16] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'|grep -vE '^[0-9]{4} ' [2016-08-15 18:17] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text.sub!("\n", " ")}' [2016-08-15 18:18] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text.sub("\n", " ")}' [2016-08-15 18:18] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'|grep -vE -e '^[0-9]{4} ' -e '^details$' [2016-08-15 18:19] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'|grep -vE -e '^[0-9]{4} ' -e '^details$'|sort |uniq -c|sort -n And most of it was fumbling around with thinking that the year/details links were one link with a newline in them, when they are in fact two separate links. For a grand total of 11 minutes; generously. My point is: Don't be afraid to play around with your tools and data! Have fun! Not everything needs to be production quality! And being written primarily in bash doesn't necessarily mean that it isn't production quality!
- dangravell 10y agoOr, for a lot of the structured elements, you could use DBPedia.
- opensourcedude 10y agoI appreciate this for the novelty factor, but, somebody show this dude how to use a spreadsheet!
- junke 10y agoHonestly, what novelty factor?
- lacksconfidence 10y agoBecause i was randomly curious, can extract this data from the structured html with some dom selectors in a similarly haphazard way: Start with: https://en.wikipedia.org/api/rest_v1/page/html/List_of_Olympic_medalists_in_judo https://en.wikipedia.org/api/rest_v1/page/html/List_of_Olymp... Run this js one-liner: [].slice.call(document.querySelectorAll('table[typeof="mw:Transclusion mw:ExpandedAttrs"] tr td:nth-child(n+2) > a:nth-child(1), table[typeof="mw:Transclusion mw:ExpandedAttrs"] tr:nth-child(3) td > a:nth-child(1)')).map(function(e) { return e.innerText; }).reduce(function(res,el) { res[el] = res[el] ? res[el] + 1 : 1; return res; }, {}); The result is an object with the medalists as keys, and the count as values. JS objects are unordered so sorting is left as an excercise for the reader.
- merpnderp 10y agoWish I could upvote this more. That is some serious queryselector fu.
- austinjp 10y agoNice. Nasty, but nice :) I can't help but notice a small bug.... Driulis Gonzalez for example has medalled 4 times, but your script gives his count as only 3. Similarly Gévrise Émane isn't listed by your script. Something to do with split tables cells I suspect. Still, it's inspiring. I've often used bash and perl to scrape data from web pages. I'll definitely consider JS in future.
- yarrel 10y agoA couple of years ago I found Perl was fastest at processing Wikipedia dumps. It also didn't require having a JVM preloaded to make startup times acceptable during development (naming no other tools). I do use shell tools to process data, a lot. They're particularly good for exploratory programming and initial analysis of new datasets.
- Steeeve 10y agocut, awk, grep, and perl can churn through an initial data dump like nobody's business.
- turtlebits 10y agoAn xpath like `//table/tr/td[2]/a[1]/text()` seems like it would be a lot simpler.
- san_dimitri 10y agoThis is my goto approach every time I have to parse html or XML. I still don't understand why people don't use something as simple as google spreadsheets and write a simple xpath to load tabular data using =IMPORTHTML().
- hbogert 10y agoIsn't this a poster child example for the semantic web?
- orfix 10y agoMy 2 cents: the cut/grep lines could be replaced by a sed/awk one-liner such as: sed -n 's/.flagIOCmedalist|\[\[\([^]|]\).*/\1/p'
- oxymoron 10y agoAgreed. I was a long time abuser of cut, but has moved to relying on sed instead. I find that it's generally a lot more robust if you think through your expressions. For certain cases awk will also do the job. Perl oneliners do seem convenient but that has never been my cup of tea.
- mickael-kerjean 10y agoYou should try wikidata for any type of query that can't be answer using google and where all the information itself is already on wikipedia. it's way faster (if you know about sparql) and way more powerfull and flexible. It only seems surprising there isn't more people talking about it, triplestore are awesome
- ShakataGaNai 10y agoerrrrrrrrrk. Extracting raw wiki-markup and trying to use it? Not the greatest of idea. The only true parser of that language is mediawiki. Doing it yourself is a recipe for a massive headache.
- ClayFerguson 10y agoI was thinking the same thing. Wikimedia makes all of wikipedia available for download. You don't need to screen scrape. LOL. I guess they had some other wikis they want to get data from but the "main" worldwide wikipedia site, that everyone thinks of as wikipedia makes the data freely downloadable, and I've downloaded it before.
- betolink 10y ago...or we can just use SPARQL and dbpedia!(http://wiki.dbpedia.org/ http://wiki.dbpedia.org/) There are questions where you'll have to scrap more than one page to get an answer and things could get really complicated with shell commands. dbpedia is a triple-store that allows us to perform simple queries against wikipedia data like listing music bands based on a particular city: SELECT ?name ?place WHERE { ?place rdfs:label "Denver"@en . ?band dbo:hometown ?place . ?band rdf:type dbo:Band . ?band rdfs:label ?name . FILTER langMatches(lang(?name),'en') } or queries that involve multiple subjects, categories etc.
- crypto5 10y agodbpedia provides very low coverage of wikipedia information.
- lacksconfidence 10y agoI looked at dbpedia, but it was non-obvious to me what statements to use. We can also use SPARQL with wikidata, although the coverage isn't particularly great.I threw together an example query for medalists and michael phelps doesn't make the list, because he doesn't have the appropriate participant of/award received statements: https://query.wikidata.org/#SELECT%20%3Fhuman%20%3FhumanLabel%20%3Fcount%20WHERE%20%7B%0A%20%20%7B%0A%20%20%20%20SELECT%20%3Fhuman%20%28COUNT%28%2a%29%20as%20%3Fcount%29%20WHERE%20%7B%0A%20%20%20%20%20%20%3Fevent%20wdt%3AP31%20wd%3AQ18536594%20.%20%23%20All%20items%20that%20are%20instance%20of%20Olympic%20sporting%20event%0A%20%20%0A%20%20%20%20%20%20%3Fmedal%20wdt%3AP279%20wd%3AQ636830%20.%20%20%20%20%23%20All%20items%20that%20are%20subclass%20of%20Olympic%20medal%20%0A%0A%20%20%20%20%20%20%3Fhuman%20p%3AP1344%20%3FparticipantStat%20.%20%23%20Humans%20with%20a%20participant%20of%20statement%0A%20%20%20%20%20%20%3FparticipantStat%20ps%3AP1344%20%3Fevent%20.%20%23%20..%20that%20has%20any%20of%20the%20values%20of%20%3Fevent%0A%20%20%20%20%20%20%3FparticipantStat%20pq%3AP166%20%3Fmedal%20.%20%23%20..%20with%20the%20award%20received%20qualifier%20of%20any%20of%20the%20values%20of%20%3Fmedal%0A%20%20%20%20%7D%0A%20%20%20%20GROUP%20BY%20%3Fhuman%0A%20%20%7D%0A%20%20%0A%20%20SERVICE%20wikibase%3Alabel%20%7B%20bd%3AserviceParam%20wikibase%3Alanguage%20%22en%22.%20%7D%0A%7D%0AORDER%20BY%20DESC%28%3Fcount%29%0ALIMIT%20100%0A https://query.wikidata.org/#SELECT%20%3Fhuman%20%3FhumanLabe... SELECT ?human ?humanLabel ?count WHERE { { SELECT ?human (COUNT(*) as ?count) WHERE { ?event wdt:P31 wd:Q18536594 . # All items that are instance of Olympic sporting event ?medal wdt:P279 wd:Q636830 . # All items that are subclass of Olympic medal ?human p:P1344 ?participantStat . # Humans with a participant of statement ?participantStat ps:P1344 ?event . # .. that has any of the values of ?event ?participantStat pq:P166 ?medal . # .. with the award received qualifier of any of the values of ?medal } GROUP BY ?human } SERVICE wikibase:label { bd:serviceParam wikibase:language "en". } } ORDER BY DESC(?count) LIMIT 100 EDIT: The dbpedia search should be something like: http://dbpedia.org/sparql?default-graph-uri=http%3A%2F%2Fdbpedia.org&query=SELECT+%3Fhuman+%3Fcount+WHERE%0D%0A%7B%0D%0A++%7B%0D%0A++++SELECT+%3Fhuman+%28count%28*%29+as+%3Fcount%29+WHERE+%7B%0D%0A++++++%3Fevent+rdf%3Atype+dbo%3AOlympicEvent%0D%0A++++++%7B%0D%0A++++++++%3Fevent+dbo%3AbronzeMedalist+%3Fhuman+.++%0D%0A++++++%7D+UNION+%7B%0D%0A++++++++%3Fevent+dbo%3AsilverMedalist+%3Fhuman%0D%0A++++++%7D+UNION+%7B%0D%0A++++++++%3Fevent+dbo%3AgoldMedalist+%3Fhuman%0D%0A++++++%7D%0D%0A++++%7D%0D%0A++++GROUP+BY+%3Fhuman%0D%0A++%7D%0D%0A%7D%0D%0AORDER+BY+DESC%28%3Fcount%29%0D%0ALIMIT+100&format=text%2Fhtml&CXML_redir_for_subjs=121&CXML_redir_for_hrefs=&timeout=30000&debug=on http://dbpedia.org/sparql?default-graph-uri=http%3A%2F%2Fdbp... SELECT ?human ?count WHERE { { SELECT ?human (count(*) as ?count) WHERE { ?event rdf:type dbo:OlympicEvent { ?event dbo:bronzeMedalist ?human . } UNION { ?event dbo:silverMedalist ?human } UNION { ?event dbo:goldMedalist ?human } } GROUP BY ?human } } ORDER BY DESC(?count) LIMIT 100
- davidgerard 10y ago... there's an API making half of this superfluous. You can do pretty much any MediaWiki reading or writing through it. (All Wikipedia bots are required to use it, for instance.) https://en.wikipedia.org/w/api.php https://en.wikipedia.org/w/api.php The article text is a raw blob of wikitext you have to process, but you don't have to go to stupid lengths trying to parse HTML without a browser.
- gkbrk 10y agoBut he didn't parse any HTML in the article.
- Washuu 10y agoThere is the other option to use Parsoid. https://github.com/wikimedia/parsoid https://github.com/wikimedia/parsoid That is MediaWiki's official off wiki parser that can turn wikitext into HTML or HTML back into wikitext. It would be reasonably simple to hook into its API and use it for data extraction instead.
- rspeer 10y agoIs converting Wikitext to HTML/RDFa really going to help with this task? I'd say it's actually clearer how to get the data out of the original Wikitext.
- tpetricek 10y agoExtracting data from Wikipedia with type providers: http://evelinag.com/blog/2015/11-18-f-tackles-james-bond/ http://evelinag.com/blog/2015/11-18-f-tackles-james-bond/
- jpatokal 10y agoThis is you-can't-parse-HTML-with-regex [1] level hideous, only worse, because Mediawiki markup is essentially a Turing-complete programming language thanks to template inclusion, parser functions [2], etc. The only remotely sane way to do this is to use the Mediawiki API [3] to get the pages you want, then use an actual parser like mwlib [4] to extract the content you need. Wikidata and DBpedia are also promising efforts, but both have a long way to go in terms of coverage. [1] http://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags http://stackoverflow.com/questions/1732348/regex-match-open-... [2] https://www.mediawiki.org/wiki/Help:Extension:ParserFunctions https://www.mediawiki.org/wiki/Help:Extension:ParserFunction... [3] https://www.mediawiki.org/wiki/API:Main_page https://www.mediawiki.org/wiki/API:Main_page [4] https://www.mediawiki.org/wiki/Alternative_parsers https://www.mediawiki.org/wiki/Alternative_parsers
- taneq 10y agoThis isn't for production-level data migration. It's for smooshing some source text into a shape which is useful to you. Parsing HTML with regexps is fine if you're just curious roughly how many images are in a page. It's great for quick command line experiments. It's just not good when you need to be "doing it properly".
- jpatokal 10y agoI used to work with Wiki markup for a living. The time you think you'll save with regex hackery is quickly chewed up by the time wasted eternally tweaking your regexes to catch yet another corner case -- it's much better just to parse for real from the get go, just like it's much better to use a real HTML/XML parser than trying to do the same job badly with regexes.
- ivanhoe 10y agoI used to make scrappers for living and trust me that it all depends on the particular situation and your requirements. Real HTML parsers are easier and safer for general type of work, but they quickly get very heavy on memory when parsing big DOM trees. If all you need from a page is a few strings, like e.g. just a product price (very common task), using regexes is far superior approach performance-wise. It's both faster and uses less memory (so you can run more parallel workers) and also if you write it well it's immune to many small html/design changes as long as the pattern you look for is not changed.
- kasperset 10y agoLarge part of Bioinformatics data processing involves these commands. They seem little cryptic but gets job done. I would also like to mention Datamash: https://www.gnu.org/software/datamash/ https://www.gnu.org/software/datamash/
- CydeWeys 10y agoThere is an active project sponsored by the Wikimedia Foundation called PyWikiBot that I've been a contributor to and user of for over a decade now. If you want to do anything and everything with Wikipedia, look no further than: https://github.com/wikimedia/pywikibot-core https://github.com/wikimedia/pywikibot-core
- lovelearning 10y agoPython has excellent packages like mwparserfromhell and wikitables for this kind of processing.
- loige 10y agoI actually added some of your alternative solutions to the bottom of the article, thanks for commenting :)