5 ms·
Wikipedia has an API. This is just fetching the HTML and parsing it with beautifulsoup, which seems like a terrible idea when an API is available.
by scandinavian 5y ago
Wikipedia has an API. This is just fetching the HTML and parsing it with beautifulsoup, which seems like a terrible idea when an API is available.
- akamhy 5y agoThey could have used https://github.com/wikimedia/pywikibot https://github.com/wikimedia/pywikibot (python API interface) or https://github.com/goldsmith/Wikipedia https://github.com/goldsmith/Wikipedia
- octoberfranklin 5y agoWebsite APIs tend to sprout "API key" requirements with little or no notice, which seem like a terrible idea. This kind of breakage is far less likely to happen to the HTML endpoint, because it gets about a zillion times more usage. It is actually sensible to stick to the interface that the majority of users are using, because that interface is the least likely to break. Spontaneous API key hoop-jumping is a form of breakage.
- martneumann 5y agoAre you saying that it's more likely API key requirements appear at some point than the interface changing? I'm not that informed, but intuitively, I'd doubt that.
- howenterprisey 5y agoAs someone who's reasonably informed on this stuff, I would be absolutely shocked if Wikipedia ever introduced that requirement.
- samatman 5y agoThis is an example of good (well, ok) general advice, which absolutely does not apply to this specific instance. As general advice it's only ok, because a) the details of served HTML can change a great deal and b) some APIs might have this problem, some are unlikely to. Wikipedia is an extreme example of the latter.
- deleted 5y ago[deleted]
- creativeCak3 5y agoWas thinking the same thing...
- kragen 5y agoIt's probably easier to parse the HTML than the Wiki markup.