8 ms·
This looks very useful, big fan of all the ^[a-z]+q$ utilities out there. But as a user, I would probably want to use XPath[0] notation here. Maybe that is just
by dfederschmidt 5y ago
This looks very useful, big fan of all the ^[a-z]+q$ utilities out there. But as a user, I would probably want to use XPath[0] notation here. Maybe that is just me. A quick search revealed xidel[1] which seems to be similar, but supports XPath.
[0]https://en.wikipedia.org/wiki/XPath https://en.wikipedia.org/wiki/XPath
[1]https://github.com/benibela/xidel https://github.com/benibela/xidel
- chriswarbo 5y agoMy web scraping tends to start with xidel. If I need a little bit more power I'll use xmlstarlet. If neither of those is enough, I'll use Python's beautifulsoup package :)
- mikepurvis 5y agoI like xmlstarlet too, if only because it's old enough that I can reliably get it in package repositories and the dependency footprint is tiny (less an issue now with this tool written in Rust, but previously I was comparing to NPM- and PyPI-based affairs).
- spiralx 5y agolxml is one of the most pleasing to use Python libraries ever, managing to wrap a hot mess of XML APIs in a consistent and Pythonic fashion that you rarely need to escape. IIRC I used beautifulsoup to parse the HTML of a site, and then lxml and either find items and fields by CSS in IPython for quick and dirty data munging, or knock up an XSLT file to transform what I'd scraped into good data in an XML file :)
- akie 5y agoI'd like to state my support for the author's choice of CSS selectors in this particular use case. I think it's a natural fit for this domain and already very well known, perhaps even known better than XPath.
- berkes 5y agoI'd like to add my support here too, but with a note. When scraping and parsing (or writing integration test DSL), I always start out with CSS selectors. But always hit cases where they lack or require hoop-jumping and then fall back on Xpath. I then have a codebase with both CSS-Sel and Xpath, which is arguably worse then having only one method. I suspect here, one uses this tool untill CSS selector limitations are getting in the way, after which one switches to another tool(chain)
- Jenk 5y agoI've not had much friction using either, they are "close enough" that the time to (re)write a query from one to the other is not very significant.
- alpha_squared 5y agoDo you mind giving an example? I'm having trouble following where CSS is limited for selection.
- vlunkr 5y agoWell, the big one is selecting a parent from the child.
- androceium 5y agoYou could do this with the :has() CSS psuedo-class[0], though inverted (select a parent that _has_ the child matching a selector). Looks like that psuedo-class has not been implemented in the kuchiki library that htmlq uses though. [0]: https://developer.mozilla.org/en-US/docs/Web/CSS/:has https://developer.mozilla.org/en-US/docs/Web/CSS/:has
- spiralx 5y agoYou can do it either way in XPath thanks to how you can use a path expression and/or predicates almost everywhere in a query # Find all elements li and select the parent element for each //li/.. # Find all element nodes with a child element named li //*[li] # Non-abbreviated queries /descendant::li/parent::* /descendant::*[child::li] # CSS using :has :has(> li)
- exyi 5y agoThanks, this looks more powerfull. Support CSS, XPath and XQuery. Maybe I could learn a bit of XQuery when I have a use case for it :)
- dmit 5y agoWell, here’s your first lesson then: if you prepend (: to your comment it will become a valid XQuery document! (: XQuery comments are marked by mirrored smilie faces, like this. :)
- spiralx 5y agoEverything that isn't a (: happy comments :) is a FLWOR: <users> { for $user in //users let $comments = //comment[@uid = $user/@id] where count($comments) > 0 order by $user/lastName, $user/firstName return <user id="{ $user/@id }"> <name>{ concat($user.firstName, " ", $user.lastName) }</name> <comments count="count($comments)"> { for $c in $comments return <comment id="{ $c/@id }" /> } </comments> </user> } </users> It's the bastard child of SQL and XPath 2 lol. http://www.stylusstudio.com/xquery-flwor.html http://www.stylusstudio.com/xquery-flwor.html
- phlummox 5y agoI kinda liked XQuery, but it seemed to never have got much traction.
- waynenilsen 5y agopart of the problem with this is that HTML is mostly not valid XML
- lilyball 5y agoThis looks really neat! It supports a bunch of different query types, and can even do things like follow links to get info about the linked-to pages! It's also in nixpkgs, though for some reason the nixpkgs derivation is marked as linux-only (i.e. not Darwin). (Edit: probably because the fpc dependency is also Linux-only, with a linux-specific patch and a comment suggesting that supporting other platforms would require adding per-platform patches)