3 ms·
It's silly to use BeautifulSoup to parse the page when you could use a simple RegEx: <td class=\"title\"><a href=\"(.?)\"(.?)>(.?)</a>(.?)</td>
by sprizzle 13y ago
It's silly to use BeautifulSoup to parse the page when you could use a simple RegEx:
<td class=\"title\"><a href=\"(.?)\"(.?)>(.?)</a>(.?)</td>
- sprizzle 13y agoArg, there should be asterisks after every period.
- joshbaptiste 13y agoSome people, when confronted with a problem... bah you know the rest.
- kaeawc 13y ago"HTML tags lea͠ki̧n͘g fr̶ǫm ̡yo͟ur eye͢s̸ ̛l̕ik͏e liquid pain" http://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags http://stackoverflow.com/questions/1732348/regex-match-open-...
- michaelmcmillan 13y agoI am willing to sacrifice my soul and everything that is holy.
- karangoeluw 13y agoRegex to parse HTMl is probably the single worst thing you can do.
- lloeki 13y agoCrafting a wide purpose regex to parse whatever HTML comes in is bad. Building a regex to extract relevant data from simple, fixed-form page data, bypassing tags irrelevant to the problem at hand is not.
- untothebreach 13y ago...until the HTML changes. I haven't look at their parsing code, so I have no idea if it is any better than using a regex, but if the regex assumes too much, simply reordering the attributes in a tag (or something similar) could break a regex-based solution.
- Goranek 13y agoBeautifulSoup is great, as long as you're using open source HTML5 parser from Google. https://github.com/google/gumbo-parser https://github.com/google/gumbo-parser