4 ms·
Thanks a lot for your reply! Scraping websites can be quite the messy business, since some websites change their document structure more often than others. No
by giu 6y ago
Thanks a lot for your reply!
Scraping websites can be quite the messy business, since some websites change their document structure more often than others.
Nonetheless, it's still a very instructive activity and you can build quite the pipeline around it (scraping multiple websites, joining datasets, efficiently storing the data, etc.).
- iagovar 6y agoYeah, when data piled up I had to think about how to store it, RAM, and a bunch of other things that I didn't have to consider with sample data. Specifically RAM and how to transform data without so much need of it was a concern for some time.
- rohan_shah 6y agoI am also currently learning to scrape forums. And I am a philosophy student. Could you point to some resources that helped you learn it better?
- iagovar 6y agoAre you looking for something specific? Most tools have documentation you can bang your head against.
- jmt_ 6y agoLearning CSS selectors and HTML structure, inspect element and the other dev tools builtin to your browser, and something like BeautifulSoup (for static/non-JS heavy pages) and Selenium (JS and other complicated pages) is pretty key imo. My background in web dev helped me with the HTML stuff. Basically, you fire up the page in a browser, inspect element to see how you can use CSS selectors to uniquely identify that data, then using BeautifulSoup or Selenium to parse and interact with the DOM will cover most web scraping cases.