4 ms·
If you like web-scraping, and you're not an idiot, you immediately have an information advantage for anything you decide to do. Do not underestimate this advant
by j0rd 11y ago
If you like web-scraping, and you're not an idiot, you immediately have an information advantage for anything you decide to do. Do not underestimate this advantage.
I would suggest you leverage your interest in web-scraping and use it to customers in what ever niche you believe you can sell them something.
Collect your "beta" users through web-scraping, figure out a way to reach them at scale. Build/sell a product they want (you can leverage your information advantage to figure out what this is)
This is personally what I do and have created/worked on many bootstrapped companies over my career. Niche for me is fairly irrelevant, as long as I have a pool of interested customers ahead of time. I primarily use my information advantage to figure these out.
The additional skills you will need is:
* learn to think as a "user/customer"
* minor copy-writing skills (or at least understand what shitty copy is and how to improve it)
* UX
* Data mining & Analytics, Analytics, Analytics
* A/B Testing, Iterations, Incremental Improvements (see point above)
* Hypothesis Driven Development (https://www.thoughtworks.com/insights/blog/how-implement-hypothesis-driven-development https://www.thoughtworks.com/insights/blog/how-implement-hyp...)
My one suggestion is to avoid any projects which do not have a clear monetization strategy. If you're following my blue-print companies that don't make money from day will only incur costs as you reach your pool of customers at scale.
- deskamess 11y agoCan you provide an example or two of a project that involved web scraping? (without divulging client info) I can see how it could be used for price scraping in web retail, but am drawing a blank in other industries.
- TheLogothete 11y agoText analytics, for one.
- mtmail 11y agoAddress parsing from a large web corpus https://github.com/openvenues/libpostal https://github.com/openvenues/libpostal The author uses http://commoncrawl.org/ http://commoncrawl.org/ data, 250TB/3.6b URL and a 100 machine cluster of 8-core machines.