3 ms·
Ah thank you, I modified it away from scraper (just thought that would help people understand how it functioned)
by Jabbs 2y ago
Ah thank you, I modified it away from scraper (just thought that would help people understand how it functioned)
- andai 2y agoHow are you extracting the data?
- Jabbs 2y agoCurrently 4 worker services run headless chrome browsers and visit company websites (about 10k URLs per service per day). If certain keywords and phrases match a valid dataset of job listings and titles then the page content is saved and parsed w various tags (locations, technologies, etc.)
- andai 2y agoThanks. I was wondering if you're using a LLM to get unstructured data? Or how do you recognize that something is a location for example?
- Jabbs 2y agoAh yea no LLMs but probably some of the similar ideas were used by having a validating dataset. I just have ~800 common locations (city, state, countries) that are parsed out of the text with validating logic wrapped around each one. Far from perfect but it is now pretty accurate. I found the index of the location matters a lot for accuracy. Its still a WIP