9 ms·
Show HN: Link.fish – API to extract data from websites as JSON
- etewiah 9y agoI've been checking out link.fish for a while now - awesome product! My interest is in scraping real estate websites and it seems to do quite a good job with many of the sites I've tried. Already mentioned by others but I suggest 1. Concentrate on a specific segment (like real estate) 2. Consider a browser extension (helps mitigate problem of too many requests coming from one central server) I have long planned on building an open source real estate website scraper but just haven't found the time to do it.
- pen2l 9y agoIt's been quite a while since I last did web-scraping (I used to use BeautifulSoup, more than a decade ago). I'm just wondering, since a lot of people are using fairly advanced cloud-hosting solutions with, I assume, tools offered by their respective hosting place to fight spam, is web-scraping a lot different from what it used to be about a decade ago? What steps do you guys take to prevent being identified as a bad actor by the place that you are scraping? And on the other end, if you have a data-rich website, what are your feelings toward aggressive scrapers?
- twblalock 9y agoCDNs like Distil Networks and Cloudflare make scraping more difficult than it used to be. If you get caught by them, you can end up blocked from all of the sites they protect, not just the one you were scraping.
- always_good 9y agoWriting some scrapers this week, I noticed it's also common for the origin server to just check if the request is coming from VPN/VPS IP address range. For example, the exact same request will work from your home connection where it doesn't work from EC2.
- mfrye0 9y agoIt's gotten pretty challenging from what it used to be. A lot of small things... but basically if you load from an actual browser (headless) and cycle IPs, it's pretty hard for a site to pinpoint you as a bot vs a user.
- janober 9y agoHi just launched this API. So would love to get feedback like what could be improved or what is missing. Is a first version so any kind comments are welcome!
- wiradikusuma 9y agoI tried a link to some product in some ecommerce (http://www.lazada.com.my/samsung-galaxy-note-8-6gb-ram64gb-rom-original-samsung-malaysia-set-black-91710731.html?spm=a2o4k.home.recommended-items_31972.18.34aa7d86A4adxN&lzd_rec_event_src=&lzd_rec_event_dest=SA356ELABILOGBANMY&strat=rec_global_top_prods&pa=home_page.recommended_items http://www.lazada.com.my/samsung-galaxy-note-8-6gb-ram64gb-r...), and it does extract the content.. but I only care about the "hero" item. Is it safe to just always take the 1st item in mainEntity>offers>offers[0]?
- linkfish 9y agoIt did actually just extract the data of the "hero" item. The thing is that it gets offered by multiple companies for different prices. So all the prices are valid and none is right or wrong. So really depends what you want. If you want simply "a" price, you can take the first. If you want the cheapest one you would have to itterate over them to find it.
- JanKoenig 9y agoThis looks very helpful. Would love to use some of this structured data to create a sample voice app for Jovo
- lawl 9y agoI tried a random reddit thread. Did not fetch comments, only information about the submission. Then I tried it with HN. Same. Then I tried it with a github issue. Same. Then I tried it with the first link I got from news.google.com which was nytimes. No article text included. Maybe I'm misunderstanding the purpose of this? Or was that just a string of bad luck?
- Treegarden 9y agoSame here. Really curious though if there is any service/api to harvest news articles from websites to experiment with text analysis with? Havent found any after some search, just apis that provide meta information but not an actual corpus of text.
- janober 9y agoNo is not just bad luck. Currently did not concentrate so much on "text pages" like blogs or articles yet. Mainly on pages which contain more data like prices, geo coordinates, social media profiles, ... That said support for the mentioned pages can simply be added via our point and click GUI by any user. Do sadly not have time right now, but can add support for this pages by tomorrow.
- lawl 9y agoAh thanks. Currently I have already written scrapers for the stuff I need. I was mainly just curious. I'll bookmark it to look at again when I need something next time.
- weego 9y agoSo your paid for product for scraping and structuring information from a webpage cannot actually return most content off a webpage as it is now? Wouldn't that be a more important vertical slice of a product for an MVP than having a fully thought out pay tier system?
- nedwin 9y agoMaybe? It's all about tradeoffs. I could see why you would want to figure out how much value the MVP is creating, and $$$ is an honest way to do that. It sounds like two things are happening with the MVP: - emphasis on more complicated sites with more data (higher propensity to pay) - this functionality is actually possible but a user needs to take the time to set it up via a the GUI. Feels like a pretty good tradeoff to me.
- sjs382 9y agoIn the pricing, different tiers have different "priorities". It would be helpful to know what these "priorities" mean in the real world. If I submit a request as a low-priority user, should I expect a response in 1 second? 1 minute? 1 hour? Something else? And how consistent is the amount of time I should expect to wait?
- janober 9y agoThe time it takes, in general, is mainly dependent on how fast the page loads. The priority simply means, that if for some reason there are more requests at a given time then we can handle, that the people with the higher plans get served first. With other words for 99% of the requests, there should not be any time difference at all.
- sjs382 9y agoI understand that people with higher priority get served first. The question is how that will affect someone's real-world use.
- janober 9y agoLike written above should it normally not make a difference at all. But sure is possible that if there is an unexpected huge surge the people in the lowest plan suddenly have to wait a few seconds longer.
- sjs382 9y agoI understand. I realize that I wasn't clear but my second comment was feedback wrt/ your marketing page.
- mmahemoff 9y agoLooks good, but why complicate the pricing with "credits" when 1 credit==1 request. The tiers already reference parallel "requests", so you could just say N requests instead of N credits.
- janober 9y agoThe reason is that we also offer that pages can be rendered with a full browser (to execute all JavaScript and make screenshots) which takes much more resources.
- noinput 9y agoa 404 on the TOS & Privacy policy isn't the best to build confidence.
- linkfish 9y agoVery sorry for that! Worked everywhere but should not have used a relative link on the free account signup page. Got fixed.
- TenJack 9y agoWasn't this submitted already? https://news.ycombinator.com/item?id=15099041 https://news.ycombinator.com/item?id=15099041 https://news.ycombinator.com/item?id=14522439 https://news.ycombinator.com/item?id=14522439 How are you able to do 'Show HN' in such recent succession?
- janober 9y agoIs the same domain but different products. The one posted before is the bookmark manager for mainly B2C. The product I posted now is the B2B version which uses the same technology behind the bookmark manager but allows access via API.
- profalseidol 9y agoNope, nothing useful. rather build your own parser.. https://www.pmu.fr/turf/02112017/R4/C4 https://www.pmu.fr/turf/02112017/R4/C4
- kashprime 9y agoIs this similar to Apify? https://www.apify.com/ https://www.apify.com/
- linkfish 9y agoYes is similar to it in the regard that with both tools data from websites can be extracted.
- mericsson 9y agoWell done! How does this compare to diffbot's website extraction? https://www.diffbot.com/ https://www.diffbot.com/
- linkfish 9y agoHard to say and compare in what regard exactly? Price wise, much cheaper. Data wise, it depends. Mainly on in what kind of data you are interested in. Will probably return better results on text-heavy pages, but the data is probably often less "deep". So really depends on the use case. If you have a question to your use case, you can simply write to api@link.fish .
- raresp 9y agoAt a first look it's much cheaper than Diffbot :)
- dwynings 9y agoDisclaimer: I work at Diffbot Major differences I can see (OP feel free to correct if I'm wrong): Link.fish * doesn't provide a web crawler * relies heavily on microdata, schema.org, RDFa, etc * relies on manual parsers for sites that don't have microdata embedded * doesn't full-render pages by default (Diffbot renders every page, so it can use computer vision to automatically extract the data) * doesn't support proxies * doesn't support entity tagging Probably plenty more, but that's what jumps out to me at first blush. -- Since I see other people have mentioned price as a concern, we're always willing to help out bootstrapped startups. Just shoot me an email: dru@diffbot.com
- danielvinson 9y agoThe feature I'd be looking for in this is to be able to recursively scrape for contact information (email, phone number, etc.)... that doesn't seem possible with this?
- linkfish 9y agoYes right now not. However an endpoint for exactly that is actually planned. So you can simply write to api@link.fish and we can inform you when it is ready.
- swaraj 9y agoDoes this use schema.org / og meta tags or are you trying to infer object types yourself
- linkfish 9y agoIt uses a combination of everything to extract information incl. custom parsers. The data returned to the user is always in schema.org.
- callmeed 9y agoHaving done a ton of scraping in the past (especially around ecommerce and products), this looks pretty cool. A couple comments in general: 1. Personally I think its better to be great at extracting one kind of data instead of average at many types. It makes sales and growth efforts easier. Pick one of those things (products, recipes, social, etc.) and just focus on that and get great at it. 2. I don't think you need the credit <-> request abstraction. Anyone using an API knows what a request is (I hope). Now, a few comments regarding products specifically: 1. I got 500 errors on a couple random product URLs. 2. On an Amazon product that's on sale, I got back the original price but not the sale price. 3. If you truly want to be GREAT at scraping products, the 2 things most people in this space can't do are: (a) extract ALL high-res images for a product, and (b) extract a product's options and variant data (colors, sizes, etc. and availability for each combination) Personally I think there are a ton of opportunities in this space. This is a good start and I wish you the best.
- RussianCow 9y agoI think the credit concept was created solely for this reason (from the page): "There is only one exception if the page should be rendered with a full browser (not headless). In this case, 5 credits get charged."
- linkfish 9y agoThanks a lot for all the feedback! 1. Yes will think about it. For just the API it would make definitely sense. However because the same technology currently also powers the bookmarking service which has to support as much as possible does the API also. 2. Exactly what RussianCow said. Honestly not a big fan of it either but that was the best I could come up with to accommodate that. About the product. 1. It logs all requests which had issues with the more descriptive cause. Always go through all of them and fix the issues. The more people use it the more stuff breakes and the product can be improved. So I guess will get way better in the next days ;-) 2. Will also check and fix. 3. Will definitely look into that! If you run into more issues or have more comments would love to hear them here or at api@link.fish . Thanks again!
- tomascot 9y agoWhat could cause that that you mention on point 2. Are they "caching" responses or that offer is tailored to your user/cookie?
- jotaen 9y agoBug report: when I try out the service on your frontpage, the URLs seem to get converted to lowercase internally. So if I try to fetch this URL https://goo.gl/DKukBD https://goo.gl/DKukBD (which points to this very HN submission), it actually queries https://goo.gl/dkukbd https://goo.gl/dkukbd (which points to some random website)
- linkfish 9y agoThanks a lot! Will investigate and fix. If you run into any other issues please keep them coming!
- tylerpachal 9y agoThere is a small typo in the "Why link.fish API?" section on the homepage: > Additionally, do we have a growing collection of custom parsers for websites and website independent parsers for specific data. I don't think you need the "do".
- linkfish 9y agoAs not native english speaker I am always happy to get help in that regard ;-) Thanks!
- laktek 9y agoNice work! I also built something very similar called Page.REST (Show HN thread: https://news.ycombinator.com/item?id=15189099 https://news.ycombinator.com/item?id=15189099) Page.REST supports extracting contents using CSS selectors. oEmbed and OpenGraph tags. Also, do you plan to support extracting from client-side rendered pages (a la React)? BTW, I'm interested how you decided the pricing? I went with a $5 one time fee as most people use such tools for ad-hoc purposes. General question to readers: What do you think of the Schema.org format? Is it easy to consume? (from a language library perspective)
- linkfish 9y agoThanks, dito! Is actually already supported when a special parameter is set(on the API-Test-Box on the landing page it is not set). About the pricing. Was a longer process. Mainly involved what other similar services charge and much more important a price which makes the service viable in the long term.
- tomerbd 9y agomay i ask how do you process the payments? looks nice and clear. thanks.
- laktek 9y agoStripe. I recently integrated their Payment Request Button https://stripe.com/docs/stripe-js/elements/payment-request-button https://stripe.com/docs/stripe-js/elements/payment-request-b...
- TamDenholm 9y agoI just tried purchasing a token, didnt work. I also tried the on-site chat and it didnt seem to work either... Wanna gimme an email? Check my profile.
- laktek 9y agoUhh, weird. I just sent you an email.
- attacomsian 9y agoCan't access the site. Is it down?
- janober 9y agoVery sorry for that. Thought a smaller server could handle the load because of Cloudflare but was apparently wrong. Is up again btw.
- adventurer 9y ago520 error here. Hug of death or they took your site down.
- janober 9y agoVery sorry for that. Thought a smaller server could handle the load because of Cloudflare but was apparently wrong. Is up again btw.
- holtalanm 9y agoThis looks almost exactly like the functionality provided by https://page.rest/ https://page.rest/
- nailer 9y agoTrying a site with a simple HTML schedule: https://www.pineapple.uk.com/studio/index/filter/ https://www.pineapple.uk.com/studio/index/filter/ didn't work. I'd love it to be able to do this and would pay a small per-API-call fee.
- janober 9y agoClicked something together very fast but then did not work in the end because the domain of the website (the library I use did not know uk.com). Will fix that issues tomorrow. You can either simply check again tomorrow or contact me at api@link.fish .
- nailer 9y agoIt works, but only on that site - a similar site (http://studio68london.net/work/timetable/ http://studio68london.net/work/timetable/ fails). I want to be able to point something at an arbitrary URL and have it extract the tabular data. I'd write this myself, but I'd rather pay someone else to maintain it.
- ojanik 9y agohttps://en.wikipedia.org/wiki/The_Expanse_(TV_series) https://en.wikipedia.org/wiki/The_Expanse_(TV_series) Error: 500 Message: { "status": 500, "message": "Internal Server Error" }
- janober 9y agoYes it seems like you found a bug;-) Is gonna get fixed tomorrow. Thanks!
- gbrits 9y agoSo where's the point & click GUI to select items from a page? Signed up, but can't seem to find it on first glance
- janober 9y agoSorry yes, have to make that clearer or write in the email. Is on the top of the page under "Plugins" -> "Data Selector".
- nl 9y agoI've been playing around with a way of extracting information from the text on websites (eg, finding names of people or price ranges in a textual story rather than in a table). I've got as far as something that works much better and is much more flexible than things like UoW OIE or just using Stanford named entity recognition. Is this a thing others need or would find useful?
- purplepotato 9y agoI like this article.
- taivare 9y agolike your logo
- nreece 9y agoLooks good! * shameless plug * Our little startup, Feedity - https://feedity.com https://feedity.com, also helps create custom RSS feeds for any webpage.
- palani666 9y agoI tried a link to some product in some ecommerce (http://www.lazada.com.my/samsung-galaxy-note-8-6gb-ram64gb-r... http://www.lazada.com.my/samsung-galaxy-note-8-6gb-ram64gb-r...), and it does extract the content.. but am only care about the "hero" item. Is it safe to just always take the 1st item in mainEntity>offers>offers[0]?
- janober 9y agoActually answered that question already yesterday. For convenience here again: It did actually just extract the data of the "hero" item. The thing is that it gets offered by multiple companies for different prices. So all the prices are valid and none is right or wrong. So really depends what you want. If you want simply "a" price, you can take the first. If you want the cheapest one you would have to itterate over them to find it.