5 ms·
Please give us feedback
by adibalcan 10y ago
Please give us feedback
- deleted 10y ago[deleted]
- klausjensen 10y agoNice and simple idea, pretty well executed with not many bells and whistles. I like it. :)
- adibalcan 10y agoThanks!
- cookiecaper 10y agoThis is probably illegal in the United States. Expect to receive cease and desist demands.
- pbhjpbhj 10y agoIllegal on what basis?
- cookiecaper 10y agoThe CFAA, which states "exceeding authorized access" to a computer system is both a crime and a tort, and the Copyright Act, which has been interpreted to mean that copies of HTML pages, even if they exist only for microseconds in RAM, are subject to copyright and thus, copyright infringement claims can be brought against anyone who downloaded the page. It's also breach of contract (which I'm labeling separately from "illegal" to avoid nitpickers, even though it could be included) due to the individual ToS on each site, which almost always include boilerplate forbidding the access of the site by "automated means" in addition to forbidding "commercial" or other non-personal use. Before you raise the common counterarguments, please know that others have done so before you, and the courts have generally sharply disagreed with them. There is no respect for non-Google data scraping in the judiciary.
- adibalcan 10y agoWhat about import.io?
- cookiecaper 10y agoFirst IANAL. But there are people building businesses on this kind of thing even though it's almost universally not allowed. Google is the most prominent -- the judiciary has decided that a different set of rules applies to Google v. smaller companies, so Google gets away with it. That's because they are massive and were probably able to pay bribes to the judges to get them to agree with them. Usually if you C&D upon receipt of C&Ds, you won't get sued unless you actually damaged the site. I assume companies like import.io and Scrapinghub adhere to those. I know particularly in the case of Scrapinghub, they won't go behind a login to get something in order to try to avoid heat. Fundamentally, however, anything that makes scraping a key component of its business model is a high-risk business under current US law. People can and have sued scrapers out of existence, often leaving a trail of screwed customers, laid off employees, and dejected founders owing huge liability judgments.
- pbhjpbhj 10y agoI know Google have been treated as if it's a sui generis case, but surely Archive.org and many others are scraping and not being shutdown. It's weird that the comparison between posting a robots.txt and a "no trespassers" sign hasn't been upheld? Or has it. IMO the tort of copyright hasn't kept up well with tech changes, but transient cache copies are handled in EU laws IIRC. Similarly with your mention of the CFAA (UK CMA has some similar terms), they're very loosely drafted. Not accounting for the need to communicate the limits of allowed access is silly though; there's a presumption of allowed access online IMO (that isn't mirrored offline) and going against that presumption should require explicit withdrawal of consent. If I were drafting the law ...
- xexers 10y agoI just tried this link and it gave me a price of 154: https://www.amazon.ca/Samsung-Smartphone-32-GB-Unlocked-International-Warranty-Silver/dp/B01D4GD98G/ref=sr_1_2?s=wireless&ie=UTF8&qid=1472655909&sr=1-2&keywords=Samsung+Galaxy+s7 https://www.amazon.ca/Samsung-Smartphone-32-GB-Unlocked-Inte...
- malcolmhere 10y agoI worked on a product like this for a long while, so I can appreciate how hard the problem is you are trying to solve. Some things I realized along the way: * Price capture needs to happen in a headless browser (e.g. PhantomJS), rather than just capturing the HTML with a GET. Too many sites use JavaScript to make raw HTML analysis feasible. * You can get > 50% of the pricing information with fairly simple matching on the class/id value in the HTML tag. But you need a headless browser to make sure the tag is visible. And since most product pages contain multiple prices, you need some heuristic to determine the relevant price. Oh, and watch of out for "reduced from" prices too (e.g. "Old Price: $50, New Price: $35". * It doesn't hurt to be able to override the general heuristic on a domain-by-domain basis, saved me a lot of headaches. * You need to be honest with yourself about how reliable the price capture algorithm is, and built up a regression database of known good pages, so when you change the algorithm, nothing else breaks. Also, you need to keep ahead of site redesigns! * Product URLs tend to look messy, but tend not to change very often, if at all. I was worried about retailers e.g. changing product identifiers, but changing URLs hurts their SEO, so they don't do it. You will find "zombie" products, though - things which appear to be still on sale, but aren't linked anywhere on the site. Deciding when a product is sold out is tricky. * The best user experience presents the items the user is watching as a "shopping basket". (I took design cues from Pinterest.) For a really slick experience, you should pick out the product name and image (Facebook meta-data helps here) and include them in you "pinned" products. * Cutting-and-pasting URLs is a hassle. Consider writing a browser extension or a bookmarklet - users don't like to have the browsing flow interrupted by having to click across tabs. Having the price capture done inline on the page really impresses people. Best of luck with this! I'm yet to see someone solve this problem well, and I eventually moved on other things after losing a lot of my hair. :-)
- adibalcan 10y agoReally interesting feedback. Thanks!
- mikeash 10y agoI'd suggest a more descriptive title here. From the title, I thought/hoped it would give me price info for local stores, not just online. One of my constant annoyances is the fact that it's difficult if not impossible to find the lowest price for grocery items without physically visiting the stores you want to survey.
- adibalcan 10y agoYou are right. Thanks for feedback.