13 ms·
Show HN: Ladder, open source alternative to 12ft.io and 1ft.io
Hey there
I made a opensource alternative for these services. Although these worked very well, I was not so confident what they do. So I made my own and opensourced it.
It is written in Golang and is fully customizable.
- SigmundurM 3y agoYou mention 13ft as another open source inspiration. How is Ladder improving on what 13ft does?
- 2cpu1container 3y agoI did try 13ft. But it misses several points. The ladder applies custom rules to inject code. It basically modifies the origin website to remove the Paywall. It rewrites (most of) the links and assets in the origins HTML to avoid CORS Errors by routing thru the local proxy. The ladder uses Golangs fiber/fasthttp, which is significantly faster than Python (biased opinion) . Several small features like basic auth ...
- withinboredom 3y ago> The ladder uses Golangs fiber/fasthttp, which is significantly faster than Python I have a feeling that this performance difference is practically imperceptible to regular humans. It's like optimizing CPU performance when the bottleneck is the database.
- ComputerGuru 3y agoNot for any publicly hosted instance, it’s not. We’re not talking about the time it takes to perform one request but the scalability it affords a small vm to handle so many requests in parallel when it is being used by the general public.
- oh_sigh 3y agoIf the paywall is implemented in client code, then usually just disabling javascript for the site is enough to let you view it. If it is implemented server side, then there usually isn't a way around it without an account.
- roydivision 3y agoIs it just me or has 12ft become less and less effective? I rarely get through with it these days.
- user_7832 3y agoTheir policies have apparently… changed. They accept donations to not have your website bypassed. Archive.org is much better. Edit: apparently it is down now. 402: PAYMENT_REQUIRED Code: DEPLOYMENT_DISABLED ID: fra1::8wkv2-1699275385535-39dedae23d6a
- jdiff 3y agoIs it donations they accept or legal threats?
- ProllyInfamous 3y agoYes.
- i67vw3 3y agoArchive.today never fails compared to Archive.org or various browser extensions To remove paywalls 12ft<Archive.org<Archive.today is my opinion.
- giancarlostoro 3y agoFor some reason all the alternative "archive.XYZXDHWIQHDQ" type of sites always give me a captcha page, and I am never able to proceed. I'm assuming its to do with the cloudflare DNS, well if they don't care to fix it on their end, I don't care to use their service.
- i67vw3 3y agoThere is a bit of 'tussle' going on between the two of them for quite a few years as you pointed out. https://x.com/archiveis/status/1018691421182791680?s=20 https://x.com/archiveis/status/1018691421182791680?s=20 https://news.ycombinator.com/item?id=36971650 https://news.ycombinator.com/item?id=36971650 https://news.ycombinator.com/item?id=19828702 https://news.ycombinator.com/item?id=19828702 https://news.ycombinator.com/item?id=36971552 https://news.ycombinator.com/item?id=36971552
- fader 3y agoFor folks like me who have no idea what 12ft.io or 1ft.io are, they appear to be services for bypassing paywalls on websites.
- alberto_ol 3y agoPrevious dicussions of the service on HN: https://hn.algolia.com/?q=12ft.io https://hn.algolia.com/?q=12ft.io
- 2cpu1container 3y agoThose were Paywall bypassing tools. 12ft.io was shut down one week ago and 1ft.io still works. But I feel a bit unconfident to let someone inject code to sites i view.
- deleted 3y ago[deleted]
- ktpsns 3y agoI got the feeling that these features should be part of a browser extension the same way as there are AdBlock extensions. I guess the reason it is not is "personal preference" of the author, or is there some technical reason?
- bilekas 3y agoI don't know for sure, but I would imagine there are more severe actions taken against circumventing paid material (content behind a paywall) than there is for free content supplemented by advertisements.. Edit : The Digital Millennium Copyright Act (DMCA) prohibits circumventing an effective technological means of control that restricts access to a copyrighted work. I guess that would apply here.
- mckirk 3y agoGiven how liberally the DMCA is applied, you definitely don't want to be on the wrong side of that. I remember some guy that wrote a WoW bot and got sued using the DMCA, with the argument that his bot was circumventing the anti-cheat and the anti-cheat could be seen as a 'mechanism protecting copyrighted material', because it was safeguarding access to the game servers, the servers were generating parts of the game world (such as sounds) dynamically, and those were under copyright... Wild stuff.
- judge2020 3y agoAs far a I know section 1201 has never been prosecuted. Distribution of the copyrighted material is what's focused on.
- mckirk 3y agoThis seems a good summary of the case I was talking about: https://massivelyop.com/2020/02/28/lawful-neutral-cheating-copyright-law-and-the-wow-glider-lawsuit/ https://massivelyop.com/2020/02/28/lawful-neutral-cheating-c...
- 3y ago
- some1else 3y agoRelevant: 12ft.io was banned by Vercel, taking down the developer's entire account with multiple other hosted projects & domains: https://twitter.com/thmsmlr/status/1718663563353755982 https://twitter.com/thmsmlr/status/1718663563353755982 Edit: Access to other projects & domains was apparently restored some time after: https://twitter.com/thmsmlr/status/1719480558932148272 https://twitter.com/thmsmlr/status/1719480558932148272
- deleted 3y ago[deleted]
- abofh 3y agoLovely, the Google classic "ban the world" approach -- I've been desperately trying to move my client off of vercel, this might just be the gasoline.
- rgrieselhuber 3y agoAside from this (which is already very shitty and would cause the same response in me) what are the issues you’re running into with Vercel?
- abofh 3y ago- Support is failing us - I want my team to use you for vercel support, but it isn't there. - Support is failing our customers - when you fail, I end up reverse-depending your repo to tell us why it's failing -- just give us a clear answer, we all move away happy, bullshit and I go to lambda where I just accept it. - EOD: Vercel makes engineers happy to bullshit, but gives operations teams nothing acceptable - I want a deliverable product.
- judge2020 3y agoFYI you have to use two line breaks to start a new line with (HN's) Markdown.
- 3y ago
- janejeon 3y agoReally dummy question: how do services like this work? As in, how do they bypass these paywalls? The obvious thing is to mock Googlebot, but site owners can check that the request isn't coming from a Google-published IP and see that it's a fake, right?
- narinxas 3y ago> site owners can check that the request isn't coming from a Google-published IP and see that it's a fake, right? just because they can doesn't mean they will... also most "site owners" are (by this point) a completely different people than "site operators" (who I take to be the 'engineers' who indeed can check this IP things)
- calflegal 3y agorelated: If this is how they work, why doesn't google offer a private service to allow publishers to have content indexed while still protected?
- matsemann 3y agoIt used to be against guidelines to serve different content to google vs what users would see. Not sure if still the case, but I don't think it's in google's interest to give a result that the user actually can't access.
- ComputerGuru 3y agoI’m not aware that this policy has changed. What has changed is that Google will rank results it can’t (officially) index without showing their content. I’m guessing they do shadow index them but use the whole “if you outwardly can’t tell they did then it’s as if they didn’t” C++ compilers use to get away with insane optimizations.
- Fnoord 3y agoSome possible clues: > https://github.com/kubero-dev/ladder#environment-variables https://github.com/kubero-dev/ladder#environment-variables > USER_AGENT User agent to emulate Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html http://www.google.com/bot.html) > X_FORWARDED_FOR IP forwarder address 66.249.66.1 > RULESET URL to a ruleset file https://raw.githubusercontent.com/kubero-dev/ladder/main/ruleset.yaml https://raw.githubusercontent.com/kubero-dev/ladder/main/rul... or /path/to/my/rules.yaml
- pacifika 3y agoOpen source makes it easy for the cat in the cat mouse game, right?
- lucideer 3y agoThere's no real cat & mouse game here (yet*) - sites don't do anything to mitigate this. Sites deliberately make their content available to robots to gain SEO traction: they're left with the choice of allowing this kind of bypass or hurting their own SEO. * I say "yet" because there could conceivably be ways to mitigate this, but afaik most would involve individual deals/contracts between every search engine & every subscription website - Google's monopoly simplifies this somewhat, but there's not much of an incentive from Google's perpsective to facilitate this at any scale.
- tiagod 3y agoGoogle publishes IP ranges for GoogleBot. You can also reverse-lookup the request IP address - the resolved domain should in turn resolve to the original address.
- ForkMeOnTinder 3y agoDoes anyone else remember 10 years ago when Google would penalize sites for serving different content to GoogleBot than to normal users? Those were the days.
- omoikane 3y ago> Google would penalize sites for serving different content to GoogleBot than to normal users Listed under spam policies: https://developers.google.com/search/docs/essentials/spam-policies#cloaking https://developers.google.com/search/docs/essentials/spam-po... "Cloaking refers to the practice of presenting different content to users and search engines with the intent to manipulate search rankings and mislead users" The top of the pages says sites that violate the policies may "rank lower or not appear in results at all".
- fyzix 3y agoI'm very new to this kind of service, but do you have to write your own rulesets for each site you want to bypass? The repo doesn't seem to include much...
- 2cpu1container 3y agoYes, the one i provide is still pretty empty yet. I plan to build one that can be used as a starting point or as a default.
- szaboat 3y agoNot relevant to the project but I usually check for earlier versions of the paywalled pages in the wayback machine (~75% success). I felt bad using these services (paywall removers), and just feeling a bit better checking in archive.org.
- jwmoz 3y ago12ft was really good!
- 2cpu1container 3y agoIn deed it was. Sad it's gone. One single downside was the intransparency. It was not clear which code was added or removed on the site you where looking at.
- KoftaBob 3y agoCreate a browser book mark and set this as the URL of the bookmark: javascript:window.location.href="https://archive.is/latest/"+location.href https://archive.is/latest/"+location.href It will usually open up the archived version of article without the paywall.
- gumby 3y agoThe README says "The author does not endorse or encourage any unethical or illegal activity." Is it actually illegal anywhere to bypass a paywall?
- 2cpu1container 3y agoNot sure about the paywalls. But it might be used for "drive by attacks" or phishing.
- qingcharles 3y agoCertainly in Illinois it would be a crime to violate the TOS of a website. Misdemeanor for first offense, felony for second, IIRC.
- quickthrower2 3y agoCan’t be that simple. What if TOS has ridiculous shit in it. Stuff about life long servitude to the webmasters pet goldfish, for example?
- muttled 3y agoBelieve it or not, straight to jail.
- quickthrower2 3y agoLink to a story like that?
- qingcharles 3y ago(720 ILCS 5/17-51) (was 720 ILCS 5/16D-3) Sec. 17-51. Computer tampering. (a) A person commits computer tampering when he or she knowingly and without the authorization of a computer's owner or in excess of the authority granted to him or her: (1) Accesses or causes to be accessed a computer or any part thereof, a computer network, or a program or data; (a-10) For purposes of subsection (a), accessing a computer network is deemed to be with the authorization of a computer's owner if: (1) the owner authorizes patrons, customers, or guests to access the computer network and the person accessing the computer network is an authorized patron, customer, or guest and complies with all terms or conditions for use of the computer network that are imposed by the owner;
- JustinGoldberg9 3y agoI still miss outline.com I use txtify.it
- donohoe 3y agoI use services like this as I often skip news site paywalls because I just can't afford, nor is it practical, to have so many subscriptions. That said, I work in news media (and have been involved in building paywalls at different orgs - NYT and New Yorker). I know how money for these directly support journalism - salaries and the costs with associated with any story. If you are skipping paywalls a lot, I would encourage you to pay for a subscription to at least one or two news sites you respect - bonus points if its a small or medium local newsroom that benefits! For me that has been; NYTimes, New Yorker, Wired, Teen Vogue, and my wife's hometown paper in Illinois.
- mejthemage 3y agoThere's a huge need for subscription bundles. I'd gladly pay $20/mo for access to a bunch of big names, even if I'm limited to like 60 articles per month combined across those sources. Instead I just don't pay anyone, turn back when I encounter a paywall and look for someone's summary if I'm really interested.
- orpheansodality 3y agoIsn’t that the value-prop of Apple News?
- stetrain 3y agoIn my experience Apple news relies entirely on the app for reading articles, so if you are on a computer without the app or following a link that doesn't auto-redirect to the app then you still hit the paywall. I'd rather have a system that was just a cross-website web account.
- basch 3y agoApple News is a bit of a discovery pain in the neck. If I have a 5 year old Atlantic article, I can’t just click the article and have it open in Apple News. I can’t search for it. If the article is any older, the magazine won’t appear at all.
- 3y ago
- rounakdatta 3y agoGiven a very different paywall model for Substack, what exactly would work for bypassing their paywalls? Wouldn't we always require a paid account to cache the HTML through (the SciHub model)?
- arendtio 3y agoSounds great, not just for paywalls, but for removing CORS as well: > Remove CORS headers from responses, assets, and images ...
- j-a-a-p 3y agoIn the README there is a WHY paragraph: > Freedom of information is an essential pillar of democracy and informed decision-making. While media organizations have legitimate financial interests, it is crucial to strike a balance between profitability and the public's right to access information. The proliferation of paywalls raises concerns about the erosion of this fundamental freedom, and it is imperative for society to find innovative ways to preserve access to vital information without compromising the sustainability of journalism.
- j-a-a-p 3y agoFor me this is grotesque. Democracy is in dispair so is journalism. What exactly is this software doing to support journalism or democracy?
- 2cpu1container 3y agoWe live in a world, where we have more misinformation and poor journalism every day, and less money in the pockets of the people to afford paying for good journalism. So this might start a more open discussion on how to finance journalism. And while discussions are still going on, people can inform themselves with good journalism, which supports the democracy.
- j-a-a-p 3y ago> And while discussions are still going on, people can inform themselves with good journalism, which supports the democracy. That is a looters mentality, sorry to say that. Paywall jumping software is like robbing a disabled old veteran in public transport. It are the last blows to finish off what was once good journalism.
- deleted 3y ago[deleted]
- boplicity 3y agoSlightly edited "Why": Access to private property is an essential pillar of democracy and the safe proliferation of ideas. While property owners have legitimate financial interests, it is crucial to strike a balance between property and the public's right to access property. The proliferation of locks on doors raises concerns about the erosion of this fundamental freedom, and it is imperative for society to find innovative ways to preserve access to people's homes and workspaces without compromising the sustainability of property ownership.. In a world where property should be shared and not commodified, locks should be critically examined to ensure that they do not undermine the principles of an open and informed society.
- deleted 3y ago[deleted]
- 2cpu1container 3y agoNice analogy. But with ladder, you're just wearing a Google shirt and get invited.
- benatkin 3y agoThis reminds me of the thread when 12ft was taken down. Does anyone have any insight into how it would take Vercel hundreds of hours of support time? https://twitter.com/rauchg/status/1718680650067460138 https://twitter.com/rauchg/status/1718680650067460138
- someotherperson 3y agoMy assumption here is that affected websites sent multiple, persistent support tickets and engaged in back and forth communication, as well as updates to the client, support team contacting engineering/legal/management/meetings on how to deal with 12ft.
- deleted 3y ago[deleted]
- TanguyN 3y agoI have noticed that on a lot of websites, if you stop the page loading at just the right moment (you have to be quick), the whole content will display without the paywall. And that's without any external tools. These kinds of tools seem, of course, much more convenient.
- serial_dev 3y agoFirst of all congrats on the project and thank you for open sourcing it. > Freedom of information is an essential pillar of democracy However, this reads like this tool saves democracy by letting you bypass a crappy pay wall on a site you visit once a year, and that whoever wants to get paid for their published content online is an enemy of democracy.
- UncleEntity 3y agoIronically, the rest of the paragraph you quoted from gives their reasoning why they believe this tool is needed beyond "whoever wants to get paid for their published content online is an enemy of democracy". Double-plus democracy and all that...
- deleted 3y ago[deleted]
- karaterobot 3y agoIt seemed to me like 12ft.io was useful for a couple of months, but then stopped being useful as they agreed to blacklist more and more URLs. I thought everybody switched to archive.is, which (so far) works 100% of the time, even if it is sometimes a pain in the butt.
- Axsuul 3y agoIs there an open source version of archive.is?
- raxi 3y agoJust webrecorder + magnolia and you'll get something similar. Maybe even better: magnolia outperforms archive.is on paywalls
- metadat 3y agoThe operator of archive.is must constantly re-up on hacked credentials for wsj and nyt. Given this is a critical aspect of the service, it is not really feasible/useful to open source it.
- deleted 3y ago[deleted]
- cooper_ganglia 3y agoReally great and easy to use. I was trying to read an article that was on the front page of HN and couldn't due to paywall. Downloaded the binary and was reading it within 30 seconds. Awesome and very useful tool, thanks!
- nfriedly 3y agoThe docker image, and on the upside is fairly easy to get running. But I'm downside, I'm zero for two actually using it. I tried a Bloomberg article which gave me a "suspicious activity from your IP, please fill out this captcha" page, only the captcha was broken and didn't load. Then I tried a WSJ article which loaded basically the same couple of paragraphs that I could get for free, but did not load any of the rest of the content.
- zippytyro 3y agodamn, thanks man