4 ms·
I don't think it is that. I work for an org with close ties to arXiv, and just like us they are getting a lot more demand due to AI crawling. As a primary sourc
by specialp 1y ago
I don't think it is that. I work for an org with close ties to arXiv, and just like us they are getting a lot more demand due to AI crawling. As a primary source of information there is a lot of traffic. They do have technical issues from time to time due to this demand, and I think their stability is just due to the exceptional amount of effort they take to keep it going. They are also getting more submissions and interest.
Kubernetes does add complexity but it does add a lot of good things too. Auto scaling, cycling of unhealthy pods, and failover of failed nodes are some of them. I know there is this feeling here sometimes that cloud services and orchestrated containers are too much for many applications, but if you are running a very busy site like arXiv I can't see how running on bare metal is going to be better for your staff and experience. I don't think they are naive and got conned into GCP as the OP alludes to. They are smart people that are dealing with scaling and tech debt issues just like we all end up with at some point in our careers.
- CharlieDigital 1y agoSound like all they needed was a CDN if the problem is AI crawlers. Adding auto-scaling compute just increases costs faster.
- specialp 1y agoCDN is one part of strategy to deal with load. But it is not the only solution unless your site is exclusively static content. Their search, APIs, submission pipelines, duplicate detectors and a lot of other things are not going to be powered by CDNs.
- motorest 1y agoThank you for the insight. It's very easy to prescribe simple solutions when we are oblivious to the actual problems being solved.
- coliveira 1y agoAll these services can be throttled to deal with AI. I don't see this as a justification. The idea that a service like arXiv should be run as a startup is, simply put, foolish.
- Imustaskforhelp 1y agopardon me but cloudflare workers seem better for this approach. If we can get for the fact that we require javascript to run it, aside from that. Cloudflare workers is literally the best single thing to happen at least to me. With a single domain, I have done so many personal projects for problems I found interesting and I built so many projects for literally free, no Credit card. No worries whatsoever. I might ditch writing other languages for server based like golang even though I like golang more just because cloudflare workers exists.
- anelson 1y agoI too am impressed by Cloudflare Workers’ potential. However Workers supports WASM so you don’t necessarily have to switch to JavaScript to use it. I wrote some Rust code that I run in Cloudflare Functions, which is a layer on top of Cloudflare Workers which also supports WASM. I wrote up the gory details if you’re interested: https://127.io/2024/11/16/generating-opengraph-image-cards-for-my-zola-static-blog-with-cloudflare-functions-and-rust/ https://127.io/2024/11/16/generating-opengraph-image-cards-f... JavaScript is most definitely the path of least resistance but it’s not the only way.
- GTP 1y agoBut, under the assumption that the problem are indeed AI crawlers, of the things you listed only the search would be under increased load.
- specialp 1y agoI am sure their pages are not entirely static either. The APIs are used by researchers and AI companies too. Also with search you end up having people trying to use it for RAG with AI. I have dealt with this all and there is no one dead simple solution to deal with things. AI crawlers are one part of things, but they also have increasing submissions, have to deal with AI generated spam papers, and all sorts of stuff. There's always this feeling here on HN that oh it is dead simple you just do "X" as if the people that are dealing with it don't know that.
- deleted 1y ago[deleted]
- miyuru 1y agothey already use fastly. https://blog.arxiv.org/2023/12/18/faster-arxiv-with-fastly/ https://blog.arxiv.org/2023/12/18/faster-arxiv-with-fastly/
- JackC 1y ago> I work for an org with close ties to arXiv, and just like us they are getting a lot more demand due to AI crawling Funny, I also work on academic sites (much smaller than arXiv) and we're looking at moving from AWS to bare metal for the same reason. The $90/TB AWS bandwidth exit tariff can be a budget killer if people write custom scripts to download all your stuff; better to slow down than 10x the monthly budget. (I never thought about it this way, but Amazon charges less to same-day deliver a 1TB SSD drive for you to keep than it does to download a TB from AWS.)
- Imustaskforhelp 1y agoI don't understand, why don't you use cloudflare? Don't they have an unlimited egress policy with R1? Its way more predictable in my opinion that you only pay per month a fixed amount to your storage, it can also help the fact that its on the edge so users would get it way faster than lets say going to bare metal (unless you are provisioning a multi server approach and I think you might be using kubernetes there and it might be a mess to handle I guess?)
- mcmcmc 1y agoCould have something to do with Cloudflare’s abhorrent sales practices.
- keepamovin 1y agoCan you tell me more? I think my business needs some abhorrent sales practices. That's how it's done, right?
- mcmcmc 1y agoOne example https://robindev.substack.com/p/cloudflare-took-down-our-website https://robindev.substack.com/p/cloudflare-took-down-our-web...
- keepamovin 1y ago
- Funes- 1y ago>AI crawling Can you not reliably block crawlers in this day and age?
- masklinn 1y agoAI crawlers are a plague, they are intentionally badly behaved and designed to be hard to flag without nuking legit traffic. That’s why projects like nepenthes exist.
- specialp 1y agoYou can to some degree with Cloudflare and other solutions. But, do you want to block them all? AI is a very useful tool for people to discover information and summarize results. Especially in scholarly publishing where one would have to previously search on dumb keywords, and have to read loads of abstracts to find the research that pertains to their interests. So by blocking AI crawlers and bots completely, you are shutting off what will probably end up being the primary way people use your resource not too long from now. arXiv is a hub of research, and their mission is to make that research freely available to the world.
- Imustaskforhelp 1y agoDude, I don't want to sound cloudflare advocate because my last 2 comments on this thread are just shilling cloudflare... but I think cloudflare is the answer to this thing as well.. (Sorry if I am being annoying) (Cloudflare isn't sponsoring me, I just love their service so much)
- evrythingisfin 1y agoGC seems like a bad choice to me. The GC CLI and tools aren’t terrible, but their services in-general have historically been not been friendly to use and their support involves average documentation and humans that talk at users instead of serve them. Google as a company is not what it was, either. A lot of their funding was driven from advertising, and between social media, streaming, and LLMs, that seems like it may be starting to dry up. Azure? Microsoft as a company is still the choice of most IT departments due to its ubiquitousness and low cost barrier to entry. I personally wouldn’t use Azure if I had the choice, because it’s easy and cheap at the surface, with hell underneath (except for products based on other things, like AD which was just a nice LDAP server, or C# which was modelled after Java). I’d have gone with AWS. EKS isn’t bad to setup and is solid once it’s up. As far as the health of Amazon itself, China entering the their space hasn’t significantly changed their retail business, though eventually they’ll likely be in price wars. The greatest risk to any cloud provider I think would be a war that could force national network isolation or government taking over infrastructure. And the grid would go down for a good while and water would stop, so everyone would start migrating, drinking polluted water, then maybe stealing food. At that point, whether or not they chose GC doesn’t matter anymore.
- Kwpolska 1y agoIsn't Azure noticeably more expensive than AWS?
- surajrmal 1y agoEverything has pros and cons. Just because their calculus came out different from yours doesn't mean they made the wrong decision for their situation. Hundreds of thousands of organizations have made similar conclusions to arxiv. Google cloud profitable these days and advertising or other income streams drying up will only entice Google to further invest in cloud to ensure they are more diversified. Google isn't going to go away overnight and cloud is perhaps the least risky business they operate in.
- sitkack 1y agoThey don't need K8S, containerization yes, but not K8S.