5 ms·
Rate limits on GitLab.com are changing
- tempest_ 16d agoI assume this is because of LLM scraping.
- nijave 16d agoPresumably but I wish they'd also focus on optimizing code/making pages fully cacheable instead of rate-limits and blocking
- swatcoder 16d agoMaking requests is inherently cheaper than delivering responses, even with caches. Efficiency improvements can buy a little time on a given resources but won't solve the problem of bot saturation now that everybody can spawn a custom bot in about 12 seconds and is being encouraged to do so. Rate limits, blocking, and pay-per-use are the only roads out and even those might not last as models get better at hacking and masquerading. The internet we want to use LLM's with is simply not one that can support LLM's, and with LLM's not going anywhere, the whole experience of the internet is going to be forced into some radically less open and more expensive paradigm. Policies like this just represent the beginning of the transition.
- jayd16 16d agoThe limits are for API requests, no? Or is this just an unrelated performance complaint?
- nijave 16d agoLast I checked, they just blanket added `cache-control: max-age=0` to everything. Loading https://gitlab.com/gitlab-org/gitlab https://gitlab.com/gitlab-org/gitlab is showing 14 API calls so presumably significantly more DB queries for a public page
- Frieren 16d agoWe need a new non-commercial version of the internet. Free of bots, free of ads, ... you pay to access social media optimized to be interesting enough to be worth paying instead of addictive enough to keep you scrolling to show you more ads. It may look impossible right now. But what is impossible for real is to continue as we are. The damage that internet does to society is increasing by the day while its value is reduced (economic value, social value).
- phoe-krk 16d ago> We need a new non-commercial version of the internet. Except it's non-commercial, therefore valuable, therefore commercialized, therefore commercial. You'd need a force strong enough to prevent it from falling prey to this tragedy of the commons, and that force would need to be stronger than the incentives to commercialize it. And that's where plenty of contemporary scraping-based salaries lay.
- immortalist 16d agoImpossible to create
- OtherShrezzing 16d agoI'm not so certain. I used to buy a broadsheet newspaper, which was full of ads. Now I pay a few hundred a year, and get the same newspaper online, with no ads at all. So, the precursor to online media has already gone through this paradigm shift.
- pixl97 16d agoI mean, what you're saying is "If I pay for a product ads go away" Which is partially true, but it only shifts the distribution of the problem. Once your service gains enough popularity network effects cause it to gain value. You have to worry about high priced buyouts of the entire service (great for the site owner, terrible for the users).
- Zambyte 16d ago
- ddtaylor 16d ago> A request that arrives with no credentials gets 60 requests per hour per IP address. One request per minute.
- Jaxan 16d agoDoesn’t help with scrapers though. They use a unique IP for each query.
- yjftsjthsd-h 16d agoBut it will hit any human behind a CGNAT:)
- mplanchard 16d agoAverage, yes, but the way they phrase it, it could be a token bucket or similar, where you can do 60 quick requests and then be blocked for a bit while the bucket refills
- sandeepkd 16d agoMore like you have 60 requests, you can exhaust them in a single second or spread them differently as per your choice.
- MeetingsBrowser 16d agoHopefully they bump this up. Browsing open issues or reviewing a few PRs will easily use more than one request per minute. The limits are based on the average user but I wonder if the most common interaction is to view a readme and bounce. I don’t know that putting a paywall up to learn from or even consider contributing to public projects is a good thing.
- silverwind 16d agoOr a lot more with IPv6.
- ddtaylor 16d agoMost connections or VPS have a shared 64bit prefix that acts similar to an IPv4 address in that it's easy to block. Yes, you get 64 more bits to make whatever addresses you want, but the prefix is still your fingerprint.
- sparkling 16d agoI noticed that recently Github.com has some kind of weird bot detection on public repos. I have a browser extension for switching User Agents for a specific legacy site, sometimes i forget to turn it off and Github will require me to login to view public repos. All of this is most likely due to mass scraping by LLMs. Welcome to the total shitification of the web.
- Macha 16d agoThe post is about gitlab but on the subject of GitHub, the rate limit I seem to have for viewing commit history seems to be 0 for logged out users and viewing source files seems to be about 5/hour.
- 296012 16d agoCongrats on making the world worse with AI. All this performative data scraping and uploading and no progress at all.
- micromacrofoot 16d agowhat are you talking about? it's progressing a lot of money into specific people's pockets
- speedgoose 16d agoNo progress?!
- IhateAI_3 16d ago[dead]
- deleted 16d ago[deleted]
- ephemerally16c4 16d agoMaking the privileged money and giving them the power to manipulate the mass is progress to some.
- jtwaleson 16d agoI think it's because people are building agentic flows, reducing the amount of developer seats needed. It's the first step towards usage based pricing.
- cush 16d agoProviding kickbacks to the repos being scraped would be a good way to help fund open source projects and pay creators like streaming services do. Seems like they're headed in this direction - it would be a massive product differentiator over GH
- latexr 16d agoThree minutes after kickbacks were announced there would be a flurry of new repos being created with bots repeatedly scraping them just to get those kickbacks.
- aprentic 16d agoMy first reaction was that I really like this idea. If we had a system where people who access projects pay and popular FOSS developers get paid for it we'd have much better alignment. My second thought was that bots would immediately try to circumvent such a plan. They'd probably spam Gitlab with fake repos to try to harvest those payouts.
- cush 16d ago> They'd probably spam Gitlab with fake repos to try to harvest those payouts. Yeah I wonder if the math would shake out to make that make any sense. Each bot would require a paid subscription, so the only incentive for them to do this would be if there was some discoverability algorithm or SEO that that traffic helped push the content to real users
- serhack_ 16d agoI would spend thousands of dollars for gitlab in terms of: 1) better UX for admin panel, I'm not sure what I've enabled and what not. Several buttons do not disable the rest of the settings, leaving me with some doubts (e.g. if I disabled grafana, why is there a setting that talks about where/how I store?) 2) a minimal version of gitlab without all the AI
- theokrueger 16d agoGithub would hit four nines if they followed suit. no clue why the dont try
- pkaye 16d agoDoesn't GitHub already have rate limiting especially if you are not logged in?
- Jaxan 16d agoYes. A lot is not accessible if not logged in.
- theokrueger 16d agoit's too generous and contributes to their reliability woes
- TiredOfLife 15d agoIt is the same 60 requests per hour
- shimman 16d agoBecause that would go against MSFT's wishes on pushing LLM driven development if they started acknowledging that these tools are more damaging than helpful.
- RomanKornev 16d agoGitHub's "paid" rate limits are the same as GitLab's free tier, so no, it doesn't help
- Retr0id 16d agoI understand why they're doing this, but the anticausative title kinda rubs me the wrong way.
- bearjaws 16d agoI am honestly surprised they aren't going lower at this point. Gitlab must pay a fortune to bot traffic, most of which is malicious or garbage at best.
- demibabs 16d agoDamn, we’re even having Claude write important press releases now
- geodel 16d agoDunno, for folks around here, Claude is important, Gitlab is important and banal press release is important. So all of them together makes it obviously important.
- vips7L 15d agoClaude sucks.
- vips7L 16d agosad days
- zb3 16d ago"Write a post about decreasing limits so that this information is as obfuscated as possible."
- Jeremy1026 16d agoI put the text in to gptzero's AI checker and it came back with being Highly Confident it was 100% AI written. I don't think I've ever seen it that confident that the entire thing was AI before.
- rcxdude 16d agoClaude's style is obvious enough I don't think you really need the checker.
- Jeremy1026 16d agoYes it is, but 100%? Typically someone goes through and touches something to clarify or add a detail.
- 16d ago
- jmclnx 16d ago> The requested URL was not found on this server. Getting that so I do not know exactly what they are doing. From the title I am guessing they are restricting or throttling if downloads exceeds some value.
- bob1029 16d agoIf you are using LLMs to interact with sites like GitLab and GitHub, and you have the option to use a GraphQL API, you should jump on it immediately. GraphQL is absolutely terrible for human developers to interact with, but it's like Facebook could see into the future back in 2012. I cannot imagine a more perfect API surface for agents. With the REST API on GitHub, you can consume maybe 10 issue JSON blobs before your context window is blown out. With GraphQL constraining the results you can easily read hundreds in the same token budget. Additionally, the # of requests your agents need to make can be reduced in many cases since GraphQL can join across types whereas REST APIs cannot. You essentially get savings in two dimensions here. Quota and raw token volume per logical response.
- aschobel 16d agomine just is the gh cli. is the advantage of graphql that they can compose a query that would take multiple cli invocations?
- mattkrick 16d agoCLI still has the possibility of being a little more token efficient, at worst it may use graphql behind the scenes. For our company, we advertise the graphql schema to bots and they can one-shot whatever task they're trying to do. I've found that it's so good that I cancelled building an an MCP server and any skill. Just a well documented GQL schema. It's pretty remarkable.
- saasrivals 16d ago[flagged]
- enormousness 16d agoI think you might also have a GQL schema that's particularly well suited to what those bots need. Either that or your backend schema is relatively simple and the GQL schema covers all its possible data compositions.
- agentdev001 16d ago
- MeetingsBrowser 16d agoWorst case, this could be the start of a paywall to learn from, contribute to, or host open source projects. Hopefully they find some kind of carve out for OSS projects while still blocking the egregious offenders.
- rkagerer 16d ago60 requests per hour per IP if you haven't signed in... well that's unfortunately low.
- lateatdesk 16d ago60 requests an hour per IP seems low for a school or office network. A few people browsing issues and source files could use that up quite fast.
- throwitaway222 16d agoWalk back in 10...9...
- mschuster91 16d ago> You get 429 Too Many Requests with RateLimit-* headers and a Retry-After. Wait the interval it gives you, then retry. Is there a test endpoint where one can validate the behavior of their ratelimit detection? Basically I do not want to cause excessive load on your servers just to test my implementation.
- PaoloBarbolini 16d agoIs this going to apply just to `https://gitlab.com/api/v4/:rest_of_the_url https://gitlab.com/api/v4/:rest_of_the_url` endpoints, or also to the API-ish endpoints like `https://gitlab.com/:org/:repo/raw/HEAD/:path https://gitlab.com/:org/:repo/raw/HEAD/:path`?
- xyst 16d agoThis is why I moved my code to self hosted forgejo instance. Private and guarded behind self hosted OIDC instance. No more worrying about "rate limiting," subscription hell, or random extended outages (ie, github). If LLM wants access, might implement payment layer and use 402 http status code and redirect them to payment page ;). Wonder how many people just give agents carte blanche physical (credit card) and virtual access
- nkapias 16d agoI forgot a superfluous free tier proof of concept pipeline and it ran every 6 hours for three months before I remembered to shut it off, sorry. I expect to not be the only one, it certainly drives usage KPIs up and lead to this kind of decisions.
- solatic 16d agoImportant buried context: 60/hour unauthenticated, but 5,000/hour on the free plan. 60/hour sucks. 5000/hour (a little more than one per second) is totally fine. I'm chalking this up alongside Docker's decision to restrict unauthenticated pulls. Unauthenticated anything went the way of the dodo some time ago. If you want unauthenticated access, go run your own mirror.
- 0xbadcafebee 16d agoWhich is fine, and how these businesses should've been run from the start. Require auth, use the sign-in to increase conversion, that increases revenue, that revenue's used to increase capacity. When free plans become a drain on paid users' experience you're shooting yourself in the foot. If you want to support increasing free use you need to increase paid use. Do that later and the users revolt; do it early and it's normal.
- htrp 15d agoblog post should've been titled, have your agents sign in
- colek42 16d agoCI needs to be decentralized, and the agents should run the test and verify the proof. This avoids all those API calls.
- Tatendaz 16d agoim self hosted so im somewhat safe i guess
- weedfroglozenge 16d agoTo anybody saying this is brought on due to AI scraping - It isn't, they are just doing this at a time that can be seen as a valid excuse. Countering AI scraping is a solved problem. This has been rolled out to bring in more subscriptions and more dollars.
- TheRealPomax 16d agoI mean, with a completely new C-suite gutting the company, that sure isn't the only thing that's changing.
- momed-0 15d agodoes this rate limit also apply to the images which they hosted on registry.gitlab.com? They do have a couple of scanners that we use and we pull it without any authentication
- accountrequired 15d ago429 gang represent