23 ms·
Chess.com stopped working on 32bit iPads because 2^31 games have been played
- prh8 9y agoReal world example of why Apple is killing 32 bit apps on iOS.
- dfox 9y agoThis has nothing to do with CPU architecture.
- mikeash 9y agoIt's a little related. The languages typically used for iOS programming encourage the use of data types whose size matches the CPU architecture's bitness. Thus, careless programmers will end up using 32-bit integer types on 32-bit devices, and 64-bit types on 64-bit devices. I really doubt this is in any way linked to Apple's reasons for dropping 32-bit, though.
- _puk 9y ago"The reason that some iOS devices are unable to connect to live chess games is because of a limit in 32bit devices which cannot handle gameIDs above 2,147,483,647."
- paulddraper 9y agoThat is a combination of code and architecture, not architecture alone.
- BinaryIdiot 9y agoThis is simply because they're taking an integer from their database as an auto incrementing ID and converting it directly into a native integer on the iOS device thus breaking it. They could work around this any number of ways. It's a pretty lame bug, to be honest and certainly something easily foreseeable as this wasn't an overnight occurrence.
- remus 9y agoHow so? A 32 bit platform doesn't mean you're limited to 32 bit ints....
- vxxzy 9y agoHow many other examples like this have occurred throughout computing history?
- SamReidHughes 9y agoPlenty, Twitter for example.
- dfox 9y agoIIRC at least slashdot and twitter had to do some non trivial database migrations because they hit maximum value of some ID field.
- mikeash 9y agoTwitter's was called the twitpocalypse and it was a pretty big deal at the time. I don't know how bad it was internally, but it hit a lot of third-party apps that stored tweet IDs in 32-bit integers. One interesting aspect of it was that Twitter realized what was going to happen in advance, and artificially pushed their IDs over the edge at a preplanned time so they could have as many people available as possible to work on any problems that appeared.
- MichaelGG 9y agoI think Instagram also recently (year?) passed the 31 or 32 bit mark, as some client lib I was using started failing.
- cm2187 9y agoWell, famously youtube's views counter overflow from Gangnam Style...
- 0003 9y agoThat was a joke though and they were prepared.
- nisse72 9y ago
- SomeHacker44 9y ago"This was obviously an unforeseen bug that was nearly impossible to anticipate..." Snarky... Except that there were probably years of games to notice that you were approaching a "magic number" like 2^31.
- CGamesPlay 9y agoI actually read that quote as fully sarcastic.
- blktiger 9y agoAs expected, sarcasm always translates correctly into textual form.
- i_cant_speel 9y agoIt's weird that you say that, because I always felt sarcasm didn't translate well to text.
- jazoom 9y agoI think it was sarcasm
- mirimir 9y agoSarcasm is recursive.
- brlewis 9y agoYours is the first comment in this chain that I can say pretty confidently isn't sarcasm. So it kind of breaks the chain, making the sarcasm in this chain non-recursive. Which means maybe you were being sarcastic after all? Actually I don't even know if my own comment is sarcasm or not.
- ivrrimum 9y agoWhats why kids you should use cryptic identifiers.
- rasz 9y agowere they ever expecting negative number of games? why signed integer?
- SamReidHughes 9y agoIt's very reasonable. This way they overflow into invalid values instead of zero.
- MereInterest 9y agoAssuming you're working in a language that defines signed integer overflow. Depending on the language, you can result in undefined behavior, instead. For that reason, I'd go with an unsigned counter, with the first million IDs being invalid or reserved for future use. That way, you get well-defined overflow into an invalid region.
- dpark 9y agoWhen's the last time you worked with an architecture that didn't use twos complement and roll into negatives on overflow? Your reserved bottom range is a perfectly good solution. But rolling into negatives seems fine, too.
- cm2187 9y agoSelf-confidence as a programmer is when starting a new project, storing the transaction ID as a long rather than an int...
- trendia 9y agoOr UUID...
- derefr 9y agoI pick UUIDs because experience has taught me that even for the smallest workloads, I'll inevitably have to shard my DB (to partition a shared public cloud from N isolated "enterprise" deployments), and then will inevitably want to do statistical-analysis things that involve ingesting rows from the shards (or log entries referencing those rows), and deduplicating them by ID, without generating false-positive collisions in the process. The simplest way to do that is to just throw UUID at all problems from the start. (https://github.com/alizain/ulid https://github.com/alizain/ulid s are better, but there aren't libraries to generate them in literally every language + RDBMS.)
- eru 9y agoI don't agree with all their reasoning---but ulid still seems like a good idea. (Though the main difference you care about in programming is how they are generated---via timestamps plus randomness, not that they have a different serialization format.) For some applications you don't want to leak the time. Choose wisely.
- rdtsc 9y ago> Self-confidence as a programmer is when starting a new project, storing the transaction ID as a long rather than an int... uint64_t even Or a UUID as others have suggested. Technically C spec doesn't really say exactly how many bits int, long and long long should be. If you want specific sizes and your code to be somewhat portable use the specific bit sizes to make that clear. There are also types for size-like things (size_t) and pointer and offset like things.
- ericfriday 9y agoThis reminds me YouTube changed its view counter from 32-bit integers to 64-bit integers due to the popularity of 'Gangnam style' https://www.wired.com/2014/12/gangnam-style-youtube-math/ https://www.wired.com/2014/12/gangnam-style-youtube-math/
- smitherfield 9y agoThat was a joke; it was always a 64-bit integer.
- akerro 9y agoWHAT?
- wnoise 9y agoDo you have a source for that?
- devrandomguy 9y agoI visited that video specifically because the view counter was jammed at UINT_MAX. There were comments confirming that everyone was now visitor number 4,294,967,295. In fact, it might have been an HN post that brought it to my attention; I totally didn't get sidetracked on YT and end up watching K-pop all afternoon.
- eponeponepon 9y agoIt's fascinating... the Y2K problem never came to fruition because - arguably - of the immense effort put in behind the scenes by people who understood what might have happened if they hadn't. The end result has been that the entire class of problems is overlooked, because people see it as having been a fuss over nothing. I sometimes think it would've been better if a few things had visibly failed in January 2000.
- mkempe 9y agoMy long-term active retirement plan involves building a year-2038 consultancy.
- deleted 9y ago[deleted]
- glhaynes 9y agoSimilar to how much effort goes into dealing with things like dangerous strains of bird flu, only to have people complain about how much money was spent on "nothing" when an outbreak doesn't occur.
- sbov 9y agoThis is a whole class of problem - I wonder if there's a name for it. More examples include: talking about welfare being unnecessary because no one is starving. Or people on medications stopping because they feel better (while still on them).
- nkrisc 9y agoThe "It's Working" problem.
- moufestaphio 9y agoLike when people say to IT: "Everything's working, what am I playing you for?!"
- throwaway2016a 9y agoReminds me of the havoc that was caused when Twitter tweet IDs rolled over. Resulting in every third party developer to update their apps (and at the time there were a lot of those). Twitter saw it coming and forced the issue. By saying that at a certain date and time they would manually jump the ID numbers rather than wait for it to happen at some unpredictable time.
- syncsynchalt 9y agoThey didn't roll over, they exceeded 2^53-1 which is the max Number which doesn't truncate when treated as an integer in js. The solution was to treat it as a string. (Or we're thinking of different events, I apologize if so)
- gilgoomesh 9y agoTwitter must have been misleading when they communicated the reasons for this change since they did not exceed 2^53-1, nor do they expect to exceed this in the near future. From a (former) Twitter dev: > Given the current allocation rate, they'll probably never overflow Javascript's precision nor get anywhere near the 64-bit integer space. https://twittercommunity.com/t/discussion-for-moving-to-64-bit-twitter-user-ids/9890/2 https://twittercommunity.com/t/discussion-for-moving-to-64-b...
- syncsynchalt 9y agoYour link discusses 2^64, which applies to languages that have native integer types. The 2^53 problem was for Javascript, which has no native integer type, and is thus limited by the mantissa size of Number (which is defined as an IEEE double-precision float). Twitter ids are unsigned 64-bit, since they're generated using Snowflake. That link must pre-date the move to snowflake ids, and is speaking to the count of tweets instead.
- MikeHolman 9y agoI'd be incredibly surprised if they overflowed 9 quadrillion tweets. That's like a million tweets per person on earth.
- shurcooL 9y agoDo we know when chess.com launched? If so, we can calculate the average number of games being played per second.
- CDRdude 9y agoWikipedia says "June 2007", which I'll approximate to 10 years. That gives us 6.8 games per second.
- jakub_g 9y agoLong long time ago, I created a poll on a small website I was maintaining. I didn't expect much traffic and, so, not thinking too much about it, I put the ID column to be a TINYINT (i.e. max value = 255)... That was a valuable lesson. (I actually generated most entries myself while testing stuff - live in prod of course - and while there were probably fewer than 255 votes, the AUTO_INCREMENT did its job and produced an overflow).
- ryandrake 9y ago> "Long long time ago" Seems you have learned your lesson :-)
- mattkenefick 9y ago"Obviously unforeseen.. impossible to predict." Really? You don't know how to properly store ID numbers? IMPOSSIBLE to predict.
- phailhaus 9y agoThat's clearly sarcasm.
- phonon 9y agoAnd I was just reading Heroku/Django discussing the same issue this morning! https://groups.google.com/forum/m/#!topic/django-developers/imBJwRrtJkk https://groups.google.com/forum/m/#!topic/django-developers/...
- abalone 9y agoSo they probably just need to use longs instead of ints. But I'm curious, if you were really stuck with a 32-bit limit on data types, what's your preferred workaround? I'm thinking I'd add another field that represents a partition. Are there other "tricks"?
- reformatt 9y agoIf you could only use 32-bit data types, you can get 64-bits from using 2 integers together like a long number. So the right integer would hit the max, start over at 0, and increment the left integer. Then, using this idea you can create a class of numbers that can have however many bits you want by using more ints.
- abalone 9y agoCool, yeah, that's what I meant by partitioning, which I guess is more of a database term.
- chesserik 9y agoHey all. Thanks for noticing :P Obviously this is embarrassing and I'm sorry about it. As a non-developer I can't really explain how or why this happened, but I can say that we do our best and are sorry when that falls short. - Erik, CEO, Chess.com
- Aloha 9y ago2 billion is a very large number that was probably not envisioned as reachable in the near future - as a programmer I'd argue this is a pretty easy mistake to make, and that while (slightly) embarrassing, its a good learning moment. It's also really awesome that you're here, and that you guys were so honest about the nature of the bug - this is really something that should be encouraged.
- chesserik 9y agoMaybe we should start a blog about all of the interesting bugs and challenges we encounter. It certainly is white-knuckle pretty often when running at scale. The number of devices, connections, features... I'm aging prematurely :P
- eru 9y agoA few articles would definitely be appreciated. Might even help with recruiting fresh blood.
- mishoo 9y agoAgree with Aloha. I wouldn't be too hard on the programmer (also, if I understand correctly it's not a database issue, but only with the 32-bit iOS client). I'd pat him on his back and say “you didn't think we'd get this big, eh?” ;-)
- nandemo 9y ago> 2 billion is a very large number that was probably not envisioned as reachable in the near future I disagree. Simple napkin calculation: 100 million players playing 40 games each per year (about 1 per week) over 5 years = 10 billion unique games. As others pointed out it was likely not a miscalculation, just a lack of calculation. The bug occurred only in the client and the decision to use a smaller data type was likely not a conscious one. In any case, I wouldn't hold it against an individual programmer. But arguably this sort of bug indicates your development process has flaws (not enough testing, code reviews, etc).
- pram 9y agoI recently experienced a nasty bug with BLOB in MySQL. The software vendor was storing a giant json which contained the entire config in a single cell. It ran fine for months, and then when it was restarted it totally broke. Reason was: the json had been truncated the entire time in the database, so it was gone forever. It was only working because it used the config stored in memory on the local system. Nasty!
- fpgaminer 9y agoFor those, like my, that didn't recall off the top of their heads: BLOB in MySQL can usually only hold ~64KB. EDIT: Though I am curious why MySQL doesn't throw an error when you try to store more than 64KB in BLOB?
- frandroid 9y agoThe BLOB size limit is a nuisance, the silent truncation is the bug...
- rsynnott 9y agoIt'll give a warning. Many/most MySQL drivers won't interpret it as an error, though.
- wnoise 9y agoHuh, I thought BLOB was a backronym for Binary Large OBject, not Binary Medium Object.
- p4lindromica 9y agoThere are LONGBLOB, TINYBLOB, etc just like the equivalent TEXT fields
- DCoder 9y agoMySQL's silent data truncation is such a nuisance. It's off by default in 5.7, and can be disabled in earlier versions by adding STRICT_ALL_TABLES/STRICT_TRANS_TABLES to sql_mode [1]. I inherited a system where, among other things, the entire response body from a payment gateway callback is saved into a text field using utf8 character set, despite the fact that most of the supported payment gateways send data in iso-8859-x (and indicate the used charset inside the body itself, how's that for a chicken-and-egg problem). Of course when the data gets truncated due to not actually being utf8, nobody notices. Fun times. [1]: https://dev.mysql.com/doc/refman/5.7/en/sql-mode.html#sql-mode-strict https://dev.mysql.com/doc/refman/5.7/en/sql-mode.html#sql-mo...
- callumjones 9y ago> For f sake how are we supposed to Anderstand that. I suppose your French fry maker is broken ? Didn't expect Chess.com and YouTube to have a crossover of users? Surprised there isn't active moderation on a site this size.
- Zanta 9y agoIn my experience the chat on chess.com harbors a similar demographic to that of most video games. You'd think that chess would attract a more mature player base, but nope.
- marcelluspye 9y agoThe only time I've observed people in real life acting like people in video games is at a chess tournament. Constantly trash-talking until they lose, then accusations of cheating. You certainly don't see this type at all (or most) chess events; I think the lack of entry free drew them out into the open.
- cooper12 9y agoAt my local library the loudest people aren't those on their phones or laptops, but the chess players. It really surprised me considering chess can be played completely non-verbally other than calling "check". Every time I'm there, they constantly argue (most of the time it's because someone wants to take back a move), trash talk to get on each others nerves, yell across the table to other players in games, and talk loudly as if they were in a park. On one hand I think it's great that the library provides a community space and lets people use their chess sets, but on the other hand as someone who goes there for quiet, it's very irritating. (I wish they had a game room or something where they could go wild) Once upon a time libraries had mythical status as a place of silence, to the point where people would shush each other for the smallest noises... I actually stopped going to that library because of noise issues and in general because of its size and limited seating.
- Bakary 9y agoChess played on a computer fits the literal definition of a video game, and one that has near infinite replay value.
- deleted 9y ago[deleted]
- key8700 9y agoeBay (almost) had this problem and I cannot find any articles about it online. They were rapidly approaching 2^31-1 auctions. So they switched to a larger integer, the switchover went badly, and they were mostly down for 4 days, if my memory serves. This would be like 10+ years ago I think.
- mtkd 9y agoThese are always the best problems to have
- _pmf_ 9y agoThat's the most successful reason for failure.
- russellbeattie 9y agoThis problem is more related to a programming underestimation than the actual limitations of a 32bit CPU (which can happily process numbers or IDs that arbitrarily big if you have the memory for it and program it correctly). That said, this is definitely indicative of what's going to happen in just 20 years, 6 months and 20 days from now. I mean, we're still cranking out 32bit CPUs in the billions, running more and more devices, and devs still aren't thinking beyond a few years out. I know of code that I wrote 12 years ago still happily cranking away in production, and there may be some I wrote even longer than that out there... and I guarantee I hadn't given two thoughts about the year 2038 problem back then, and I doubt many devs are giving it much thought today. It's truly going to be chaos.
- protomyth 9y agoThe sad part is people are going to look at the lack of a year 2000 event and assume 2038 is going to be a "dud", when they fail to see all the damn work that went into making sure Y2K was a dud including a significant portion of IT hours and probably a lot of extra support laid in. I expect 2038 to be a rare hell because of the nature of the devices. Y2K was an IT problem, but 2038 will be an embedded system problem and that's going to be a much more painful thing to audit. Moving from the server room to inside equipment and walls is going to be fun.
- isolli 9y agoWhy 2038?
- fsiefken 9y agoWill the Lichess app and platform have this issue? And if not, why not?
- shmageggy 9y agoLooks like lichess is using strings for IDs so they will not have this issue https://github.com/ornicar/lila/blob/master/app/controllers/Game.scala https://github.com/ornicar/lila/blob/master/app/controllers/...
- chesserik 9y agoFun to read some of other stories where this bit them too (PacMan, WoW, and eBay)! Anyway, new app has been approved by Apple and should be rolling out soooooooooon.... Thanks for all the comments! Always lots to learn from.
- spullara 9y agoThe other one to watch out for is the 53-bit javascript integer limit. That caused the twitpocalypse when Twitter tweet IDs hit it. They had to switch to strings in the JSON representation.
- gilgoomesh 9y agoThe 2009 Twitpocalypse concerned overflow of 31-bit precision. Twitter has not yet hit 53-bits for raw number of Tweets, in fact, they passed 32-bits in 2014 and might not have reached 33-bits, yet. Moving to strings for Javascript was really just safety planning for the future since: > Given the current allocation rate, they'll probably never overflow Javascript's precision nor get anywhere near the 64-bit integer space. https://twittercommunity.com/t/discussion-for-moving-to-64-bit-twitter-user-ids/9890/2 https://twittercommunity.com/t/discussion-for-moving-to-64-b...
- spullara 9y agoThat discussion is about user IDs, not tweet IDs. Tweet ID from today: 875423039323688960 Number of bits of precision necessary to represent it exactly: 60 Overflowed 53-bit precision long long ago. You can read about it here: https://dev.twitter.com/overview/api/twitter-ids-json-and-snowflake https://dev.twitter.com/overview/api/twitter-ids-json-and-sn...
- inieves 9y agoThe title is probably wrong, off by one. You probably mean 2^31 -1.
- Piskvorrr 9y agoHelp me, Off-By-One Kenobi! You're one of my two last hopes!
- nicky0 9y ago> an unforeseen bug that was nearly impossible to anticipate Hmmm... :)
- yoz-y 9y agoWhat would be the best way to test for this kind of issues in advance. Testing at theoretical limits at all endpoints?
- damagednoob 9y agoI'm not sure this falls under testing. If you start with an empty database each time you start the test you may never hit this issue. I think this is more of a capacity planning issue.
- yoz-y 9y agoMy concern is that even if one plans for a sufficient capacity, there still needs to be testing done to verify that the code actually works if the capacity is nearing the theoretical limit. In this example the database id was transformed into a 32 bit integer somewhere in the application code. Usually when I hit some sort of unexpected bug in production I try to think about what type of testing will prevent similar problems in the future.
- vitomd 9y agoA lot of comments but no one said the great time that we are living for chess. So many games online, ready to be analysed and learn from them. After deep blue people thought that it was the end of chess, but it´s only getting better. Computers helping players to improve. Chess.com is a great site, also lichess.org and chessable.com if you like chess you should check them.
- cwfrank 9y agoIssues like this are not uncommon on Chess.com. I've been playing there since 2008 or 2009. If you read recent comments about issues as they pertain to the recent "v3" release ... as much is to be expected.