4 ms·
If you think it's a hard problem, you aren't thinking about it correctly. When you say "@foo Hi, foo!" You're sending a message to some person named foo. That'
by jamess 18y ago
If you think it's a hard problem, you aren't thinking about it correctly.
When you say "@foo Hi, foo!" You're sending a message to some person named foo. That's a push. If you start thinking of it as a pull, then you get in to the sort of trouble you're mired in.
In a distributed system the process goes like this:
1) Sender queries name server: WHOIS foo?
2) Nameserver responds foo is http://www.foo.com/mytwitterfeed http://www.foo.com/mytwitterfeed.
3) Sender pings recipient http//www.foo.com/mytwitterfeed?ping=messageid&origin=bar
4) Recipient queries name server: WHOIS bar?
5) Name server responds http://bar.com/mytwitterfeed http://bar.com/mytwitterfeed
6) Recipient requests message text, http://bar.com/mytwitterfeed?id=messageid http://bar.com/mytwitterfeed?id=messageid and verifies it actually is addressed to him.
From then on, you have to worry about spam but that's a solved problem (killfiles, Bayesian filtering, etc.) No caching is involved. Recipients store messages addressed to them, just like any other messaging system. This isn't at all a hard problem.
- tptacek 18y agoThis appears to be a system that allows strangers to coerce my client into contacting their server.
- jamess 18y agoWhat's your point? You might as well say HTTP is a method that allows strangers to get files from my server! While technically true, it doesn't say much. Yes, people can send you false pings, which is why you need to go back through the name server and do a reverse lookup, and once you fetch the message you need to parse it and make sure it within bounds (less than 160 characters or whatever) and addressed to you (has @foo in the message text.) and is not spam (the name isn't in your killfile, is in your whitelist, akismet says it isn't spam, etc.) Rate limiting helps too, real people aren't going to send you 100 messages a second.
- Retric 18y agoIMO: They are dealing with a tiny amount of bandwidth say there are 30million 1kb text messages a day that's 900 gb / month which is well easy for a single core CPU to handle. Personally I would go with: User Adam sends message to Bob: 1) Adam sends message to Twitter service. "Bob>Hi." 2) Twitter check to see if Bob knows Adam. 3) Twitter sends "Adam>Hi." to Bob. (With optimal storage so Bob and Adam can see their old conversations) Your system let's random spammers people find out people's address which IMO is bad. PS: Or Adam could send a message to "All>I like this soup." and Twitter then sends the message to everyone that cares about Adam including Bob. Edit: It looks like Twitter is sending around 2million Tweets a day.
- brooksbp 18y agoYou're thinking too qualitatively; imagine the spikes they get... Also most apps are I/O restricted, not CPU... Also, all these proposed messaging architectures are somewhat flawed. Twitter isn't really "sending" messages to other people. It's more like a person's message history is bound to their account, and appropriate privileges are applied. Then, when people try to "read," privileges are obeyed and information is produced... Am I the only one who believes the solution to Twitter does not involve a massive distributed system? Twitter is inherently centralized... people don't see that.
- tptacek 18y agoThe most popular messaging systems are all centralized, too.
- Retric 18y agoOk I can't edit the old post but they are basically just reinventing mailing lists. Anyway, poling is fine if it's initiated by a user action or it's a vary limited number of systems but if you have a few million users you can't uses it per user. As to reading old log's that's a trivially parallizable problem see database replication for some hints. So you separate the active messaging that automatically adds new Tweets to a user’s cell phone from the way users read old messages with the webpage and your home free. PS: Text is cheep EX: Slashdot is running off of ~4 computers. And bandwidth is not a scaling issue until you need a single system to handle more than ~1GB of bandwidth per second. (Saturating an OC-3 line costs a lot of money but you don't need to change your architecture to do so.)
- whatusername 18y agoAren't the scaling issues more in the send-to-all cateogry than in the person-to-person? Aka - a scoble broadcast out to 20,000 different endpoints causes challenges - whether it's pushed or pulled..