3 ms·
that's a good point... as mentioned in the post we use ssh tunnels and have experienced quite a few glitches, but in this case it has been running without a hic
by solso 14y ago
that's a good point... as mentioned in the post we use ssh tunnels and have experienced quite a few glitches, but in this case it has been running without a hiccup for 3 weeks (proof of nothing of course).
I suspect (guess) that not having an idle connection help the tunnel to not randomly drop.
- semiquaver 14y agoCan you talk more about what the glitches were like? We're a few months away from production with our first product that uses redis, so this is very interesting. Have you simulated a network partition? How does redis handle the reconnect?
- ironchef 14y agoDepending on length of time the network is "out" there's a avariety of things that could occur. In general, the master would kick off a bgsave and the slave would re-bootstrap once it received said snapshot.
- solso 14y agothat's totally correct. In the case of a redis replication, the ssh tunnel failing would be quite costly, since the whole database will have to be send again since the slave would do a SYNC op. And that can be multiple GB in our case. But so far, we haven't experience it.
- solso 14y agoThe glitches with the ssh tunnel were not for the redis replication setup. This has been running fine for 3 weeks already. However we also use ssh tunnels for our continuous integration (jenkins) and the irc bots that we use for development (http://3scale.github.com/2012/06/29/irc-driven-development-part-2/ http://3scale.github.com/2012/06/29/irc-driven-development-p...). In this setup we have the autossh going awol every one or two weeks. But as the blog post mentions, the ssh tunnel runs between our HQ (fiber) and Amazon, not very reliable. Furthermore, there is a lot of idle periods, which seems to trigger most of the issues. We cannot give more specifics since we forcefully just restart the daemon with a monit/munin combo.
- alexfoo 14y agoFirewalls (which may be beyond your control and that you may not even be aware of) between you and the other end often drop connections that have been idle for a while (anything from 5 minutes to a couple of hours), especially if the firewall is busy tracking tens of thousands of active connections. The outbound firewall of our company drops my outbound ssh connections if they are idle for just 5 minutes. Adding:- Host * ServerAliveInterval 60 in my ~/.ssh/config file prevented most of the drops (the rest are explained by proper network outages or my DSL connection at home flapping.) A corresponding ClientAliveInterval setting in the sshd_config file also helps mask the problem [EDIT] if I ssh home from a machine at work that doesn't share my home directory ssh config file.