11 ms·
I've been having major internet issues lately (Seattle area), have had 4 techs come try to figure it out. Yesterday's tech finally correctly diagnosed the probl
by rococode 7y ago
I've been having major internet issues lately (Seattle area), have had 4 techs come try to figure it out. Yesterday's tech finally correctly diagnosed the problem as happening before the connection reaches our home but was unsure of the cause. He called his supervisor to investigate, and they found that the capacity for our neighborhood's node was nearly at 100%, while ideally it should always be under 80%. Fortunately they said they'll be able to fix it within a few weeks by doing a node split. The tech mentioned he'd never heard of capacity issues before in his ~20 years as a tech and that some smaller ISPs have been having issues keeping their internet up and running at all.
I've been tracking the performance with PingPlotter, if you're curious how bad it is right now here's the last 10 minutes: https://i.imgur.com/AnUqv3j.png https://i.imgur.com/AnUqv3j.png (red lines are packet loss) Pretty interesting how current circumstances are pushing even tried and tested infrastructure to their limits.
- martin_bech 7y agoI’ve worked at a major ISP, for a decade, and spotting something like this should be so easy to spot. There are tools on monitoring of load all the time, and areas are routibely getting split etc. to improve bandwith, so I think your ISP are basicly amateurs..
- cpitman 7y agoAlternatively, load has gone up across the board in a short period of time, so that preventive scaling has fallen behind and are in recovery mode.
- martin_bech 7y agoYes it can, but why would it take several techs, to spot something like load, which is the first thing you would do, it should take no more than 10s to look it up in a tool.
- sulam 7y agoA "last foot" tech might not even have access to those tools, much less know how to use them.
- rwbhn 7y agoRolling out that tech has got to be more expensive than checking the load first.
- toyg 7y agoDunno how it is in the States, but here in UK rolling out the tech is basically the first thing they do after the unavoidable "have you tried turning it on and off again" phone call. They just don't trust the customer to have any clue and maybe don't want to waste time doing troubleshooting at their end when it's "probably" a downstream issue.
- TheSpiceIsLife 7y agoNetwork Operations should be raising known problem issues to front line call centre staff. Network congestion issues shouldn't be handed off to field techs to check local loop (last mile) and CPE (Customer Premises Equipment.
- wtallis 7y agoI'm pretty sure it's standard practice at these companies to never let front line call center staff acknowledge known problems. Sometimes, the automated phone menu will give you a recorded generic message that they are currently experiencing a service issue, but that's intended to convince you to hang up and patiently wait for them to sort their shit out. I've never had a front-line rep be at all useful in diagnosing a real problem.
- TheSpiceIsLife 7y agoYeah true. I guess I need to remember the ISP I work for here in Australia (front line tech support, and then network operations physical security and infrastructure) was widely recognised as the best ISP in Australia multiple years running, so I shouldn't use it as a baseline expectation.
- op00to 7y agoNot amateurs, liars.
- dawnerd 7y agoSo Frontier?
- fibers 7y agohow can i as a subscriber find out whats the capacity?
- lxgr 7y agoIt used to be possible to determine the downlink capacity and even current usage with a DVB-C receiver and some Linux software, since DOCSIS is essentially just IP encapsulated in MPEG transport streams on a digital TV channel. More recent versions of DOCSIS have moved away from that layer of backwards compatibility, so you would probably need some specialised equipment, if it is possible at all (I don't know at what layer exactly encryption happens).
- kitteh 7y agoThe problem is that most companies aren't going to tell you that their peering circuits are running hot or that their internal network or access layers to the end user are running warm at peak. ISPs all do stat muxing and the line is "we make money when customers don't use the service". They'll be happy to deal with the last mile segment, but anything beyond that is murky and most companies I know aren't going to share much. Helps to have friends on the inside leak some graphs, though.
- notyourday 7y ago> I’ve worked at a major ISP, for a decade, and spotting something like this should be so easy to spot MRTG graph, ISP circa 1995. Colorized. See a flat line? that's congestion. Now figure out where it is coming from. Sorry, we have been doing this for thirty years so I'm kind of cranky. It is not a rocket science.
- the8472 7y ago> There are tools on monitoring of load all the time On some days my connection resets 5 times within an hour, which is quite annoying since retraining the connection takes a minute or two. When I call support about it they have zero monitoring in place that would let them know about the recent history of the connection quality, they can only do spot tests of SNR on demand, which of course doesn't show any transient events. According support forum posts of other users they'd have to explicitly enable "long term monitoring" based on user input to get that information. Of course SNR line quality is an issue separate from congestion, but still, automatic monitoring appears to be limited.
- texthompson 7y agoIf you didn't know, that 80% number is probably the result of Little's Law. That's the result where if your demand is generated by a Poisson process, and your service has a queue, 80% utilization of the service is where the probability of an infinite queue starts to get really high. People Here's a nice blog post about the subject: https://www.johndcook.com/blog/2009/01/30/server-utilization-joel-on-queuing/ https://www.johndcook.com/blog/2009/01/30/server-utilization...
- lxgr 7y agoThis law does not apply to queueing as encountered in routers. It assumes unbounded queues and a poisson arrival process (i.e. a memoryless channel); both assumptions don't hold for packet routers and senders using congestion control (TCP or otherwise). There is, however, a high chance of encountering buffer bloat if countermeasures are not taken at the chokepoint: https://en.wikipedia.org/wiki/Bufferbloat https://en.wikipedia.org/wiki/Bufferbloat Modern cable modems, for example, are required to implement such countermeasures. My ISP is at over 90% capacity and round trip times are still mostly reasonable. (Bandwidth is atrocious, of course.)
- drewbailey 7y agoHow do you monitor this? The 90% over capacity, would like to see where mine is at
- lxgr 7y agoThere might be a way using a cable TV receiver (see my other comment on this thread), but in my case, a sales rep of my ISP just told me on the phone.
- mikepurvis 7y agoI have an older modem (DCM476) and it definitely doesn't have this or doesn't have it enabled. I have to use/tune queue management myself on the router side.
- kitteh 7y agoComcast's last mile network in Seattle has been struggling in some areas from the morning until around 4 to 5 PM. It's not massive loss, but enough to disrupt video conference. Run a mtr towards an Internet dest and you'll see loss at the first hop and everything behind it.
- op00to 7y agoMtr isn’t a reliable measure of packet loss. Routers drop “extra” packets like ping before they drop “paying” packets.
- lxgr 7y agomtr uses UDP data packets, as far as I am aware. Yes, the ICMP response packets could still be skewed, and the effect you mention is definitely real, but on a good connection, usually there should not be much to drop at all, neither TCP/UDP traffic nor ICMP packets.
- cthalupa 7y ago>mtr uses UDP data packets, as far as I am aware. Doesn't matter what it uses (though by default MTR does use regular old ICMP Echo - you have to specify -u or -t to get it in UDP or TCP mode). When TTL expires it still requires an ICMP TTL Exceeded be sent, regardless of whether or not you were sending ICMP through it. Traceroute implementations in general are probably telling most everyone in this thread a lot less than they think, even without icmp deprioritization being taken into account. https://archive.nanog.org/meetings/nanog47/presentations/Sunday/RAS_Traceroute_N47_Sun.pdf https://archive.nanog.org/meetings/nanog47/presentations/Sun... is worth a read for most anyone that's ever attempted to use traceroute to troubleshoot networking, because they're almost certainly doing it wrong.
- kitteh 7y agoYes I'm well aware of routers policing TTL=1 packets, but if you see consistent loss all the way down it's usually a sign. This compared to seeing individual spikes on intermediate routers which are usually control plane policing.
- op00to 7y agoThe tech was full of shit. This happens literally all the time. You probably won’t get a “node split” unless more people loudly complain. It’s cheaper for them to roll a tech and hope you get fed up than it is to actually fix the problem.
- lxgr 7y agoMy ISP has been playing the same game with me for months. I finally cancelled the contract when it was about to renew, and I got a very interesting winback call from sales: Not only did the rep freely share the utilization numbers with me (80% during the day and 90% at night), he also mentioned that things would not get better until end of the year when they would do a node split. As consolation, they offered me 10x the download speed for half the price. I'm not really sure how that would help congestion...
- iagovar 7y agoI work in this field in Spain. Margins in this sector are slim, deployment is expensive. EVERYONE works with simultaneity rates, it's the only way to have cheap connections. In fiber connections is actually not that expensive to split a fiber after a CTO, you can actually sort of daisy chain it, but you want to keep everything as standard as possible.
- VHRanger 7y agoMargins are not slim at all in the USA
- ThePowerOfFuet 7y agoYou think they're fat in the US? Look north.
- richjdsmith 7y agoShh, you'll upset the Great Robelus[1] and they may start euthanizing animals.... [1] https://www.thebeaverton.com/2020/03/telus-threatens-to-euthanize-animals-if-crtc-approves-mobile-virtual-network-operators/ https://www.thebeaverton.com/2020/03/telus-threatens-to-euth...
- chrisseaton 7y ago> have had 4 techs come try to figure it out Doesn't sound ideal for distancing.
- martyvis 7y agoThe way the story is written it sounds like their attendance was there was serial temporal distance involved (they didn't come at the same time)
- chrisseaton 7y agoI mean inviting four different people into your home sounds silly if they're there at the same time or not! I guess people need internet to earn a living though.
- TheSpiceIsLife 7y agoHow does it come about that the ISPs Network Operations team didn't know they were saturating a link? Last ISP I worked at would have email and SMS notifications going to On Call staff.
- kitteh 7y agoBecause the NOCs may not be all that competent. I remember talking to the Cablevision IP NOC back in the mid 2000s about their internal backbone circuits they were running hot that went to a POP we peered with them. I had Cablevision at home and the congestion was breaking my VPN to work. The NOC said "an OC45 was down" (no such thing, it's an OC48) and that congestion is okay because TCP will work with it okay and there won't be a problem. I shutdown the peering session with them force traffic around a diff city (sent it to Chicago). I remember talking to the eng team at Cablevision about their NOC and they had a good chuckle and admitted they're only good for the simplest of operations (link down, go fix). In some parts of the world running links at 95 percent is okay because look 5 percent left (totally ignorant of buffers or microbursts etc.).
- 1_player 7y agoA free alternative to PingPlotter: https://www.thinkbroadband.com/broadband/monitoring/quality https://www.thinkbroadband.com/broadband/monitoring/quality My connection: https://www.thinkbroadband.com/broadband/monitoring/quality/share/696c05a52149ab5ae899496280fae698730cdcc2-19-03-2020.png https://www.thinkbroadband.com/broadband/monitoring/quality/... In case anyone is shopping for broadband in the UK, I only have great things to say about Zen pictured above. It's so good I just called to upgrade my 80 Mb to a 300 Mb just for fun, meanwhile my quarantined Italian friends are suffering awful internet now that everybody's at home streaming Netflix. I used to have Virgin fibre and my average ping was 80ms with a ton of jitter. The plot above is my internet while downloading at about 2MB/s average over the past 24 hours, and surprisingly stays the same even at peak download.
- wh1t3n01s3 7y agoNot friend of yours italian quarantined enjoying 1Gb/s here. Never used Netflix ;-)
- collinmanderson 7y agoThat's a neat alternative to PingPlotter. I like that it pings from outside, so no client required. I'll check it out, however, I'm in the US, so I bet it's always going to be high latency.
- nvarsj 7y agoI’m being pedantic, but that’s not really Zen, it’s the BT Openreach backend which has really great stability and latencies. I tracked my BT Openreach connection for many years and I never got more than a few ms of jitter, really amazing. However the speeds are not great (70/20), and the coverage is also fairly poor - I'm in a dead zone right now between two local exchanges. So unfortunately I'm forced to use Virgin, which has gotta be the worst ISP in the history of the world (and I have had Comcast!). Terrible network and terrible customer service - I don't know how this company exists.
- mulmen 7y agoWhat ISP? I’m on Comcast “Business Class” in Seattle and experiencing occasional slowdowns as well.
- dottenad 7y agoSame
- solsticedev 7y agoCurious, what ISP do you have? Currently moving to a new place in Seattle and have to decide between Wave G or Atlas Networks.
- rkeene2 7y agoFWIW, the reason nodes typically don't get to 100% is due to something called WRED (Weighted Random Early Detection). As the outbound/inbound queue on your "node" approaches fullness, it randomly selects packets to drop. This signals TCP on the sender to back-off. The closer to full-ness it gets, the higher the probability (weight), so the sender knows to slow down to the slowest link's speed. I've written more about this problem here [0]. [0] https://rkeene.org/projects/info/wiki/176 https://rkeene.org/projects/info/wiki/176
- eru 7y agoThanks for the write-up! I wonder how TCP BBR would react here. If I understand it right, it wouldn't need RED to back off: the increased latency of buffers filling up would do that automatically. But BBR also wouldn't let the occasional dropped packet make it back off.
- rkeene2 7y agoFrom what I understand about TCP BBR from reading about it the past few minutes, it would compute a new link speed as a result of impacts from WRED and then use that for the connection baseline speed. TCP BBR would still rely on RED/WRED to compute the connection rate estimate initially, then it would attempt to send below that rate to avoid packet loss. If packet loss is detected it would recompute the estimated connection rate. I found this page [0] useful, especially the graphs. [0] https://blog.apnic.net/2017/05/09/bbr-new-kid-tcp-block/ https://blog.apnic.net/2017/05/09/bbr-new-kid-tcp-block/
- cmauniada 7y agoHo lee sh, that is absolutely crazy. I am sure its affecting you internet speed, what sorts of tasks are you generally doing now that the entire is state is pretty much on lockdown? Here in Alberta, although we are told be socially distant, there is no full lockdown and I want to know what kind of issues would I be expecting to run into in the up coming weeks/months?
- jtokoph 7y agoThis happened to me years ago near the University of Illinois campus (UIUC) with Comcast. I had multiple techs come out but they would only come in the morning when the connection was fine. I finally escalated to corporate who finally told me they needed a node split. I made them give me 100% free internet until the split was complete about 6 months later.
- illiilliiililil 7y agoWhen I have connectivity issues during a pandemic I make sure at LEAST 6 techs come to make sure I have perfect connectivity to Netflix and chill.
- geniium 7y agoThanks for mentioning PingPlotter, I'll try it out to monitor our connexion.
- hendry 7y agoPing plotter looks like a SaaS https://hub.docker.com/r/linuxserver/smokeping/ https://hub.docker.com/r/linuxserver/smokeping/
- teekert 7y agoSince I have been at home I practically live in MS Teams, with constant video chats. Yesterday I did a presentation with 140 people connecting watching my ppt and camera. That's got to be unusual. I imagine most of my colleagues going through this routine daily.
- C1sc0cat 7y agoMakes me glad I went with the Business version of Vodaphone in the UK - which is ironically £1 cheaper a month than the consumer. I suspect its the services that relay on super low prices and don't have excess capacity Talk Talk etc that are really going to feel the pressure in the UK
- abjKT26nO8 7y agoYou're describing an issue specific to US ISPs. It doesn't apply to Europe. From what I read even before the pandemic the US ISPs offered rather crappy services. In Europe, particularly in Poland, I don't have and haven't heard about anyone having any issues with connectivity right now, even though the country is in lockdown, schools and universities are closed, restaurants work only in delivery/take-out mode, companies switched to remote work, ... And still no issues at home nor at work. Don't make decisions about the European infrastructure based on American problems.
- pawelk 7y agoI'm in Poland as well and I've been working remotely for over three years. Since the lock down started I feel that everything is a bit slower and less stable, but I haven't experienced major issues during usual work hours doing work-related things (maybe except MS Teams acting up). However Netflix is broken most of the time during afternoon hours (when I want to keep kids occupied with cartoons for an hour or so to get things done). Luckily other streaming services work fine.
- terramex 7y agoIn contrast, my internet connection finally started working great since lockdowns started. I suspect my ISP (small local company in central Poland) got some additional bandwidth or somehow finally fixed their infrastructure when they saw increased internet usage among their clients.
- abraae 7y agoI work virtually from New Zealand with my colleague in Lombardy Italy. Today I noticed some more serious degradation in video call quality for the first time. But mostly I'm amazed how well the internet is working given the circumstances.
- nicoburns 7y agoHaving issues with the internet here in the UK today. Unsurprising given that half of the world has suddenly discovered video calling. Mobile network seems more stable.
- the8472 7y ago> I've been tracking the performance with PingPlotter, if you're curious how bad it is right now here's the last 10 minutes: https://i.imgur.com/AnUqv3j.png https://i.imgur.com/AnUqv3j.png Is your own connection idle though? Pings are also affected by the congestion on your own router†, especially if you don't have good AQM (such as CAKE). Dumb queues will just drop all packets equally, smart queues will do flow isolation and penalize the bulk flows first while keeping the trickle ones (ping, ssh, voip, ...) untouched. † and anything else along the path to your ping target
- catalogia 7y agoAround the time you posted this, my internet in Seattle was down for near around 12 hours yesterday. I'm not fond of my ISP, but that's unusual even for them.