7 ms·
There were 84 errors for GCP, but the breakdown says 74 409s and 5 timeouts. Maybe it was 79 409s? Or 10 timeouts? I suspect the 409 conflicts are probably fro
by rwiggins 4y ago
There were 84 errors for GCP, but the breakdown says 74 409s and 5 timeouts. Maybe it was 79 409s? Or 10 timeouts?
I suspect the 409 conflicts are probably from the instance name not being unique in the test. It looks like the instance name used was:
instance_name = f"gpu-test-{int(time())}"
which has a 1-second precision. The test harness appears to do a `sleep(1)` between test creations, but this sort of thing can have weird boundary cases, particularly because (1) it does cleanup after creation, which will have variable latency, (2) `int()` will truncate the fractional part of the second from `time()`, and (3) `time.time()` is not monotonic.
I would not ask the author to spend money to test it again, but I think the 409s would probably disappear if you replaced `int(time())` with `uuid.uuid4()`.
Disclosure: I work at Google - on Google Compute Engine. :-)
- Cthulhu_ 4y agoI've naively used millisecond precision things for a long time - not in anything critical I don't think - but I've only recently come to more of an awareness that a millisecond is a pretty long time. Recent example is that I used a timestamp to version a record in a database, but it's feasible that in a Go application, a record could feasilby be mutated multiple times a millisecond by different users / processes / requests. Unfortunately, millisecond-precise timestamps proved to be a bit tricky in combination with sqlite.
- account42 4y agoTo put this into perspective, a game (or anything) running at 60 FPS only as a bit over 16 milliseconds to render each frame. These days higher frame rates are common enough, often putting you down into single-digit milliseconds per frame. Not think about how many things there are simulated and rendered in each frame for common games. Way more than 16.
- sitharus 4y agoThis is a very good point - AWS uses tags to give instances a friendly name, so the name does not have to be unique. The same logic would not fail on AWS.
- smashed 4y agoWhich makes 2000% sense. Why would any tenant supplied data affect anything whatsoever? As a tenant, unless you are clashing with another resource under your own name, I don't see the point of failing. aws S3 would be an exception, where they make that limitation on globally unique bucket name very clear.
- deleted 4y ago[deleted]
- vineyardmike 4y ago> unless you are clashing with another resource under your own name, I don't see the point of failing. Is that not the conclusion? The tester was clashing with their own names?
- jsolson 4y agoIdempotency. You're inserting a VM with a specific name. If you try to create the same resource twice, the GCE control plane reports that as a conflict. What they're doing here would be roughly equivalent to supplying the time to the AWS RunInstances API as an idempotency token. (I work on GCE, and asked an industry friend at AWS about how they guarantee idempotency for RunInstances).
- bushbaba 4y agoDo you really need idempotency for runVM though.
- jsolson 4y agoI mean, it's kinda nice to know that if you reissue a request for an instance that could costs thousands of dollars per month due to a network glitch that you won't accidentally create two of them? More practically, though, the instance name here is literally the name of the instance as it appears in the RESTful URL used for future queries about it. The 409 here is rejecting an attempt to create the same explicitly named resource twice.
- mempko 4y agoWhat are your thoughts on the generally slower launch times with a huge variance on GCP?
- kevincox 4y agoFWIW in our use case of non-GPU instances they launched way faster and more consistently on GCP than AWS. So I guess it is complicated and may depend on exactly what instance you are launching.
- Crash0v3rid3 4y agoThe author failed to mention which regions these tests were run. GPU availability can vary depending on the regions that were tested for both Cloud providers.
- fomine3 4y agoThis is a big missed point.
- ayewo 4y agoThe author linked to the code at the end of the post. The regions used are "us-east-1" for AWS [1] and "us-central1-b" for GCP [2]. 1: https://github.com/piercefreeman/cloud-gpu-reliability/blob/4cfd6817c9233e1c01482f6233d39fa8529dc7f1/gpu_reliability/cli.py#L39 https://github.com/piercefreeman/cloud-gpu-reliability/blob/... 2: https://github.com/piercefreeman/cloud-gpu-reliability/blob/4cfd6817c9233e1c01482f6233d39fa8529dc7f1/gpu_reliability/cli.py#L36 https://github.com/piercefreeman/cloud-gpu-reliability/blob/...
- valleyjo 4y agoJust remember this is for GPU instances. Other vm families are pretty fast to launch.
- jhugo 4y agoAt work we run some (non-GPU) instances in every AWS region, and there's pretty big variability over time and region for on-demand launch time. I'd expect it might be even higher for GPU instances. I suspect that a more rigorous investigation might find there isn't quite as big a difference overall as this article suggests.
- okdood64 4y agoHope icyfox can try running this with a fix.
- brianpan 4y agoTime is difficult. Reminds of me this post on mtime which recently resurfaced on HN: https://apenwarr.ca/log/20181113 https://apenwarr.ca/log/20181113
- jrumbut 4y ago> (3) `time.time()` is not monotonic. I just winced in pain thinking of the ways that can bite you. I guess in a cloud/virtualized environment with many short lived instances it isn't even that obscure an issue to run into. A nice discussion on Stack Overflow: https://stackoverflow.com/questions/64497035/is-time-from-time-h-monotonic https://stackoverflow.com/questions/64497035/is-time-from-ti...
- hansvm 4y agoA lot of the places you use time.time and would be bitten by non-monotonicity you probably want something like time.perf_counter, which is useless for measuring absolute time but perfect for calculating time elapsed.
- rkangel 4y agoYes. When people write `time.time()` they almost always actually want `time.monotonic()`.
- flutas 4y ago> I just winced in pain thinking of the ways that can bite you. Something similar caused my favorite bug so far to track down. We were seeing odd spikes in our video playback analytics of some devices watching multiple years worth of video in < 1 hour. System.currenTimeMillis() in Java isn't monotonic either is my short answer for what was causing it. Tracking down _what_ was causing it was even more fun though. Devices (phones) were updating their system time from the network and jumping between timezones.
- jrumbut 4y agoThat's a bad day at the office when you have to go and say "hey remember all that data we painstakingly collected and maybe even billed clients for?"
- flutas 4y agoLuckily I worked for the company that made the analytics tools and consumed them! Bosses actually came to us because our analytics team was trying to figure out who was causing it, because it had been caught by the team doing checks against the data. (a playback period should never have had > 30s of time)