4 ms·
I wonder if this inspired the VMWare VM record-and-replay functionality that came out in 2008. They discontinued it in 2011, but it's important to me because we
by roca 2y ago
I wonder if this inspired the VMWare VM record-and-replay functionality that came out in 2008. They discontinued it in 2011, but it's important to me because we used it at Mozilla to great effect and that made it easier for me to get Mozilla to support the development of rr, which started in 2011.
- icholy 2y agoI don't get how rr isn't more popular.
- soamv 2y agoIt might have, I remember attending a talk by Peter Chen when I was at VMware around that time, and I know there was some kind of collaboration. (I wasn't involved in record-replay at VMware, but I was interested in it due to some ancient work I did with userspace record-replay debugging, https://lizard.sf.net https://lizard.sf.net). rr is fantastic work, mad props! And the multiprocess stuff in pernosco looks super neat.
- purple-dragon 2y agoYes. There’s a whole family of papers by Chen, his students, and other VMware folks in this vein.
- gwd 2y agoYes, at some point when I was a student I went out to VMWare and gave a presentation to a team of their engineers. The thing about execution replay is that, in the general case, it's incredibly fragile. I happened to chat to some people from VMWare in 2008 and they said that there was a team of 10 people whose only job it was to fix execution replay when it broke. The main thing I learned from my PhD on execution replay was how to debug insane bugs. Presumably rr, focusing on processes, offers a more constrained environment that's easier to log & replay.
- mark_undoio 2y ago> Presumably rr, focusing on processes, offers a more constrained environment that's easier to log & replay. As someone who's worked elsewhere on time travel debug, I'm really curious on @roca's take on this - because I'd have expected a full-VM solution to be easier to make reliable. Hardware-level behaviour sounds harder but it's well-constrained. The behaviour an OS can rely on from the hardware is large but well-documented and slow to change. In contrast, the process boundary is really ill-defined and permeable. Also, when you need things to be precisely the same, you notice bits of kernel behaviour leaking into the user space ABI in unexpected ways.
- roca 2y agoMy instinct is the same as yours, but I'm also well aware of the "grass is greener on the other side of the fence" cognitive bias :-). So without actually having written a record-and-replay hypervisor, I hesitate to judge.
- Veserv 2y agoIf you are doing a single core virtual machine, the hypervisor is properly implemented, and you can change the hypervisor then it is fairly straightforward. You just funnel everything through a single guest memory write function and then instrument that. Unfortunately, virtual machines are basically only ever used in multicore configurations which are fundamentally unconstrained multithreading which makes record-replay based (versus full trace based) implementations impossible (or at least very challenging) unless you serialize. The device driver interface is most likely going to be trapping accesses and communicating back via asynchronous shared memory writes if not pass-through or paravirtualized. I suspect most hypervisor implementations will not even allow you to intercept/trace those operations externally, and if they do it is likely going to be hard to sequence them relative to VM execution/instruction stream as they are rooted in a shared memory interface. If they are paravirtualized, or just in general, the hypercall interface is also usually very ad-hoc and ill-defined and also likely not possible to intercept or trace. The scope is probably going to be more like trying to represent the effects of every random vendor-specific ioctl.
- 2y ago
- roca 2y agoThe interface between userspace processes and the kernel and CPU is pretty complicated. It's hard for me to tell whether it's more or less complicated than the interface between guests and the hypervisor (given most record and replay scenarios don't require virtualizing exotic hardware). It probably is easier to debug rr replay bugs than VM-level replay bugs, partly because rr is lighter-weight. rr replay bugs are still insane, but it's all relative :-).
- gwd 2y agoMy first prototype of execution replay I implemented was for Linux processes (in like, 2000 or something), which I did by modifying Linux. The primary thing I did there was to instrument copy_to_user(). The final version of execution replay that I did was for Xen, which had the asynchronous device drivers. In that case, the goal of the drivers was to have the device do DMA directly into the guest's memory space, and thus had to be modeled as an external processor. If you were content to modify your virtual devices to always doing copying, it would be significantly easier. But yeah, in both cases you just have so much that the thing underneath might be doing, it's hard to tell which is going to be more fragile. :-) The final version also did multi-processor VMs, using the paging hardware to track sharing across virtual cpus [1]. The key take-away from that was, "It's slow, but often not as slow as you might think." [1] https://www.semanticscholar.org/paper/Execution-replay-of-multiprocessor-virtual-machines-Dunlap-Lucchetti/b83955e88528cce2b6715e59c7e6f690b8e6b73e https://www.semanticscholar.org/paper/Execution-replay-of-mu...