3 ms·
DWARF is terribly designed. It is far more complicated than it needs to be. For example, to map binary locations to line numbers -- it uses its own opcodes and
by Peaker 9y ago
DWARF is terribly designed.
It is far more complicated than it needs to be.
For example, to map binary locations to line numbers -- it uses its own opcodes and registers in a custom VIRTUAL MACHINE(!!)
When it could have had a simple table like:
Line Offset
1 0xa
2 0x14
4 0x2a
...
Then just compress that with an ordinary compression algorithm rather than inventing silly virtual machines.
The debug info itself is encoded in a very cumbersome encoding of debug-information-entries that point to one another.
A much simpler method would have been RLP, for example.
- the_mitsuhiko 9y agoThe line number programs of DWARF are pretty well working for the problems they should solve. Out of all things in DWARF that are weird, these are not the issue. We use DWARF at Sentry just fine and I am quite a big of a supporter of the format as ypu can guess. In particular it was designed and specified unlike sourcemaps which are a random google docs page and don’t even solve basic problems such as finding out which function a token belongs to.
- Jasper_ 9y agoDWARF solves a much harder problem than mapping binary offsets to line numbers. DWARF solves the problem of transforming optimized code into something debuggable. For instance: https://godbolt.org/g/cHvnCh https://godbolt.org/g/cHvnCh This was compiled with debug info turned on. Where did the memcpy line go? Turns out, it can be optimized out entirely and the whole function combined. There's no possible mapping back to source code. DWARF allows you to transform this back to stupider, "unoptimized code", which can be stepped through line-by-line. This is one basic reason why Source Maps are useless: in anything but the stupidest compilers, transformation back to the source code is more complicated than table mapping.
- Peaker 9y agoWhatever information is within the line number information - it can stored in a simple, naive way - and then compressed normally. There is absolutely no justification in inventing a virtual machine with its own opcodes for this.
- the_mitsuhiko 9y ago>Whatever information is within the line number information - it can stored in a simple, naive way - and then compressed normally. Not sure how a compression is going to be better than a VM. The "vm" here is super simple and achieves significantly better compression than an actual compression algorithm. And it's easier to implement and work with. Also again this is not just line information so you really want a state machine for this or this explodes in more and more complexity. We built a system that generates out simple mappings from DWARF's line number programs to files we can mmap and it's only smaller for the case we are about (line number info). Anything else and DWARF's programs are better. So no surprised DWARF works the way it does.
- Peaker 9y agoWhen you just want line number info from DWARF -- all the existing tools are extremely slow. A simple sorted address->line table with binary search is incredibly faster. This is a very common use case. At the very least, this proves DWARF is not designed for its common use cases, not properly at least.
- the_mitsuhiko 9y ago> When you just want line number info from DWARF -- all the existing tools are extremely slow. Sure, but so are sourcemaps. If that is all the info you need then you can build tables for that which is as mentioned precisely what we do. However DWARF is more than that and DWARF is a really good standard for debug information data. You can trivially build cache files for the subset of info you need out of them.