4 ms·
Why PostScript had to be Turing-complete makes no sense to me. Loops, code-execution, functions, it all seems so unnecessary for a markup and presentation langu
by nness 3y ago
Why PostScript had to be Turing-complete makes no sense to me. Loops, code-execution, functions, it all seems so unnecessary for a markup and presentation language.
- mananaysiempre 3y agoTuring completeness isn’t really a huge problem (and not included in that part of PDF anyway). PostScript in its printer application doesn’t really have I/O except for, well, the printer, and no raw pointers, so there’s a smaller surface than, say, JavaScript (no network access).
- derefr 3y agoOn its own, the Turing-completeness of PostScript is not a vulnerability, per se. But it does mean that anything that interprets or renders PostScript might take an unbounded (or infinite) amount of time to do so. So PostScript-handling software — even things like virus scanners — have to be coded to handle timing out and interrupting/cancelling PostScript interpretation. (Where PostScript/PDF rendering is often the only format among many "document formats" or "image formats" handled by such software, that imposes this requirement — so when multi-format "reader" software first adds PostScript- or PDF-handling capability, it often finds itself suddenly needing to rearchitect from using single-threaded bounded-time blocking calls to parsers, to needing to use async, concurrency, timers, etc.) That being said, in combination with exploitable vulnerabilities in particular PostScript parsers used in software like Acrobat, the Turing-completeness of PostScript makes it much harder to detect such exploits. 0days in PDF readers are "nice exploits to have", because a PDF virus can be coded such that it's represented in code inside the PDF in an unbounded number of different ways, foiling signature-based virus sanners.
- anthk 3y agoThere's a ZMachine interpreter written in postscript: gs -dNOSAFER zmachine.ps -- yourgame.z3
- bee_rider 3y agoIt was the early 80’s, probably the flexibility seemed nice, and the riskiness of putting a full programming language in your documents was probably not as obvious. Thankfully we learned from them and didn’t repeat that mistake over and over again.
- actionfromafar 3y agoWASM?
- dotancohen 3y agoYes, and JavaScript before that. His last line was sarcasm.
- linguae 3y agoPostScript does more than markup and presentation; it is an entire 2D graphics engine. At one point PostScript served as the graphics substrate for the Sun NeWS and NeXT Display PostScript-based window systems. While I agree from a security standpoint that its Turing-completeness poses challenges, it also makes it easier to express certain constructs such as complex shapes and fonts programmatically. I’m not a PostScript expert but I’ve been reading a lot about it recently. It’s a rather fascinating system for 2D graphics.
- layer8 3y agoPostScript is deprecated in PDF 2.0 and is not the source of the issues listed in TFA.
- derefr 3y agoWell, if you think it's possible, try coming up with even the core of an architectural basis for 1. a declarative language for describing the same things PostScript describes, 2. which allows the rasterization of arbitrary shapes at arbitrary DPI (PostScript is DPI-oblivious — it's up to the printer what DPI it's printing at!); 3. and which also works for vector plotters, that will never rasterize the data you're sending at all, but will actually follow the bezier curves, like the 2D version of 3D-printer GCODE; 4. and which enables the implementation of this rasterization and/or plotting on a variety of affordable hardware architectures in the 1980s — where the 1980s was a time where CPU power wasn't too expensive, but where memory prices were at an absolute premium. So your printer might have had a CPU as powerful as your computer's in it, to crunch PostScript — but definitely wouldn't have had the memory to buffer a full rasterized page. --- To put that last constraint another way: PostScript was designed to be rasterized in a way that enabled printers to do something much akin to "Racing the Beam" (https://www.youtube.com/watch?v=sJFnWZH5FXc https://www.youtube.com/watch?v=sJFnWZH5FXc). In both the display-rendering and printing cases, this was done in 1980s hardware, because 1980s memory was too expensive for most systems to be dedicating it to hold a buffer to asynchronously pre-render into and then read from when drawing. So instead, in both cases, you must render+rasterize in one motion, programmatically and extremely efficiently. And the obvious way to do this, is by using a CPU with rasterization MMIO registers it can very quickly poke at, to change mode bits during the rendering process. A CPU whose ISA becomes, in effect, a Domain Specific Bytecode for procedurally generating raster-lines. If printer vendors of the 1980s could have been expected to agree on a single such ISA, then chip vendors would have just made printer SoCs that conform to that ISA — and we'd have ended up with some kind of "vector-drawing abstract machine" bytecode (with real hardware impls in printer ASICs, but also virtual ones on PCs) rather than PostScript. But as with RDBMS vendors in the 1980s, the printer vendors were all too invested in their own internal architectures to agree on what the low-level execution plan should look like. And a with RDBMS vendors, the solution that all these 1980s manufacturers could get behind, was a standard for a (theoretically) portsable, text-based intermediate-language standard — one that could be generated by computer software, transmitted to their system, and then, inside their system, compiled down to whatever internal representation allows the system to do things its own way. In RDBMSes, this "intermediate language" was SQL. In printers: PostScript.