4 ms·
Recently I began reverse engineering the game files for Circuit's Edge (1989). Most of the in-game text is stored in a separate file called DISKTEXT.TXT. The me
by coderdude 15y ago
Recently I began reverse engineering the game files for Circuit's Edge (1989). Most of the in-game text is stored in a separate file called DISKTEXT.TXT. The method Westwood Associates used seems to be a mix of encryption and compression, and looks very similar to the kind of encoding seen in this article.
The file basically uses a number of bytes between 128-254 to represent two-letter combinations. After an hour of tweaking a Python script I finally had the file decrypted. As someone who had never before dabbled in this sort of thing I felt very accomplished, although I quickly realized how rudimentary their methods were.
- jgrahamc 15y agoHow did you come to realize that you were dealing with bigrams?
- emillon 15y agoI did something similar with SNES roms. Sometimes the in-game text is encoded using "DTE compression", a mapping from one hex character to an alphanumeric sequence (bounded Huffman coding, basically). On certain sentences you can notice that gaps are smaller (in terms of bytes) than what should fit in there ; so you can deduce that it was a bigram (or trigram, or more).
- coderdude 15y agoTrial and error for the most part. First I tried to find sequences of characters that, without modification, already looked like real words. I found one in particular that turned out to be the name of the city in which the game takes place ("Budayeen", though it was completely garbled). That was a stroke of luck because that word appeared many times in the text and gave me some clue about the adjacent words, since there are only so many words you could reasonably put around the name of the city. I tried globally replacing the characters in DISKTEXT.TXT but was ending up with a word half as long as I thought it should be (for "Budayeen"). One of the fortunate things that happened though was that it revealed to me, by accident, a couple other words (even though the decryption wasn't correct, it made some previously indistinguishable sequences look more like real words, which I pursued). I think it really clicked that I was dealing with bigram substitution when I had to come up with a theory of how whitespace was so cleverly hidden. Here's the lookup dictionary if anyone is interested: http://pastebin.com/pxJU7q7F http://pastebin.com/pxJU7q7F And here's the script that does the substitution: http://pastebin.com/8eYrGbzQ http://pastebin.com/8eYrGbzQ You know, now that I actually think back to it, the lower 4 bits represent one character and the higher 4 bits represent the other. And there I was thinking in bytes the whole time.