7 ms·
Rot8000
- egypturnash 8y agoIt meticulously refrains from rotating emoji. Somehow this feels like failure.
- SomeCallMeTim 8y agoROT-8000 is only touching the first 65536 Unicode characters (UCS-2). Unicode has >1M code points. [0] Most emojis seem to be above the first 16 bits. [1] But there are a number of emojis in the first 16 bits, like the "frowning face" emoji at U+2639 -- it rotates just fine -- plus others in the first 16 bits. (TIL you can't paste emojis into HN comment threads. Probably all for the best.) [0] https://en.wikipedia.org/wiki/Unicode https://en.wikipedia.org/wiki/Unicode [1] https://unicode.org/emoji/charts/emoji-list.html https://unicode.org/emoji/charts/emoji-list.html
- have_faith 8y ago> TIL you can't paste emojis into HN comment threads. Probably all for the best. I'm gonna make a HN where you can only speak in Emoji! Sorta unrelated, does anyone remember the social network where you could only write in Emoji? http://emoj.li/ http://emoj.li/
- tyingq 8y agoAre the rules for what it does allow written down somewhere? I know country flags work: 🇩🇪
- codetrotter 8y agoIf your question isn’t answered, you could post all of them in a comment and see which ones remain unfiltered. https://unicode.org/emoji/charts/full-emoji-list.html https://unicode.org/emoji/charts/full-emoji-list.html https://unicode.org/emoji/charts/full-emoji-modifiers.html https://unicode.org/emoji/charts/full-emoji-modifiers.html You might need to split the text over multiple comments. Don’t remember whether or not there is a limit to the length a comment can have. Probably there is.
- tyingq 8y agoI may try that. It seems a little arbitrary. I can post, for example: ↙️ ↩️ ⌚ ⌛ ⌨ ⏏ ⏩ ⏰ ️⏱ ⏲ ⏳ ️◾
- deleted 8y ago[deleted]
- angus-prune 8y agoThere is a wonderful talk by the founders of emoj.li about how it was all a joke which got out of hand. https://www.youtube.com/watch?v=GsyhGHUEt-k https://www.youtube.com/watch?v=GsyhGHUEt-k
- ucosty 8y agoIt seems you can type them in natively, though edit: scratch that, they get stripped out.
- wl 8y ago> TIL you can't paste emojis into HN comment threads. Probably all for the best. It got in the way of me explaining the Rebus principle used in Egyptian hieroglyphs a little while ago. Then again, the topic was phalluses in Unicode (which do display!), so maybe you're right.
- deleted 8y ago[deleted]
- supakeen 8y agoFun, but outputs unprintable or non-used characters and only functions on the BMP?
- jstanley 8y agoInteresting! I made a very similar tool earlier this year. It comes with presets for various different areas of Unicode, and some example text, although the intended use case was very different, I looked at it from a steganography perspective rather than an honours-system obfuscation perspective. https://incoherency.co.uk/mojibake/ https://incoherency.co.uk/mojibake/ I initially thought it would be able to decode the rot8000 output without any modification but I think the utf-8 escaping that my tool expects (from its own output) gets confused by the output from rot8000.
- hyper_reality 8y agoI also had a similar idea a few years ago for a CTF challenge, coming at it from a "modern Caesar cipher" perspective: https://laurencetennant.com/unicode-shift-cipher https://laurencetennant.com/unicode-shift-cipher Also a crude "modern Bacon cipher" using Punycode characters as the B's and ASCII-range characters as the A's.
- brlewis 8y agoIt may also be that you're rotating by 0x8000 and this code is not. It's creating a mapping that's restricted to non-control, non-surrogate, non-whitespace characters and rotating by half the size of that mapping. https://github.com/rottytooth/rot8000/blob/master/Rottytooth.Rot8000/Rotator.cs https://github.com/rottytooth/rot8000/blob/master/Rottytooth...
- danbruc 8y agoThis will break, i.e. two consecutive rotations will no longer be the identity, if the number of valid characters in the BMP ever becomes odd. And there are still a few unallocated code points in the BMP. There is also an overflow in line 39 because of the check i <= BMP_SIZE in line 37 which, I guess, previously used Char.MaxValue instead of BMP_SIZE. But it does no harm here, U+0000 just gets filtered out twice.
- platforms 8y agoThere's a test for BMP characters being even: https://github.com/rottytooth/rot8000/blob/master/Rottytooth.Rot8000.Tests/MappingsTests.cs https://github.com/rottytooth/rot8000/blob/master/Rottytooth... More critically, if the # of valid chars changes, previously rot-8000'd text will no longer be reversible through the tool
- ninjin 8y agoThis certainly is what I would call a “neat hack”. Out of curiosity I had to check what it rotates Japanese into. Turns out, mostly Korean: “日本語はどうかな?” becomes “ື걅갿개갡걀等”.
- theophrastus 8y agoI was curious as to how one might implement this with a familiar language, and fetched up on this interesting python github script, specifically "rot32768"[0] [0] https://gist.github.com/terrorbyte/7967039 https://gist.github.com/terrorbyte/7967039
- tuttle7 8y agoNoone is concerned by the fact this is sending your text using POST requests. The guy could not use DOM/JS.
- stilldavid 8y agoI'd be more concerned if you used this for actual secrets.
- deleted 8y ago[deleted]
- Crespyl 8y agoNo, no one is concerned by this. Not every toy website needs to have JS.
- ravenstine 8y agoNah bro, it needs Webpack and a mishmash of Angular and Vue with a "sprinkling" of React along with an Elixir backend so it's fault-tolerant. Else, how is this toy site supposed to scale at all?
- Sohcahtoa82 8y agoI think the point the tuttle7 was trying to make was that this site could be implemented client-side quite easily. There's no real reason to make the translation server-side and require more server CPU resources and bandwidth. I feel the same way about https://www.base64decode.org/ https://www.base64decode.org/ . By default, everything gets translated server-side. I wonder how many people use this site on a regular basis for translating secrets. I'd bet my life that the number is greater than zero.
- richrichardsson 8y agoThat's why you write rude messages to give him/her a laugh when checking server logs.
- tuttle7 8y ago
- collyw 8y agoCan someone explain what this is doing please?
- joshuamcginnis 8y agoThe only link on the page links to the explanation.
- collyw 8y agoI missed that - meow_info doesn't really convey that its an explanation have the same noticeaion.
- Crespyl 8y agoSee: http://rot8000.com/info http://rot8000.com/info It's essentially a Unicode version of the old "Rot 13" cypher. In Rot 13, you translate each letter 13 places down (as if on a code wheel), such that 'A' becomes 'N', 'B' becomes 'O', wrapping such that 'Z' becomes 'M', and so on. This version, instead of using the simple 'A=1...Z=26' number space, uses the Unicode range and rotates by 32,768 (0x8000).
- scrooched_moose 8y agoOne key aspect you skipped over is it's self-reversible. 'A' becomes 'N', and applying it again 'N' becomes 'A'. "rot13 is reversible" -> "ebg13 vf erirefvoyr" -> "rot13 is reversible". "rot8000 is also reversible" -> "类籸籽籁簹簹簹 籲籼 籪籵籼籸 类籮籿籮类籼籲籫籵籮" -> "rot8000 is also reversible" Rot13 is English-alphabet only so it skips numbers, while rot8000 doesn't have this limitation because it uses the larger unicode set.
- TazeTSchnitzel 8y agoReminds me of the infamous 畂桳栠摩琠敨映捡獴.
- Arkanosis 8y agoFor anyone not getting it: https://en.wikipedia.org/wiki/Bush_hid_the_facts https://en.wikipedia.org/wiki/Bush_hid_the_facts
- loa_in_ 8y agoReminds me of http://base91.sourceforge.net/ http://base91.sourceforge.net/. We could go further, straight to Base8000!
- ar-nelson 8y agoAlready exists: https://github.com/qntm/base65536 https://github.com/qntm/base65536 It's actually pretty useful for compressing data in Unicode-aware environments, like Twitter. Which makes me wonder if Unicode support is universal enough now that an encoding like this could replace MIME/base64 in email.
- lifthrasiir 8y agoOkay, I have seen this 10 times or so when I tried to compare various binary-to-text encodings and basE91 is the only one without a format description. Probably it's time to directly look at the source code. Amazingly, this one turns out to be the only binary-to-text encoding with the input bits groupped by varying number of bits I have ever seen. More specifically: * The input bits are packed in the reverse order (e.g. 1A 2B 3C is packed as 0x3C2B1A) unlike most other binary-to-text encodings. The last bits are padded with preceding zeroes. * A pair of basE91 alphabets encode a number 0 through 8280. The first alphabet is least significant: `AB` encodes 91 and not 1. * 91^2 = 8281 > 2^13 = 8192, so groups of 13 bits are read and encoded as two basE91 alphabets from the least significant to the most significant. But it's not always the case. Occasionally a group of lowermost 14 bits will be read if the bits are less than 91^2. As a result, the first 8281 - 8192 = 89 values (0..88) and the last 89 values (8192..8280) actually encode 14 bits, and it includes all-zero bits. Its average overhead is therefore 22.93% (16 / lg 8281 - 1) and can reach 14.29% (16 / 14 - 1) when all bits are zero. It reminds me of Ascii85 [1] which had a shorthand for all-zero groups and all-space groups, but this one is more general. Speaking of generality, probably a binary-to-text encoding with arithmetic coding is now viable? [1] https://en.wikipedia.org/wiki/Ascii85#btoa_version https://en.wikipedia.org/wiki/Ascii85#btoa_version
- deleted 8y ago[deleted]
- platforms 8y ago籝籱籮 籺籾籲籬籴 籫类籸粀籷 籯籸粁 米籾籶籹籼 籸籿籮类 簹粁籁簹簹簹 籭籸籰籼簷 http://rot8000.com/Index?%E7%B1%9D%E7%B1%B1%E7%B1%AE%20%E7%B1%BA%E7%B1%BE%E7%B1%B2%E7%B1%AC%E7%B1%B4%20%E7%B1%AB%E7%B1%BB%E7%B1%B8%E7%B2%80%E7%B1%B7%20%E7%B1%AF%E7%B1%B8%E7%B2%81%20%E7%B1%B3%E7%B1%BE%E7%B1%B6%E7%B1%B9%E7%B1%BC%20%E7%B1%B8%E7%B1%BF%E7%B1%AE%E7%B1%BB%20%E7%B0%B9%E7%B2%81%E7%B1%81%E7%B0%B9%E7%B0%B9%E7%B0%B9%20%E7%B1%AD%E7%B1%B8%E7%B1%B0%E7%B1%BC%E7%B0%B7 http://rot8000.com/Index?%E7%B1%9D%E7%B1%B1%E7%B1%AE%20%E7%B...
- danbruc 8y agoIn case the author sees this, some comments about Rotator.cs. 1. This algorithm will break if the number of valid characters in the BMP becomes odd. EDIT: As user platforms pointed out, there is an unit test for this. 2. There is an overflow in line 39 because of the check i <= BMP_SIZE in line 37. 3. The web server at rot8000.com exposes at least some errors with stack traces, try rotating the string <script>. 4. In line 42 you are performing a linear search for every character you transform, that is very inefficient, especially with characters at the end of the BMP. At least use a hash map or even better just use an array mapping the input code point directly to the output code point. 5. rot8000.com does at the very least allow rather long inputs which paired with the inefficiency of the linear search makes a DoS attack pretty easy. I tried a 10,000 word lorem ipsum, it was not rejected and the request took a minute to complete.
- rottytooth 8y agoThanks -- added an issue for the linear search https://github.com/rottytooth/rot8000/issues/2 https://github.com/rottytooth/rot8000/issues/2 -- will place a limit on chars in that textbox as well
- andyburke 8y agoLimiting the text box will protect you against the most naive DoS attacks, but you need some kind of limit at the API level (request size, etc.). Never trust the client.
- danbruc 8y agoFor reference, I created an optimized implementation and tested it with a string containing all characters from U+0000 to U+FFFF in order and got the following times. The original implementation took 5202.766 ms, the optimized implementation took 0.079 ms for a speed-up of about 65858. That this is pretty close to 65536 is probably a reflection of the cost for the linear search through almost that number of characters and the test pattern I choose but I am not entirely sure, intuitively I would have expected a factor of 0.5 in there to account for the average case. But I am too lazy right now to do the math.
- deleted 8y ago[deleted]
- dullroar 8y agoShould also change spaces to zero-width spaces, which would then make it less obvious where the word breaks are.
- tsaoyu 8y agoReminds me 锟斤拷 due to Unicode replacement character misinterpretation problem. When placeholder 'U+FFFD' decoded using GBK it will displayed as these characters. Some of glitches can still be found online, e.g., https://docs.oracle.com/cd/E19199-01/817-4244-10/preface.html https://docs.oracle.com/cd/E19199-01/817-4244-10/preface.htm...
- dana321 8y ago籖粂 籶籸籽籱籮类 籪籽籮 粂籸籾类 籬籱籲粀籸粀籸粀籸粀粀
- ConcernedCoder 8y agoFYI: Here's a static JavaScript version I whipped-up ( as a lunch-time challenge ) that will reversable rotate everything except whitespace... https://github.com/jeffallen6767/rot0x8000 https://github.com/jeffallen6767/rot0x8000
- rottytooth 8y agoDon't see how this will work without checking for control characters, surrogates and chars above 0x10000 (try 𝄞 for instance)
- omarforgotpwd 8y agoIf you are just starting to get interested in cryptography, try and make a program that can break ciphers like this one or similar. Hint: Use frequency analysis on sample ciphertext and compare to known letter frequencies in english letter to match to plaintext. Then you can determine the offset and decrypt