10 ms·
TIL you can de-obfuscate code with ChatGPT
- Double_a_92 4y agoMore like: "ChatGPT can pretend to deobfuscate the code"
- yodon 4y agoAre you disagreeing with the deobfuscation that ChatGPT produced?
- KomoD 4y agoIt's just guessing and happened to be right once
- o_m 4y agoYou can double check the answer and still save a ton of time.
- throwawayway99 4y agoIf you're double checking every line, how much time are you saving?
- shever73 4y agoAccording to Alex's own tweet, ChatGPT originally produced the obfuscated code too. In my head, this is like asking someone to translate "Hello" into French and then asking them to translate "Bonjour" back to English. It proves nothing about capabilities or usefulness.
- twic 4y agoThis is the correct response. Especially given that the "obfuscated code" is not syntactically valid. Even if you repair the syntax, it contains a number of errors, and eventually descends into gibberish (although some of what it does is not bad!). There is nothing to deobfuscate here.
- zekica 4y agoI'm disagreeing that the obfuscated code even works. You shoudn't be able to deobfuscate code that doesn't work and get a resulting code that works.
- cinntaile 4y agoThat's amazing, I'll have to try this. Did you verify that it's correct?
- crizzlenizzle 4y agoIs it correct though? I’ve been toying around with ChatGPT for a few weeks now and I encountered a few situations in which ChatGPT was like 90% accurate at best. Things like suggesting snippets of configuration files or plugin research. It’s good to get an idea and get started somewhere, but I certainly cannot trust it blindly.
- Jeema101 4y agoDon't think so. There's clearly the beginning of a while loop near the top of the obfuscated version. There's no loops at all in the 'de-obfuscated' version.
- bobkazamakis 4y agoThat's not generally a good enough indicator; plenty of obfuscation involves loops that otherwise aren't hit.
- enjoytheview 4y agoYeah, pretty sure that loop is to modify the lookup table for strings in runtime so you can't just statically replace it in the code, it's not strong obfuscation but in a properly deobfuscated code that loop wouldn't exist
- cocomutator 4y agoHere's GPT's own explanation what the purpose of that while loop is: --- This code uses JavaScript's `eval` function to obfuscate the code by looping over an array of strings and passing them as arguments to `eval` to create a variable. It also uses an anonymous function to obfuscate the code. The code is deobfuscated by replacing the `eval` function and the anonymous function with their respective strings.
- twic 4y agoThat explanation is also not correct! The while loop is an obfuscation gadget (of sorts), but it doesn't use eval, it uses push and shift to rotate the array. The only use of eval is 'eval("find")', which is not top grade obfuscation.
- KomoD 4y agoWasn't really a particularly difficult sample... You know what it does by just looking at it.
- dom96 4y agoAlso interested about its accuracy. I built a C/C++ obfuscator[0] a while back, I'll have to make a note to see how well ChatGPT deobfuscates its output. [0] - https://picheta.me/obfuscator https://picheta.me/obfuscator
- arriu 4y agoCould this be taken to the extreme by asking it to decompile assembly?
- pathartl 4y agoI've been using it paired with Ghidra to give me a better idea of what's going on. It helped me create a no-CD crack for an old game.
- amrb 4y agoSilly idea tho would it be helpful in creating the server side code for an online only game? I've seen a project for battlefield 3 tho already have the feeling it's a team effort at minimum?
- pathartl 4y agoIt's not some magic bullet. It helps a ton with trying to give names to obfuscated function and variable names, but you have to be intelligent enough to know what the code's actually doing. It probably helps RE teams a lot, but until it can easily run across an entire codebase it's just another tool in the toolbox.
- amrb 4y agoAgreed its can lacking logic till to spell it out, example was this ctf challenge and it just could understand the hash collision till I gave it the full write up [0] w.r.t codebases I may look at some of the free models (as this gets around the cost problem) and try to feed it prompts as a block of code plus meaningful references to same under the token limit. [0] https://github.com/victor-li/pwnable.kr-write-ups/blob/master/collision.md https://github.com/victor-li/pwnable.kr-write-ups/blob/maste...
- the_mitsuhiko 4y agoI have used ChatGPT somewhat successfully to decompile assembly in to C and C++. It's making a lot of mistakes but despite all of this, it's very helpful.
- mlatu 4y agoi would rather ask it to give me the python source of a program capable to deobfuscate THAT string. anything else is just naive, wishfull thinking and a waste of time. you will have to deobfuscate the code manually anyways but at least you got on HN EDIT: OMG... the obfuscated code was generated by chatgpt AND DOESNT EVEN WORK
- garrettjoecox 4y agoAre you just now realizing these LLM’s can’t factually determine if what they output is correct?
- mlatu 4y agono, i am just now realizing what a clout hero alex is
- twic 4y agoNo, i'm just now realizing that HN users can’t factually determine if what these LLM's output is correct.
- Sakos 4y agoTurns out we're all just ChatGPTs after all.
- whamlastxmas 4y agoIt actively makes shit up in my experience. Like outright lies about basic facts.
- takeda 4y agoNow I'm starting to wonder if Elon's tweets aren't him just testing ChatGPT.
- teeray 4y agoEvery commenter is jumping in asking “is it correct?” Even if it’s not 100%, if it’s at least reasonably close, it could be a tremendous force-multiplier against obfuscation for someone with some familiarity with roughly what the code is trying to do.
- jerf 4y agoCorrectness is a big deal here. This is a security context and we can assume that the obfuscators are active attackers against legible code, not just people passively hoping that their obfuscated code is obfuscated. If this becomes a popular technique, then code obfuscation tools will simply pivot to writing code that ChatGPT gets wrong when asked to unobfuscate it. I can't even imagine that would be a particularly hard thing to do, especially if it isn't correct even before actively attacking ChatGPT! Fooling it even harder won't be terribly difficult. This is advantage attacker overall. I imagine it would be as easy as using some cognitively loaded, but wrong, terms as variable names instead of short letters and numbers. Ask ChatGPT "please unobfuscate this network code" and get back a substring search algorithm because the network code was written with a dozen variants on "haystack" and "needle" for variable names, for instance. ChatGPT being actively wrong would be a step back for such deobfuscators then, not a positive at all.
- moyix 4y agoI don't really see this as a problem. Once you have a first cut deobfuscation from this you can refine it with other methods, like comparing input/output examples between the original and the deobfuscated version, or even use something more sophisticated like symbolic execution [1] or differential fuzzing [2] to systematically look for divergence between the behavior of the two. You could even feed these back in to ChatGPT and ask it to redo the deobfuscation given a failing test case. Such testing won't be able to prove that the two are equivalent (unless it's exhaustive) but with decent coverage of the original you can get some good confidence. The goal of deobfuscation is usually understanding, so I'm not sure you need strong guarantees of perfect semantic equivalence with no human intervention/judgment. And of course, existing deobfuscators have bugs and aren't guaranteed to preserve semantics either. [1] https://en.wikipedia.org/wiki/Symbolic_execution https://en.wikipedia.org/wiki/Symbolic_execution [2] https://en.wikipedia.org/wiki/Differential_testing https://en.wikipedia.org/wiki/Differential_testing
- ZyanWu 4y agoMisleading, ChatGPT generated both obfuscated and de-obfuscated code
- worldsavior 4y agoSo what?
- sidewndr46 4y agoA reversible function is not a very interesting development in the world.
- mlatu 4y agothe "obfuscated" "version" doesnt even work. alex literally asked chatgpt to come up with a math problem and its solution, both from whole cloth. and you ask so what. well, everyone should ask "so what?" to alex.
- wittycardio 4y ago[dead]
- worldsavior 4y agoI meant "so what" by so what it generated the obfuscated code, it still de-obfusacted it which is impressive.
- mlatu 4y agofirst it generated garbage that doesnt even run. and then it generated something that looks like valid source based on that garbage. is the source it spat out runable? then it is not the same program as the input and (any way you spin it) nothing has been deobfuscated but just dreamed about the prompt a bit and then shown you its dream diary notes. wake up boy, this is statistics, not a magic swiss army knife API.
- shever73 4y agohttps://twitter.com/AlexAlexandrius/status/1617885287048425478?s=20&t=fVNqR7S_mhvIDHtRe50Wrw https://twitter.com/AlexAlexandrius/status/16178852870484254... So, all I've learned is that ChatGPT knows the obfuscated and de-obfuscated versions of code that it itself has generated.
- buster 4y agoWhich is not runnable in the first place. Interesting though, as this shows exactly the problem. It looks legit, but it's some generated fake text.
- letmevoteplease 4y agoI just tested it on a little snippet of my code obfuscated with https://obfuscator.io/ https://obfuscator.io/ and it worked almost perfectly. My original code: function resizeImage(img) { var maxHeight = 350; var ratio = 1; if(img.height > maxHeight) { ratio = maxHeight / img.height; } var width = img.width * ratio var height = img.height * ratio; var canvas = document.createElement('canvas'); canvas.height = height; canvas.width = width; var ctx = canvas.getContext('2d'); ctx.drawImage(img, 0, 0, width, height); return canvas.toDataURL("image/jpeg", 0.8); } ChatGPT's answer: https://i.imgur.com/5jgPMEd.png https://i.imgur.com/5jgPMEd.png
- sdflhasjd 4y ago
- dsabanin 4y agoI have to say, I find all the comments dismissing ChatGPT hilarious. I read them in a funny grandpa voice. However, we should look past the insignificant details. The main achievement is that we now have a really capable unstructured text-to-computer interface. We can hook it up to anything and it will give us answers with whatever properties we desire, in whatever shape we can think of.
- megous 4y ago[flagged]
- dang 4y agoPersonal attacks will get you banned here, so please don't post like this. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- megous 4y agoI understand. The response was generated by ChatGPT with a prompt "make a snarky response". I just found it amusing to use the technology itself against its own proponent who insinuated that everyone having something against the technology is somehow a tech averse "grandpa".
- nprateem 4y agoI know, all those frusty old grumpyboots who actually want the thing to return factually accurate answers and valid code. Just be happy with plausible sounding answers people. Sheesh!
- aliqot 4y agoI've had a few instances where it returned bad code or was unable to solve a challenge, however most were fixed by better prompts, or by clarifying prompts. In a way I think there is a two-way "learning" process going on here. I'm training it how to give me what I ask for, and it trains me how to ask for what I want it to give me.
- deleted 4y ago[deleted]
- RektBoy 4y agoGive it binary obfuscated with VMprotect
- rednerrus 4y agoIt's been my experience that ChatGPT gets things wrong sometimes. It's also been my experience that if you say "X isn't working as expected." it will do its best to fix the issues. It usually does a pretty good job of fixing it. I had it write a handful of scripts for me yesterday. It got about 90% of it right on the first pass and 99% of it right on the second pass. You still need to have some understanding of what you're doing so you can see when things are wrong but man if it doesn't save you a lot of time.
- 1970-01-01 4y agoSoundness implies validity, but validity does not imply soundness.
- azatom 4y agoTwo types of comments: - "the result doesn't even work..." yeah, even to be a rubberduck is amazing, what thll you expect, got your payment too? - "wow, amazing...": not really, best case it is a google without (direct) advertisement, found the original code/very similar parts.. tryied with own obfuscated code, not from the net... can not get anything from it
- jcadam 4y agoI asked ChatGPT to write me a fibonacci sequence generator suitable for entry into the obfuscated C contest. And it actually spit out something reasonably well obfuscated (used lots of macros).
- omgomgomgomg 4y agoAnd the thing does not learn, it does not save the good answers somewhere nor does it discard the bad ones. I have asked it to generate some mac os compatible vba code, oh dear, that "macro" never worked so far.
- captainmuon 4y agoWow, I wonder how many "bytes of state" ChatGPT must have internally to be able to do that transform. Or does it guess from certain sequences and just writes something plausible? It would be interesting to test if it can solve "memory hard" problems, like repeated obfuscation or hashing, that would require large internal state. Then you could map how large the equivalent memory would be.
- shagie 4y agoFortunately, there's a good repository of obfuscated code. https://www.ioccc.org https://www.ioccc.org int main(int b,char**i){long long n=B,a=I^n,r=(a/b&a)>>4,y=atoi(*++i),_=(((a^n/b)*(y>>T)|y>>S)&r)|(a^r);printf("%.8s\n",(char*)&_);} If you want to get the same (or similar) results: curl https://api.openai.com/v1/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OPENAI_API_KEY" \ -d '{ "model": "code-davinci-002", "prompt": "Deobfuscate the following C code:\n\nint main(int b,char**i){long long n=B,a=I^n,r=(a/b&a)>>4,y=atoi(*++i),_=(((a^n/b)*(y>>T)|y>>S)&r)|(a^r);printf(\"%.8s\\n\",(char*)&_);}", "temperature": 0, "max_tokens": 256, "top_p": 1, "frequency_penalty": 0, "presence_penalty": 0 }' That one loops on the response: The code is obfuscated with the following rules: * `B` is the value of `b` * `I` is the value of `i` * `T` is the value of `y` * `S` is the value of `_` Changing to: "model": "text-davinci-003", int main(int b, char** i) { long long n = b, a = *i[0] ^ n, r = (a / b & a) >> 4, y = atoi(*++i), _ = (((a ^ n / b) * (y >> 8) | y >> 7) & r) | (a ^ r); printf("%.8s\n", (char*)&_); } Which... isn't exactly helpful. Asking it to explain it instead: This code is a C program that prints out a string. The program takes two parameters, b and i, and assigns them to variables n and a. Then the program performs various calculations on those variables and the parameters, such as shifting and bitwise operations. Finally, the program prints out the result of the calculations, which is stored in the variable _, as an 8-character string.