13 ms·
You can't copy code with memcpy
- jnwatson 5y agoAt least one additional step which is required on some architectures is you must flush the data cache and invalidate the instruction cache at the location of the new code. Dynamically loading code is indistinguishable from self-modifying code, and each architecture has special steps you must take in order for it to work.
- wmu 5y ago- It looks like a virus... - But it's the antivirus! :)
- jayd16 5y agoBut Doctor, I Am Pagliacci
- dark-star 5y agoI actually did something like that on Windows x86, and it worked fine. Even I was surprised by that fact :) I used it to copy out a (forgotten) password from a password inputfield in another program, which you cannot read remotely (for security reasons). Worked fine for that one use-case, and I haven't used this trick it anywhere else ever again :)
- gavinray 5y agoHow did this work exactly? Suppose you are able to inject and execute remote code into the process containing your password in a text field How did you read the password out of it? Did you know that the password variable was stored at some particular memory location?
- kelnos 5y agoI don't know much about Windows GUI programming, but in other GUI toolkits I've used, there's usually some sort of text_field.get_current_value() function you can call. Presumably the parent injected some code that repurposes the callback of a button or something so that when you click it, it calls get_current_value() and then dumps it to console or a log file or something.
- dark-star 5y agoI don't think I have that code around anymore, at least I can't find it now. And it's been a while. Here is what I remember: Basically you use VirtualAllocEx() to allocate some memory in the remote thread. The returned pointers are in the context of the target process. You can access that remote memory with ReadProcessMemory() and WriteProcessMemory(), which uses those "remote" pointers to copy data to/from your process. You can then use these memory areas to pass global handles and other stuff around. For accessing the actual password field data, you use standard Window-Messages with SendMessage() etc.
- mwcampbell 5y agoSome Windows screen readers used, and maybe still use, the same technique to get data out of the SysListView32 common control, since the parameters and results of that control's window messages aren't marshaled by the OS. Edit: Looks like NVDA still does. https://github.com/nvaccess/nvda/blob/master/source/NVDAObjects/IAccessible/sysListView32.py https://github.com/nvaccess/nvda/blob/master/source/NVDAObje...
- heinrich5991 5y agoPermanent link (press 'y' anywhere on Github): https://github.com/nvaccess/nvda/blob/b5f82f878f344ab26004f39a291da5d65a57efb0/source/NVDAObjects/IAccessible/sysListView32.py https://github.com/nvaccess/nvda/blob/b5f82f878f344ab26004f3....
- unnouinceput 5y agoThis technique is a classic one. I remember learning about it in late 90's after I got infected with a malware that hid itself from Task Manager. That made me also write my own Task Manager. To this day (Win10/Win11) you can hide your program from Task Manager using this technique and any malware that respect itself does it.
- stevekemp 5y agoThat reminds me of the "Shatter Attack" - which basically involved posting messages to the event-loop/window of a different process: https://en.wikipedia.org/wiki/Shatter_attack https://en.wikipedia.org/wiki/Shatter_attack You could do that to call WM_GETTEXT, against an input-control, for example.
- RcouF1uZ4gsC 5y ago> I pointed out to the customer liaison that what the customer is trying to do is very suspicious and looks like a virus. The customer liaison explained that it’s quite the opposite: The customer is a major anti-virus software vendor! The customer has important functionality in their product that that they have built based on this technique of remote code injection, and they cannot afford to give it up at this point. As an aside, whenever I set up a Windows PC for me or a family member, the first thing I do is uninstall any third-party antivirus that may have come with the computer. I have found that anti-virus software likely makes my computer more insecure by having a big attack surface, not to mention slowing it down.
- kmeisthax 5y agoActual Chromium developers have similar opinions w.r.t. antivirus vendors. It's almost like the extensibility necessary to make third-party security products work requires creating entirely new attack surface for those products to work.
- AshamedCaptain 5y agoIt's not as if the first party antivirus inspires any confidence, either. It still loses in performance to some 3rd party ones, and has had its own list of security issues. For those who remember, its early versions (before the MS acquisition) required the Visual Basic runtime.
- ssklash 5y agoThe current Windows Defender is not the crappy old MS AV it used to be. It's gotten drastically better, and for normal home users I always recommend it over installing any 3rd party AV. I'm a pentester and red teamer, and yes, bypassing Defender isn't all that hard, but neither are most 3rd party AVs, and Defender does not bring in all the instability and additional attack surface. Also I think it can take advantage of newer kernel APIs that non-MS AV can't.
- floatboth 5y ago
- skim_milk 5y agoI remember doing exactly this to get code injection working on GNU/Linux systems! I made a library injection library in college for some coursework, which involved copying a C function into a code cave on a remote process and getting the remote process to execute it and return. It only works because its so bare-bones, it doesn't use try/catch and calls to other functions are possible because the function pointers are passed in through the registers and the compiled code is small enough to fit in a single page. An example function I memcpy and run: https://github.com/skimmilk/liblibinject/blob/master/src/liblibinject/external.cpp#L92 https://github.com/skimmilk/liblibinject/blob/master/src/lib...
- tomcam 5y agoWait does C now have try/catch?
- basementcat 5y agoIt’s called setjmp() and longjmp(). https://web.archive.org/web/20091104065428/http://www.di.unipi.it/~nids/docs/longjump_try_trow_catch.html https://web.archive.org/web/20091104065428/http://www.di.uni...
- tptacek 5y agoIt has setjmp and longjmp, so yeah.
- kragen 5y agoAs you are well aware, calling setjmp and longjmp doesn't make code non-position-independent in the way that Raymond is talking about, because setjmp saves the return address with which it was actually called. It doesn't rely on PC tables in the executable the way that (a common implementation of) C++ exception handling does. I mention this not to inform you of it (you already understand it better than I do) but to keep anyone else reading the thread from being misled.
- 5y ago
- leni536 5y agoOf course all of this is very much undefined behavior in standard C and C++. Some programmers really need to learn that they program the "abstract machine" when they write C or C++.
- plorkyeran 5y agoThat is probably the smallest and least interesting of the problems with this. Any method for injecting code into another process is inherently going to be platform-specific and outside the bounds of the C++ abstract machine.
- woodruffw 5y agoI don't think most of this is actually UB. The cast from a function pointer to `BYTE *` is, but the rest is "sound" (AFAICT) from the C abstract machine's perspective. The reason it fails is basically orthogonal to C.
- Kranar 5y agoThe reason it fails is due to UB. There's not much point in saying that most of it isn't UB since the part that is UB is the part that causes the failure.
- woodruffw 5y agoThe cast isn't the part that causes the failure (C says that function pointers don't have to be safely representable within `void *` or any other fundamental pointer type, but they are on both x86(-64) and Itanium[1]). The failure happens because the programmer assumed that the code is "self contained" and position-independent, both of which are concepts outside of the C abstract machine. [1]: Function pointers on Itanium are actually fat, but IIRC most compilers hide this by making the "function pointer" point to some kind of thunk instead.
- Kranar 5y agoThis is the type of reasoning that leads people to write insecure code, by trying to outguess the compiler with implementation details that are entirely inapplicable. It suggests that if the code was self-contained or position independent or satisfied any property whatsoever, that using memcpy would be a perfectly fine way to copy it, that is simply untrue, it would never be safe to use memcpy on pointers to functions under any circumstance. It's irrelevant whether you are on x86, Itanium, or any other platform, undefined behavior WILL result in incorrect and invalid program semantics and can not be relied upon to produce consistent results even within the same execution of the same program. There's literally decades worth of people trying to use undefined behavior in clever ways and ultimately failing and yet here we are...
- justicz 5y ago"This code is such a bad idea, I’ve intentionally introduced errors so it won’t even compile." I am stealing this line.
- qahaz 5y agoThis was common back in the day with ezines. They would include codes to exploit security bugs in common software but they would have intentional errors so script kiddies couldn't compile them.
- commandlinefan 5y agoSeems like running it through a compiler and fixing the errors would address that for you, though…
- behringer 5y agoScript kiddies can't debug.
- DoctorDabadedoo 5y agoWhat is this mysterious art you speak of? /s
- t-3 5y agoReally? Compiler error messages are downright friendly compared to networking tools and concepts. When I was 12, I certainly found them to be much easier, even though I completely failed to learn C, Perl, C++, or any of the other popular languages of the time.
- thecupisblue 5y agoExactly the idea. If you're a skiddie, you'll copy/paste, look for a minute, say it doesn't work. If you dive in, you gotta learn about the language, debugging it, the exploit itself and suddenly you're not a script kiddy anymore. It's like a leap of faith.
- Jyaif 5y agoDon't publicly make fun of your customers... unless they are an anti-virus company. What boggles my mind is how they went on to ask MS for help fixing their obviously wrong vulnerability-and-crash-introducing software.
- Arnavion 5y agoRaymond Chen makes fun of customers all the time. He just doesn't name them. Also, in this particular case the case is old enough that it's not necessarily any AV that's still around anyway.
- deleted 5y ago[deleted]
- phendrenad2 5y agoAs long as the code is truly position-independent, I.E. compiled with -fPIC in GCC/LLVM, then there will be no problem.
- skim_milk 5y agoBut the code being memcpy'ed is using the parent process's specific symbol relocations. When a library with PIC is loaded the executable code is copied from the file into RAM to a random offset and all references to structures in the ASM are updated to match their now random offset in memory (simplifying). Say function XYZ in library libfoo is in offsetXYZ, parent process loads libfoo at offset 0xDEADBEEF, injectee process loads libfoo at offset 0xDEC0DE. In Windows the call to function XYZ in the parent process uses the address offsetXYZ+0xDEADBEEF, but the call in the injectee process uses offsetXYZ+0xDEC0DE, causing any reference to the parent process's function to fail. GNU/Linux is very similar but library symbols are found based on an offset to a structure in memory that contains the library symbol metadata, that changes every time the program is loaded. So actually the opposite is true, if the code wasn't position-independent and was statically located, the assembly code offsets wouldn't need to be updated and you might be able to call a memcpy'd function. Position-independent code could only possibly work if you updated the reference to the symbol metadata structure in the ASM after the memcpy, but at that point you're re-implementing libdl and no longer just using memcpy.
- ssklash 5y agoThis depends on the libraries/DLLs being used. Windows loads system DLLs at the same location in every process's address space, so you can use process-local offsets in a remote process. For custom libraries of course this wouldn't work. Or if the required system library hasn't been loaded in the remote process.
- phendrenad2 5y agoYes, the example Raymond uses isn't PIC.
- userbinator 5y agoI've done stuff like this before; it works very well if you know the limitations, and I'd say that it even gives you a better understanding of how things actually work. Of course, don't bother MS or any other "official" vendor if it doesn't work, because you are on your own in debugging it.
- Terry_Roll 5y ago>because you are on your own in debugging it. "The customer reported that this code “worked just fine on 32-bit x86 and 64-bit x86”, but it doesn’t work on Itanium." "Actually, I’m surprised that it worked even on x86!" And that is probably how Spectre and Meltdown were found! Differences in the cpu's.
- wyldfire 5y agoWell, it should work in some well-controlled cases and in those cases if you're writing this code you are probably an OS vendor. Examples of legit use case: an executable loader or some bootstrapping code/bootloader. When writing code like this, as Chen says, you are bound to the architectural rules regarding how to appropriately locate code and safely invalidate code caches etc. So if you follow the rules it doesn't work then typically I figure you'd take it up with the CPU vendor. While you could potentially do it and have a good reason to in userspace code, it should be heavily scrutinized because it's so unconventional.
- basementcat 5y agoThe client certainly should have made sure their code was truly position independent. Also, the client should have embedded their code in the executable file name so they just have to jump to the appropriate offset in argv[0]. This way, future updates just require renaming the file!
- liftm 5y agoWhy use PIC if you can just write your own relocator? https://fasterthanli.me/series/making-our-own-executable-packer https://fasterthanli.me/series/making-our-own-executable-pac...
- MattPalmer1086 5y agoI imagine they were dynamically building the code to inject, or why bother with the complexity?
- truekonrads 5y agoMuch of red-team/pentest/malware code works like this and is surprisingly reliable.
- mox1 5y agoUhh yea, this is Running shellcode 101, works very well. My Red Team stuff at work all starts with a simple loader like this (with some encryption / obfuscation sprinkled in). When I was first shown this I was like 'What non virus use case does this have!?!?'
- tptacek 5y agoInstrumenting and debugging live processes is the big one.
- dotancohen 5y agoWhich is not an end-user advantage but rather a tool to better understand end-user software - equally useful for improving said software or attacking it. Assuming GP meant "virus" as in "malware" then actually this supports his point.
- ssklash 5y agoRed teamer here too, and this was my exact thought. There's lots of legit uses for DLL injection, but straight up shellcode injection? Shady as hell. So of course it was an AV vendor...
- deleted 5y ago[deleted]
- beaconstudios 5y agoreally? I've not written much shellcode at all, but what I did write wasn't generically-compiled C++ - it was always either C or ASM, specifically because you get to avoid all the platform and position-dependent stuff (except in return-to-libc payloads).
- ouid 5y ago"No, the opposite of a virus, an anti-virus". Thereby demonstrating that a virus is its own anti-particle.
- watersb 5y agoWhen I read this quote from the customer in the story, I wasn't surprised. I figured it was likely that code that aggressively scans and modifies other running executables would be written as a kludge, an unorthodox way of abusing the compiler-loader-runtime chain.
- anothernewdude 5y agoKaspersky actually has been known to travel networks.
- MonkeyClub 5y agoBit more on this, please? Sounds interesting!
- donkarma 5y agoDoes anyone else remember reading an Old New Thing article like this one before?
- donkarma 5y agoFound it, https://devblogs.microsoft.com/oldnewthing/20180615-00/?p=99025 https://devblogs.microsoft.com/oldnewthing/20180615-00/?p=99...
- WalterBright 5y ago> This code is such a bad idea, I’ve intentionally introduced errors so it won’t even compile. No problem, I just fixed the compiler to compile it!
- WalterBright 5y agoIn my 1980s version of Empire, all the global variables were kept contiguously in one source file. To save/restore the game, it just took the address of the first one, the address of the last one, and blitted it to a disk file, and blitted it back. Very fast & easy. Of course, it broke when COMDATs were introduced. I did a similar thing with my text editor. The colors were configurable. The usual way was to have a configuration file, which the editor would read upon startup. But floppy disk systems were unbearably slow. So what I did was take the address of the configuration data in the data segment. I'd work backwards to where those bytes were in the EXE file, and patch the EXE file. This worked great! Until the advent of virus scanners, which broke that. Virus scanners hated self-modifying EXE files.
- withinrafael 5y agoHow did it break? COMDAT usage is optional today. (Perhaps it wasn't when first created?)
- WalterBright 5y agoCOMDATs aren't placed sequentially by the linker. Also, zero initialized data got placed in a separate section.
- pvillano 5y agohow would you do this today? could you wrap it in a struct?
- __david__ 5y agoSee also: Emacs's old unexec() function. How do you speed up a bunch standard library lisp loading and initialization? Just do it once then write a new exe out with all your state pre-computed. Genius.
- barchar 5y agothere's also unfork(2)! Do you unfork then unexec or unexec then unfork?
- SystemOut 5y agoThis reminded me of the old days working in Windows 3.1 and my first professional project was to write a SOCKS client that could be loaded up and intercept all calls to Winsock's connect() function. It needed to do this without modifying the other programs and it had to happen at the DLL level and not the VxD layer where our IP stack ran. Turns out there was an undocumented Windows API function along the lines of "AliasCsToDsRegister" or something like that - I've tried to find a reference to it but I can't find it. It allowed me write into the code segment (the CS was global and read only - as it was shared among all processes) and replace the first few bytes of the connect function call with a jump to my code which would then put it back, make the call to the socks server, do some other magic, put my jump hook back in and the return to the caller. Good times! Kind of surprised I remember this and more so that it actually worked.
- adontz 5y agohttps://www.microsoft.com/en-us/research/project/detours/ https://www.microsoft.com/en-us/research/project/detours/
- pacaro 5y agoYeah win16 had AllocCStoDS and AllocDStoCS, one was documented the other wasn't. They also had the fabulously named PrestoChangoSelector which toggled the code/data bit in the descriptor table
- throwaway2037 5y agoCool. I did some Google searching about PrestoChangoSelector and I found this: KnowledgeBase Archive An Archive of Early Microsoft KnowledgeBase Articles Q89560: Creating Dynamic Code Segments Using PrestoChangoSelector https://jeffpar.github.io/kbarchive/kb/089/Q89560/ https://jeffpar.github.io/kbarchive/kb/089/Q89560/
- bonzini 5y agoIIRC there was a documented ChangeSelector function that was actually not implemented, so you had to use the undocumented PrestoChangoSelector instead. Visual Basic used it to implement a direct threaded code interpreter. Some other functions had similarly great names: >>Anyway, I am not sure what type of person writes a function called "BozosLiveHere" and puts it into USER.EXE >This started out life with a non-bozo name as an undocumented function in Windows 3.0. >Windows 3.1 removed the undocumented function, but we found that some programs were using the undocumented function and started crashing. >So we reluctantly put the function back, but changed its name to "BozosLiveHere" so that nobody else would use it in the future. >A similar story exists for "TabTheTextOutForWimps".
- seligman99 5y agoI wrote a "cd" replacement for cmd[1] a _long_ time ago (I only recently uploaded it to Github). It uses exactly this technique to run a thread in cmd's process to actually change the directory. It's kept working from XP on up to Windows 11 now. I am always amazed it works, I fully expect it to go boom some day, probably with an error along the lines of "Don't do that, please". [1] https://github.com/seligman/ccd/blob/master/RemoteThread.cpp#L416 https://github.com/seligman/ccd/blob/master/RemoteThread.cpp...
- rightbyte 5y agoTerrible hacks is an art form that is hugely underrated today in the name of overengineered best practice complexity monsters. Sometimes, just doing it the stupid way is simpler than doing it properly. Especially when working with propertiary systems ...
- ggm 5y agoYet, copy-on-write works well in Unix fork/exec() models and helps reduce memory pressure. Presumably, the kernel has a mechanism which presents as logistically simple "copy" but takes care of page/pointer/vm necessity.
- bzbarsky 5y agoPosition-dependent code can be copied in _physical_ memory as long as the _virtual_ addresses (which is what you observe) don't change. So yes, the mechanism you posit exists, and its the virtual memory manager.
- SeanLuke 5y agoWhen I took a compiler class back in the early '90s, the project was to write the compiled machine code into an array, then cast the array into a function, and execute it. I and another student were doing it on 68040 NeXT workstations. One other student was doing it on a Mac, one on a VAX, and the rest on PCs (the PC students largely failed!). We were mystified why, when we tried to execute our code, it was as if it wasn't there. Took us a while to realize that the 68040 had separate instruction and data caches, and even more time (and emailing people at NeXT) to determine what the cache flush procedure was.
- rkeene2 5y agoI once saw code like: unsigned short main[] = {0104525, 0xb8e5, 32, 0, 0xC35D}; and I made a tool to generate it automatically from a function [0]. [0] https://rkeene.org/viewer/tmp/ret32-maker.c.htm https://rkeene.org/viewer/tmp/ret32-maker.c.htm
- foota 5y agoYou may be thinking of https://jroweboy.github.io/c/asm/2015/01/26/when-is-main-not-a-function.html https://jroweboy.github.io/c/asm/2015/01/26/when-is-main-not...
- rkeene2 5y agoThis reminds me of the time I wanted to run binaries compiled for SSE3 on a system that lacked SSE3. I started writing a tool to emulate this [0], and one thing it could do is rewrite the executable pages with replacement instructions if there was something that would fit (using memcpy(2), naturally). This harkens back to the days when you could "download" a math coprocessor for your SX system, which was a TSR which likely did the same catching and handling of illegal instructions. [0] https://github.com/rkeene/sse3-emu/blob/master/libsse3.c https://github.com/rkeene/sse3-emu/blob/master/libsse3.c
- tjalfi 5y agoOlder versions of Windows used a similar technique to replace the PREFETCHW instruction with NOPs on processors that didn't support it.
- fork-bomber 5y agoMemcpying and executing code could also surface micro-architectural realities of the underlying CPU and memory subsystem micro-architecture that may need attention from the programmer. For example: - On most RISCy Arm CPUs with Harvard style split instruction and data caches special architecture specific actions would need to be taken to ensure that after the memcpy any code still lingering in the data cache was cleaned/pushed out to the intended destination memory immediately (instead of at the next cache line eviction). - Any stale code that happened to be cached from the destination (either by design or coincidence) needs to be invalidated in the instruction cache. - Depending on the CPU micro architecture, programmer unknown speculative prefetching into caches as a result of the previous two actions may also need attention.
- oshiar53-0 5y agoIs there anything else that should be done except flushing D-cache and invalidating I-cache? I'm genuinely curious.
- jlokier 5y agoIf you are asking about general cases: If using paging you may need to invalidate the TLB entry which contains execute permission for the page. On x86 if using segments, after changing segment attributes you need to reload the segment selectors. The execution pipeline may need to be flushed, using a serialising instruction. When modifying code in place that may be being executed by another thread on another core at the same time, some modifications may trigger CPU errata. On particular CPUs there may be other kinds of caches or state invalidation required, but hopefully the OS provides a "flush I-cache" function that covers all of them.
- _8j50 5y agoMalware authors disagree.
- weinzierl 5y agoIsn't this whole idea of injecting code on modern systems doomed because of write xor execute (W^X, NX), which is hopefully enabled?
- liftm 5y agoIf you can allocate memory in a foreign process, I'd guess you could also change the permissions on that memory… So write first, then change to executable. (VirtualProtectEx looks like it would do that. Never used winapi, not sure.)
- oshiar53-0 5y agoRelated: https://devblogs.microsoft.com/oldnewthing/20190902-00/?p=102828 https://devblogs.microsoft.com/oldnewthing/20190902-00/?p=10...
- rwmj 5y agoI think the even crazier thing is Windows having a function to allocate memory in another process (https://docs.microsoft.com/en-us/windows/win32/api/memoryapi/nf-memoryapi-virtualallocex https://docs.microsoft.com/en-us/windows/win32/api/memoryapi...). That seems like a potential source for all kinds of impossible to track bugs. > The customer is a major anti-virus software vendor! The customer has important functionality in their product that that they have built based on this technique of remote code injection, and they cannot afford to give it up at this point. Oh now it makes sense.
- jrtc27 5y agoYou can allocate memory in another process on Unix too: use ptrace to make the other process call malloc (use PTRACE_SETREGS to set PC to malloc and the first argument register to the number of bytes, then intercept the return). GDB will use this if you tell it something like `p foo("bar")`, as it needs to allocate memory for that string somewhere.
- watermelon0 5y agoWhat would be the use case of AV software needing to perform remote code injection?
- rwmj 5y agoThere's barely a case for AV software existing at all. It's the root cause for all kinds of impossible to track bugs and performance regressions.
- DarkWiiPlayer 5y agoThe best kind of antivirus is the one that stays quiet unless you specifically ask it to scan a file. Did someone say Clam AV just now?
- Darvon 5y agoWhy did you download an executable that you don't trust? Don't even scan it, just delete it and go find one you do trust.
- 0xdky 5y agoI froze for a moment seeing this article after having worked at a major anti-virus company long time back and used some low level Win32 APIs. Fortunately, I followed some of the techniques from “Programming Applications for Microsoft Windows” book and Detours project to intercept and execute custom code mostly based on loading custom DLL in target remote process and using DllMain() to execute.
- jl2718 5y agoI was competing in the Jump Trading programming competition and thought I had a pretty good implementation in AVX asm, but I was still behind one of their engineers, so I asked him after the competition. Turns out he was a Linux kernel committer and wrote a process to spawn multiple threads by copying itself, modifying the parameters, and then setting the offsets directly in the thread table, avoiding all mallocs and thread startup. So basically, his math code was just basic C loops, but his process was complete before my threads even finished allocation. Forgive me if I got it wrong, I am definitely not a Linux kernel committer.
- mrlonglong 5y agoYes, you can copy code with memcpy as long as it's position independent code!