7 ms·
Good C code will try to avoid allocations as much as possible in the first place. You absolutely don’t need to copy strings around when handling a request. You
by bluetomcat 1y ago
Good C code will try to avoid allocations as much as possible in the first place. You absolutely don’t need to copy strings around when handling a request. You can read data from the socket in a fixed-size buffer, do all the processing in-place, and then process the next chunk in-place too. You get predictable performance and the thing will work like precise clockwork. Reading the entire thing just to copy the body of the request in another location makes no sense. Most of the “nice” javaesque XXXParser, XXXBuilder, XXXManager abstractions seen in “easier” languages make little sense in C. They obfuscate what really needs to happen in memory to solve a problem efficiently.
- 01HNNWZ0MV43FF 1y agoCan you do parsing of JSON and XML without allocating?
- bluetomcat 1y agoYes, you can do it with minimal allocations - provided that the source buffer is read-only or is mutable but is unused later directly by the caller. If the buffer is mutable, any un-escaping can be done in-place because the un-escaped string will always be shorter. All the substrings you want are already in the source buffer. You just need a growable array of pointer/length pairs to know where tokens start.
- gritzko 1y agoYep, no problem. In place parsing only requires a stack. Stack length is the maximum JSON nesting allowed. I have a C dialect exactly like that.
- veqq 1y agoOf course. You can do it in a single pass/just parse the token stream. There are various implementations like: https://zserge.com/jsmn/ https://zserge.com/jsmn/
- Ygg2 1y agoTheoretically yes. Practically there is character escaping. That kills any non-allocation dreams. Moment you have "Hi \uxxxx isn't the UTF nice?" you will probably have to allocate. If source is read-only you have to allocate. If source is mutable you have to waste CPU to rewrite the string.
- lelanthran 1y ago> Moment you have "Hi \uxxxx isn't the UTF nice?" you will probably have to allocate. Depends on what you are doing with it. If you aren't displaying it (and typically you are not in a server application), you don't need to unescape it.
- mpyne 1y agoAnd this is indeed something that the C++ Glaze library supports, to allow for parsing into a string_view pointing into the original input buffer.
- deaddodo 1y agoI'm confused why this would be a problem. UTF-8 and UTF-16 (the only two common unicode subsets) are a maximum of 4 bytes wide (and, most commonly, 2 in English text). The ASCII representation you gave is 6-bytes wide. I don't know of many ASCII unicode representations that have less bytewidth than their native Unicode representation. Same goes for other characters such as \n, \0, \t, \r, etc. All half in native byte representation.
- topspin 1y ago> Practically there is character escaping The voice of experience appears. Upvoted. It is conceivable to deal with escaping in-place, and thus remain zero-alloc. It's hideous to think about, but I'll bet someone has done it. Dreams are powerful things.
- _3u10 1y agoIt’s just two pointers the current place to write and the current place to read, escapes are always more characters than they represent so there’s no danger of overwriting the read pointer. If you support compression this can become somewhat of and issue but you simply support a max block size which is usually defined by the compression algorithm anyway.
- lelanthran 1y ago> Can you do parsing of JSON and XML without allocating? If the source JSON/XML is in a writeable buffer, with some helper functions you can do it. I've done it for a few small-memory systems.
- zzo38computer 1y agoIt depends what you intend to do with the parsed data, and where the input comes from and where the output will be going to. There are situations that allocations can be reduced or avoided, but that is not all of them. (In some cases, you do not need full parsing, e.g. to split an array, you can check if it is a string or not and the nesting level, and then find the commas outside of any arrays other than the first one, to be split.) (If the input is in memory, then you can also consider if you can modify that memory for parsing, which is sometimes suitable but sometimes not.) However, for many applications, it will be better to use a binary format (or in some cases, a different text format) rather than JSON or XML. (For the PostScript binary format, there is no escaping, and the structure does not need to be parsed and converted ahead of time; items in an array are consecutive and fixed size, and data it references (strings and other arrays) is given by an offset, so you can avoid most of the parsing. However, note that key/value lists in PostScript binary format is nonstandard (even though PostScript does have that type, it does not have a standard representation in the binary object format), and that PostScript has a better string type than JavaScript but a worse numeric type than JavaScript.)
- megous 1y agoYes, you can first validate the buffer, to know it contains valid JSON, and then you can work with pointers to beginings of individual syntactic parts of JSON, and have functions that decide what type of the current element is, or move to the next element, etc. Even string work (comparisons with other escaped or unescaped strings, etc.) can be done on escaped strings directly without unescaping them to a buffer first. Ergonomically, it's pretty much the same as parsing the JSON into some AST first, and then working on the AST. And it can be much faster than dumb parsers that use malloc for individual AST elements. You can even do JSON path queries on top of this, without allocations. Eg. https://xff.cz/git/megatools/tree/lib/sjson.c https://xff.cz/git/megatools/tree/lib/sjson.c
- acidx 1y agoYes! The JSON library I wrote for the Zephyr RTOS does this. Say, for instance, you have the following struct: struct SomeStruct { char *some_string; int some_number; }; You would need to declare a descriptor, linking each field to how it's spelled in the JSON (e.g. the some_string member could be "some-string" in the JSON), the byte offset from the beginning of the struct where the field is (using the offsetof() macro), and the type. The parser is then able to go through the JSON, and initialize the struct directly, as if you had reflection in the language. It'll validate the types as well. All this without having to allocate a node type, perform copies, or things like that. This approach has its limitations, but it's pretty efficient -- and safe! Someone wrote a nice blog post about (and even a video) it a while back: https://blog.golioth.io/how-to-parse-json-data-in-zephyr/ https://blog.golioth.io/how-to-parse-json-data-in-zephyr/ The opposite is true, too -- you can use the same descriptor to serialize a struct back to JSON. I've been maintaining it outside Zephyr for a while, although with different constraints (I'm not using it for an embedded system where memory is golden): https://github.com/lpereira/lwan/blob/master/src/samples/techempower/json.c https://github.com/lpereira/lwan/blob/master/src/samples/tec...
- lock1 1y agoWhy does "good" C have to be zero alloc? Why should "nice" javaesque make little sense in C? Why do you implicitly assume performance is "efficient problem solving"? Not sure why many people seem fixated on the idea that using a programming language must follow a particular approach. You can do minimal alloc Java, you can simulate OOP-like in C, etc. Unconventional, but why do we need to restrict certain optimizations (space/time perf, "readability", conciseness, etc) to only a particular language?
- bluetomcat 1y agoBecause in C, every allocation incurs a responsibility to track its lifetime and to know who will eventually free it. Copying and moving buffers is also prone to overflows, off-by-one errors, etc. The generic memory allocator is a smart but unpredictable complex beast that lives in your address space and can mess your CPU cache, can introduce undesired memory fragmentation, etc. In Java, you don't care because the GC cleans after you and you don't usually care about millisecond-grade performance.
- jstimpfle 1y agoNo. Look up Arenas. In general group allocations to avoid making a mess.
- rictic 1y agoIf you send a task off to a work queue in another thread, and then do some local processing on it, you can't usually use a single Arena, unless the work queue itself is short lived.
- jenadine 1y agoI don't see how arenas solve the problems.
- jstimpfle 1y agoYou group things from the same context together, so you can free everything in a single call.
- lelanthran 1y ago> Good C code will try to avoid allocations as much as possible in the first place. I've upvoted you, but I'm not so sure I agree though. Sure, each allocation imposes a new obligation to track that allocation, but on the downside, passing around already-allocated blocks imposes a new burden for each call to ensure that the callees have the correct permissions (modify it, reallocate it, free it, etc). If you're doing any sort of concurrency this can be hard to track - sometimes it's easier to simply allocate a new block and give it to the callee, and then the caller can forget all about it (callee then has the obligation to free it).
- 1718627440 1y agoTo reduce the amount of allocation instead of: struct parsed_data * = parse (...); struct process_data * = process (..., parsed_data); struct foo_data * = do_foo (..., process_data); you can do parse (...) { ... process (...); ... } process (...) { ... do_foo (...); ... } It sounds like violating separation of concerns at first, but it has the benefit, that you can easily do procession and parsing in parallel, and all the data can become readonly. Also I was impressed when I looked at a call graph of this, since this essentially becomes the documentation of the whole program.
- ambicapter 1y agoHow testable is this, though?
- 1718627440 1y agoIt might be a problem when you can't afford side-effects that you later throw away, but I haven't experienced that yet. The functions still have return codes, so you still can test, whether a correct input results in no error check being followed and that incorrect input results in an error check being triggered.
- throwawaymaths 1y agois there any system where doing the basics of http (everything up to framework handoff of structured data) are done outside of a single concurrency unit?
- fulafel 1y agoThis shared memory and pointer shuffling is of course fraught with requiring correct logic to avoid memory safety bugs. Good C code doesn't get you pwned, I'd argue.
- jenadine 1y ago> Good C code doesn't get you pwned, I'd argue. This is not a serious argument because you don't really define good C code and how easy or practical it is to do. The sentence works for every language. "Good <whatever language> code doesn't get you pwned" But the question is whether "Average" or "Normal" C code gets you pwned? And the answer is yes, as told in the article.
- fulafel 1y agoThe comment I was responding to suggested Good C Code employes optimizations that, I opined, are more error prone wrt memory safety - so I was not attempting to define it, but challenging the offered characterisation.
- riedel 1y agoA long time ago I was involved in building compilers. It was common that we solved this problem with obstacks, which are basically stacked heaps. I wonder one could not build more things like this, where freeing is a bit more best effort but you have some checkpoints. (I guess one would rather need tree like stacks) Just have to disallow pointers going the wrong way. Allocation remains ugly in C and I think explicit data structures are are definitely a better way of handling it.
- self_awareness 1y agoThat mythical "Good C Code", which is known only to some people who I never met.
- pjmlp 1y agoThese abstractions were already common in enterprise C code decades before Java came to be, thanks to stuff like Yourdon Structured Method. Using fixed size buffers doesn't fix out of bounds errors, and stack corruption caused by such bugs. Naturally we all know good C programmers never make them. /s
- fsckboy 1y ago>Good C code will try to avoid allocations as much as possible in the first place. there's a genius to this: if you're going to optimize prematurely, do it right out of the gate!
- wfn 1y agoAgree re: no need for heap allocation - for others: I recommend reading thru whole masscan source (https://github.com/robertdavidgraham/masscan https://github.com/robertdavidgraham/masscan), it's a pleasure btw - iirc rather few/sparse malloc()s which are part of regular I/O processing flow (there will be malloc()s which depending on config etc. set up additional data structs but as part of setup).