4 ms·
Go did fix the C mistake: package main func foo(a []byte) { println(len(a)) } func main() { var b []byte // cann
by 2h 3y ago
Go did fix the C mistake:
package main
func foo(a []byte) {
println(len(a))
}
func main() {
var b []byte
// cannot use &b (value of type *[]byte) as []byte value in argument to foo
foo(&b)
}
- vimda 3y agoBasically every modern language has solved this...
- throwawaymaths 3y agoNot only that, but some have solved it while maintaining compatibility with null terminated strings. Null terminated strings, after all, are sometimes more efficient.
- zerodensity 3y agoCan't think of a single case where null terminated strings are more efficient. Could you give some examples?
- School-Cotton 3y agoIf you have to walk the string anyway, the null terminator has no downside.
- zerodensity 3y agoAccording to the bible (https://www.agner.org/optimize/ https://www.agner.org/optimize/) it's faster to use a loop with length than walking though a pointer so not having a length will make it slower to walk the string whole also making things like simd optimizations harder for the compiler to do.
- throwawaymaths 3y agoThat doesn't make sense. If you have loop with length you have to check both the content of the byte and the index; if you have null terminated strings you only check the content of the byte.
- GrumpySloth 3y agoWhen you have the length, you can unroll the loop, so that you e.g. do 4 iterations at a time. With NUL you can’t do that. Moreover, loop iteration can be done in parallel (instruction-level parallelism) with processing the content of the string, since there is no data dependency between the two. With NUL you introduce a data dependency.
- throwawaymaths 3y agoYou can't always do those things. Yes, pointer length is almost always faster. But it's not always faster. https://lemire.me/blog/2020/09/03/sentinels-can-be-faster/?amp https://lemire.me/blog/2020/09/03/sentinels-can-be-faster/?a...
- yakubin 3y agoYou can. Here is my version: <https://godbolt.org/z/91KecEfbM https://godbolt.org/z/91KecEfbM> Results: N = 10000 range 1218.1 sentinel 937.258 mine 672.227 ratio 1.29964 (ratio is range/sentinel, not mine/whatever) You can get even more crazy with SIMD. But for all that you need to know the length beforehand. Edit: The b4 variable should actually be called b8. That's a remnant of a previous version, where I used 32-bit chunks.
- throwawaymaths 3y agoRead the blog post. Dan lemire isn't exactly a slouch.
- 3y ago
- kevin_thibedeau 3y agoYou can modify the length of NUL strings in place by inserting a new NUL or overwriting past the end without any other bookkeeping. You can split a string on delimiters simply by overwriting them with NULs.
- jrpelkonen 3y agoI find these arguments rather weak. I don’t see how writing a NUL is any more efficient compared to updating a length. Furthermore, having the terminators in-band prevent the character data to be used for multiple substrings. E.g. in the split example the original string is no longer available.
- slaymaker1907 3y agoIt takes 4-8 bytes to represent the size of the string versus 1 byte for a null terminator. That doubles the size of the string when you embed it in a struct or pass it as an argument on the stack. In particular, remember that even today cache lines are only 64 bytes for x86-64 and while that seems like a lot, going from 64 bytes to 68 means you go from 1 cache miss to load some struct to 2 cache misses.
- GrumpySloth 3y agoSame as going from 64 to 65.
- throwawaymaths 3y ago> Can't think of a single case where null terminated strings are more efficient https://lemire.me/blog/2020/09/03/sentinels-can-be-faster/?amp https://lemire.me/blog/2020/09/03/sentinels-can-be-faster/?a...
- 2h 3y agothis ignores the fact that strings can contain a null character. for example, this is a valid Go program: package main import "fmt" func main() { fmt.Printf("%q\n", "hello \x00 world") }
- throwawaymaths 3y agoI'm not sure what your point is. We're talking about c null terminated strings, not go.