3 ms·
> Interning is a trade-off: you get decreased memory usage (if you use a lot of long-lived duplicated strings) and faster string comparisons (for pairs of strin
by uxcn 11y ago
> Interning is a trade-off: you get decreased memory usage (if you use a lot of long-lived duplicated strings) and faster string comparisons (for pairs of strings that you do intern), at the cost of extra work to create the interned strings.
For garbage collected languages the benefit isn't only memory consumption, it's performance, since using canonical objects eliminates the additional allocations and garbage collections.
The string comparison argument is a bit of a dubious one though. Comparing length is only an integer comparison, and depending on the architecture, you can compare up to eight (or more) characters per cycle.
- Someone 11y ago"since using canonical objects eliminates the additional allocations and garbage collections" That requires a quite advanced compiler. Looking at Java, the process is: - create a new string in some way. - call String.intern() to create or retrieve the interned string with the same contents. If you do String si = (s+t).intern(); or String si = s.replace('a', 'b').intern(); it would take quite a compiler to prevent the creation of an intermediate string. You could have every function return an iterator over the characters that would end up in the string, iterate over that to check whether the string already is interned, and if not, iterate again to allocate a new string, but I think it would typically be cheaper to create it and let he young generation garbage collector collect it.
- uxcn 11y ago> it would take quite a compiler to prevent the creation of an intermediate string. You could have every function return an iterator over the characters that would end up in the string, iterate over that to check whether the string already is interned, and if not, iterate again to allocate a new string, but I think it would typically be cheaper to create it and let he young generation garbage collector collect it. I can't think of a language where you would always want to canonicalize strings at a global scope. For example, consider the case where you have a large number of threads and cores. Unless the strings are explicitly allocated on their own cache lines, any thread that references a string now has to worry about false sharing. You would also need to worry about the contention on the canonical store. At a user level, in languages like Java, there's generally no reason to create any intermediate string if you're already reading from a direct byte buffer. This covers a fairly large set of use cases. There may be other techniques considering Java supports scalar interpolation now, but direct byte buffers have been the most effective in my experience.