6 ms·
> Latin 1 standard is still in widespread inside some systems (such as browsers) That doesn't seem to be correct. UTF-8 is used by 98% of all the websites. I a
by ko27 3y ago
> Latin 1 standard is still in widespread inside some systems (such as browsers)
That doesn't seem to be correct. UTF-8 is used by 98% of all the websites. I am not sure if it's even worth the trouble for libraries to implement this algorithm, since Latin-1 encoding is being phased out.
https://w3techs.com/technologies/details/en-utf8 https://w3techs.com/technologies/details/en-utf8
- fulafel 3y agoIt's the default HTTP character set. It's not clear whether the above stat page is about what charsets are explicitly specified. Also headers, mostly relevant for header values, are I think ISO-8859-1.
- ko27 3y agoSince HTML5 UTF-8 is the default charset. And for headers, they are parsed as ASCII encoded in almost all cases although ISO-8859-1 is supported.
- fulafel 3y agoI tried to find confirmation of this but found only: https://html.spec.whatwg.org/multipage/semantics.html#charset https://html.spec.whatwg.org/multipage/semantics.html#charse... > The Encoding standard requires use of the UTF-8 character encoding and requires use of the "utf-8" encoding label to identify it. Those Sounds to me like it tells you that you have to explicitly declare the charset as UTF-8, so you don't get the HTTP default of Latin-1. (But that's just one "living standard" not exactly synonymous with with HTML5 and it might change, or might have been different last week..)
- ko27 3y ago> so you don't get the HTTP default of Latin-1. That's not what your linked spec says. You can try it yourself, in any browser. If you omit the encoding the browser uses heuristics to guess, but it will always work if you write UTF-8 even without meta charset or encoding header.
- fulafel 3y agoI don't doubt browsers use heuristics. But spec-wise I think it's your turn to to provide a reference in favour of a utf-8-is-default interpretation :)
- ko27 3y agoNo it isn't. My original point is that Latin-1 is used very rarely on Internet and is being phased out. Now it's your turn to provide some references that a significant percentage of websites are omitting encoding (which is required by spec!) and using Latin-1. But if you insist, here is this quote: https://www.w3docs.com/learn-html/html-character-sets.html https://www.w3docs.com/learn-html/html-character-sets.html > UTF-8 is the default character encoding for HTML5. However, it was used to be different. ASCII was the character set before it. And the ISO-8859-1 was the default character set from HTML 2.0 till HTML 4.01. or another: https://www.dofactory.com/html/charset https://www.dofactory.com/html/charset > If a web page starts with <!DOCTYPE html> (which indicates HTML5), then the above meta tag is optional, because the default for HTML5 is UTF-8.
- bawolff 3y ago> My original point is that Latin-1 is used very rarely on Internet and is being phased out. Nobody disagrees with this, but this is a very different statement from what you said originally in regards to what the default is. Things can be phased out but still have the old default with no plan to change the default. Re other sources - how about citing the actual spec instead of sketchy websites that seem likely to have incorrect information.
- rhdunn 3y agoThe WHATWG HTML spec [1] has various heuristics it uses/specifies for detecting the character encoding. In point 8, it says an implementation may use heuristics to detect the encoding. It has a note which states: > The UTF-8 encoding has a highly detectable bit pattern. Files from the local file system that contain bytes with values greater than 0x7F which match the UTF-8 pattern are very likely to be UTF-8, while documents with byte sequences that do not match it are very likely not. When a user agent can examine the whole file, rather than just the preamble, detecting for UTF-8 specifically can be especially effective. In point 9, the implementation can return an implementation or user-defined encoding. Here, it suggests a locale-based default encoding, including windows-1252 for "en". As such, implementations may be capable of detecting/defaulting to UTF-8, but are equally likely to default to windows-1252, Shift_JIS, or other encoding. [1] https://html.spec.whatwg.org/#determining-the-character-encoding https://html.spec.whatwg.org/#determining-the-character-enco...
- rhdunn 3y agoBe aware that with the WHATWG Encoding specification [1], that says that latin1, ISO-8859-1, etc. are aliases of the windows-1252 encoding, not the proper latin1 encoding. As a result, browsers and operating systems will display those files differently! It also aliases the ASCII encoding to windows-1252. [1] https://encoding.spec.whatwg.org/#names-and-labels https://encoding.spec.whatwg.org/#names-and-labels
- pzmarzly 3y agoAnd yet HTTP/1.1 headers should be sent in Latin1 (is this fixed in HTTP/2 or HTTP/3?). And WebKit's JavaScriptCore has special handling for Latin1 strings in JS, for performance reasons I assume.
- ko27 3y ago> should be sent in Latin1 Do you have a source on that "should" part. Because the spec disagrees https://www.rfc-editor.org/rfc/rfc7230#section-3.2.4 https://www.rfc-editor.org/rfc/rfc7230#section-3.2.4: > Historically, HTTP has allowed field content with text in the ISO-8859-1 charset [ISO-8859-1], supporting other charsets only through use of [RFC2047] encoding. In practice, most HTTP header field values use only a subset of the US-ASCII charset [USASCII]. Newly defined header fields SHOULD limit their field values to US-ASCII octets. In practice and by spec, HTTP headers should be ASCII encoded.
- nicktelford 3y agoISO-8859-1 (aka. Latin-1) is a superset of ASCII, so all ASCII strings are also valid Latin-1 strings. The section you quoted actually suggests that implementations should support ISO-8859-1 to ensure compatibility with systems that use it.
- ko27 3y agoYou should read it again > Newly defined header fields SHOULD limit their field values to US-ASCII octets ASCII octets! That means you SHOULD NOT send Latin1 encoded headers. The opposite of what pzmarzly was saying. I don't disagree Latin-1 being a superset of ASCII or having backward compatibility in mind, but that's not relevant to my response.
- layer8 3y agoSHOULD is a recommendation, not a requirement, and it refers only to newly-defined header fields, not existing ones. The text implies that 8-bit characters in existing fields are to be interpreted as ISO-8859-1.
- TheRealPomax 3y agoOnly because those websites include `<meta charset="utf-8">`. Browsers don't use utf-8 unless you tell them to, so we tell them to. But there's an entire internet archive's worth of pages that don't tell them to.
- ko27 3y agoNot including charset="utf-8" doesn't mean that the website is not UTF-8. Do you have a source on a significant percentage of website being Latin-1 while omitting charset encoding? I don't believe that's the case. > Browsers don't use utf-8 unless you tell them to This is wrong. You can prove this very easily by creating a HTML file with UTF-8 text while omitting the charset. It will render correctly.
- TheRealPomax 3y agoAnswering your "do you have a source" question, yeah: "the entire history of the web prior to HTML5's release", which the internet has already forgotten is a rather recent thing (2008). And even then, it took a while for HTML5 to become the de facto format, because it took the majority of the web years before they'd changed over their tooling from HTML 4.01 to HTML5. > This is wrong. You can prove this very easily by creating a HTML file with UTF-8 text No, but I will create an HTML file with latin-1 text, because that's what we're discussing: HTML files that don't use UTF-8 (and so by definition don't contain UTF-8 either). While modern browsers will guess the encoding by examining the content, if you make an html file that just has plain text, then it won't magically convert it to UTF-8: create a file with `<html><head><title>encoding check</title></head><body><h1>Not much here, just plain text</h1><p>More text that's not special</p></body></html>` in it. Load it in your browser through an http server (e.g. `python -m http.server`), and then hit up the dev tools console and look at `document.characterSet`. Both firefox and chrome give me "windows-1252" on Windows, for which the "windows" part in the name is of course irrelevant; what matters is what it's not, which is that it's not UTF-8, because the content has nothing in it to warrant UTF-8.
- ko27 3y ago
- kannanvijayan 3y agoOne place I know where latin1 is still used is as an internal optimization in javascript engines. JS strings are composed of 16-bit values, but the vast majority of strings are ascii. So there's a motivation to store simpler strings using 1 byte per char. However, once that optimization has been decided, there's no point in leaving the high bit unused, so the engines keep optimized "1-byte char" strings as Latin1.
- HideousKojima 3y ago>So there's a motivation to store simpler strings using 1 byte per char. What advantage would this have over UTF-7, especially since the upper 128 characters wouldn't match their Unicode values?
- layer8 3y agoLatin1 does match the Unicode values (0-255).
- laurencerowe 3y ago> What advantage would this have over UTF-7, especially since the upper 128 characters wouldn't match their Unicode values? (I'm going to assume you mean UTF-8 here rather than UTF-7 since UTF-7 is not really useful for anything, it's jus a way to pack Unicode into only 7-bit ascii characters.) Fixed width string encodings like Latin-1 let you directly index to a particular character (code point) within a string without having to iterate from the beginning of the string. JavaScript was originally specified in terms of UCS-2 which is a 16 bit fixed width encoding as this was commonly used at the time in both Windows and Java. However there are more than 64k characters in all the world's languages so it eventually evolved to UTF-16 which allows for wide characters. However because of this history indexing into a JavaScript string gives you the 16-bit code unit which may be only part of a wide character. A string's length is defined in terms of 16-bit code units but iterating over a string gives you full characters. Using Latin-1 as an optimisation allows JavaScript to preserve the same semantics around indexing and length. While it does require translating 8 bit Latin-1 character codes to 16 bit code points, this can be done very quickly through a lookup table. This would not be possible with UTF-8 since it is not fixed width. EDIT: A lookup table may not be required. I was confused by new TextDecoder('latin1') actually using windows-1252. More modern languages just use UTF-8 everywhere because it uses less space on average and UTF-16 doesn't save you from having to deal with wide characters.
- syats 3y agoIn countries communicating in non-English languages which are written in the latin script, there is a very large use of Latin-1. Even when Latin-1 is "phased out", there are tons and tons of documents and databases encoded in Latin-1, not to mention millions of ill-configured terminals. I think it makes total sense to implement this.