4 ms·
No, it's different between Chinese and Thai. Lexing is very clear in Chinese. It's never the case that you look at a Chinese sentence and don't know where a ch
by eric-hu 3y ago
No, it's different between Chinese and Thai.
Lexing is very clear in Chinese. It's never the case that you look at a Chinese sentence and don't know where a character ends and another begins. Take this sentence in both languages: "good morning, how are you"
早安,你好吗
This sentence clearly has "spaces" and I'm pretty sure any person illiterate in Chinese could tell you there are 5 characters / words. Technically the third character is composed of 人 and 尔 but I don't know that anyone, even kids or beginners, would mistake those as _not_ going together.
สวัสดีตอนเช้าคุณเป็นอย่างไรบ้าง
In contrast, Thai is as you say: lexing and parsing bleed together. There are 7 words in this sentence, but you need to lex the 10 syllables and run them through your mental dictionary to recognize the possible words they could be. My Thai is very limited, but there are examples of sentences out there that actually have multiple valid readings with different semantic meanings, depending on how you group sounds together.
- Anon1096 3y ago早安 is made up of 2 characters but is a single word. If you fall into the trap of thinking 1 character = 1 word, you won't understand a thing. In this case you'd have thought it meant "early safe" instead of "good morning".
- eric-hu 3y agoOkay, you make a good point. Let's look back at the GGP's comment though: > Same with Chinese language, thus lexing and parsing requires knowing many more words than in languages with spaces between words. In English, can you get away with knowing the meaning of "good" and "morning" and not "good morning", and know that I'm greeting you instead of commenting on the quality of this morning?
- likpok 3y agoGood morning is a bad example because it has a colloquial meaning that is a least a little idiomatic. Most other words/phrases in English don’t have this effect, while many Chinese words are like 早安. 了解, for example, can’t even be pronounced without correctly parsing the word.
- eric-hu 3y agoOkay, I concede that I may have forgotten that Chinese has its exceptions too. 了解 is indeed a good example. There are plenty in English though. Even with context, sometimes I have to really pause and think whether to pronounce read as red or reed (I read it just fine, I read English just fine). Where I've had pain specifically with Thai is that I can't even know where a syllable begins and ends until I read a few "syllables" together and decide whether some vowels go with the consonant in front or behind it, and whether some an -ar should be pronounced as an -aan.
- adastra22 3y agoChinese has pretty regular rules about grouping characters into words though, as most compounds are 2-characters, or a 4-character idiomatic phrase. Even if I know only half the characters in a sentence, I can usually guess the word boundaries correctly. It's not 100% reliable, but good enough to avoid confusion.
- hnfong 3y agoI guess it really depends on "dialect". Try that with Cantonese :) As mentioned in another comment, single syllable words are much more common in Cantonese, and word combinations are much more "free" in the sense that there are a lot more ambiguity as to what counts as a "word" and what is merely two single-character-words idiomatically used together. There are also cases where grammatical constructs (and also foul words) are inserted in between a two-character word/idiomatic combo, and sometimes the characters are reversed, to the extent that it used to be a meme: https://evchk.fandom.com/zh/wiki/Y%E5%B7%B2x https://evchk.fandom.com/zh/wiki/Y%E5%B7%B2x It's gotten to a point where, after thinking about it for a couple years, I've come to believe that segmentation on Cantonese is a fool's errand... Of course, there's also classical Chinese where most of the time a character is a word.
- deadfoxygrandpa 3y agoi think you're on to something about cantonese, but it's also true of mandarin. segmentation of words in chinese in general seems inherently messier than segmentation in english. also look at stuff like abbreviations: is 北大 one word? is it an abbreviation for 北京大学 the same way Caltech is an abbreviation for california institute of technology? is it just two single character words, each of which is an abbreviation? i think its much less clear than english
- adastra22 3y agoThe same problem exists in Japanese FWIW, whose speakers like to make the same sorts of abbreviations despite not having a bisyllabic meter like Mandarin does. Japanese is somewhat helped by having multiple orthographies, however.