by Arnd Hafner
Released on July 31, 2026,
under the CC BY-NC 4.0 license.
Last update on August 8, 2026.
As a morphographic writing system, Chinese characters represent units of meaning, sometimes entire words, sometimes meaningful parts of words, in short morphemes. For a long time, this led to the misunderstanding that Chinese characters are ideographs, i.e. graphic characters that directly indicate or express meaning[45]". Actually, the morphemes and not the characters carry the meaning even in Chinese. In this regard, Chinese characters essentially function like any other script worldwide: they are a set of symbols that represent or record language. Hence, an accurate description of Chinese characters cannot avoid the question of the connection between characters and language.
Traditionally, this connection was conceived as the trinity of shape (xíng形), sound/reading (yīn音), and meaning (yì義). According to Morioka Tomohiko,
“Chinese characters (kanji/hànzì漢字) can be considered morphographic characters as one character corresponds to one morpheme in classical Chinese. Hence, they were traditionally analyzed and arranged based on the pairing of or the correspondence between various elements related to shape, sound and meaning."[46]"
It is noteworthy that Morioka limits his argument to morphemes in “classical Chinese古典中國語”. Even though the definition of “classical Chinese” is not readily apparent, a reference to language is indispensable for analyzing a set of symbols that is presumed to represent language. In my view, Morioka’s reference to “classical Chinese” could be seen as appealing to a common tradition shared by at least modern Japanese, Korean, and Chinese. In Japanese as well as in Korean, there still exists a set of readings for Chinese characters that can, just like modern Chinese readings, be more or less accurately traced back to the systematic and standardized description of medieval Chinese sounds in the Guǎngyùn廣韻, a comprehensive Chinese rhyme dictionary compiled in the early 11th century[47]". Most of the meanings of morphemes bound to these “common” readings in modern Japanese and Korean dictionaries derive from or are strongly influenced by the standardized understanding of Chinese characters displayed in the Kāngxī zìdiǎn康熙字典[48]", a standard dictionary of Chinese characters edited under the patronage of the fourth ruler of the Manchurian Qing empire.
To my understanding, this common tradition, though fuzzy as it might be, was the foundation that facilitated the integration of several national codes representing Chinese characters into one single character code, the Unicode Standard. The Unicode Consortium refers to this common ground as follows:
“For many centuries, written Chinese was the accepted written standard throughout East Asia. The influence of the Chinese language and its written form on the modern East Asian languages is similar to the influence of Latin on the vocabulary and written forms of languages in the West. …… The evolution of character shapes and semantic drift over the centuries has resulted in changes to the original forms and meanings. …… Still, the identical appearance and similarities in meaning are dramatic and more than justify the concept of a unified Han script that transcends language.”[49]"
Speaking of “transcending languages” is, of course, highly problematic. It contradicts both Morioka’s limitation to “classical Chinese” and the consortium’s own understanding of characters:
“The difference between identifying a character and rendering it on screen or paper is crucial to understanding the Unicode Standard’s role in text processing. The character identified by a Unicode code point is an abstract entity, such as “LATIN CAPITAL LETTER A” or “BENGALI DIGIT FIVE”. The mark made on screen or paper, called a glyph, is a visual representation of the character.”[50]"
Without question, the distinction between characters and glyphs does indeed clarify the dependency of characters on language. The expressions “CAPITAL LETTER A” and the “DIGIT FIVE” in the quotation would not make any clear sense without the modifications “LATIN” and “BENGALI”. The same glyph can represent a completely different character in a different language context.
The Unicode Consortium tries to define Chinese characters, or “unified Han ideographs”, by shape and meaning, excluding the reading element. Considering the shared phonetic origins, this exclusion is neither necessary nor fundamentally flawed. This phonetic feature is inherent to the common tradition that underlies the integration of various national standards. The defining readings of the characters are interconnected through their relation to the Guǎngyùn sounds. Omitting this inherent connection only makes the definition of a character overly simplistic.
Based on a three-dimensional conceptual model, the Unicode Consortium defines characters by shape and meaning. The model assigns “written elements” three primary attributes: semantic (meaning, function), abstract shape (general form), and actual shape (instantiated, typeface form), as depicted in figure 27[51]". Regarding the three dimensions, the following is said:
The semantic attribute (represented along the X axis) distinguishes characters by meaning and usage. Distinctions are made between entirely unrelated characters such as 澤(marsh) and 機(machine) as well as extensions or borrowings beyond the original semantic cluster such as 机1 (a phonetic borrowing used as a simplified form of 機) and 机2 (table, the original meaning).
The abstract shape attribute (the Y axis) distinguishes the variant forms of a single character with a single semantic attribute (that is, a character with a single position on the X axis).
……
Z-axis typeface and stylistic differences are generally ignored ……[52]"
The explanation about the semantic attribute sounds as if the Unicode Standard would distinguish 机1 and 机2, which, of course, is not true. The haracters 机1 and 机2 share the same code point “673A”. Taking practical considerations into account, this is reasonable for the Unicode Standard[53]", but it means that Unicode discerns “writing elements” merely based on abstract shapes as all characters that differ in meaning but overlap in shape are unified into one single code point. In other words, the Unicode Consortium merely provides a list of abstract shapes or glyphs, leaving the question of characters open.
From a linguistic perspective, the focus of character analysis within the context of a morphographic script lies with the morphemes. These morphemes consist of sound and meaning, often concealed behind constantly shifting shapes. This perspective is literally the opposite of the shape-centered approach of the Unicode standard.
With regard to this perspective, Qiú Xīguī裘錫圭 (1935-2025)’s argument that the Chinese script ought not to be understood as a morphemic or morphographic script might be instructive. In the Chapter Two “The Nature of Chinese Characters” of his groundbreaking book “Chinese Writing”, Qiú strives to develop his own comprehensive perspective on the nature of Chinese characters. In his final argument, he offers additional evidence for its accuracy by comparing it with morphemic views on Chinese characters.
According to Qiú, the Chinese script is a set of characters that employ three different types of symbols: semantic symbols (=yìfǔ意符), phonetic symbols (=yīnfǔ音符), and signs (=jìhào記號)[54]". Regarding “signs”, Qiú argues that:
“(I)n the process of the development of the Chinese script, due to changes in graphic shape, phonology and semantics, many semantic and phonetic symbols lost their semantic and phonetic functions and became signs.”[55]"
He provides the definition of “signs” as “arbitrarily fixed symbols that had no inherent relationship to the object represented.”[56]" The wide use of signs is considered a rather late phenomenon. Hence, Qiú assumes the development of the Chinese script from a strictly “semanto-phonetic script” to a “late semanto-phonetic script” that incorporates considerable amounts of signs.[57]"
For an accurate understanding, it is indispensable to mention that Qiú distinguishes two different levels of symbols as follows:
“Characters are symbols representing language. But characters themselves taken as language representing symbols and the symbols used by the characters are concepts belonging to two different levels.”[58]"
The terms “character” and “symbol” are renderings of the Chinese words “wénzì文字” and “fúhào符號”, respectively. In this sentence, Qiú still uses the term “symbol” in a two-fold manner, but two paragraphs later he makes an important clarification: “For the sake of clarity, the symbols used by a character will be called graphic symbols (=zìfǔ字符).”[59]" Qiú dedicates himself throughout the book to adhere to this distinction. Characters are symbols representing language whilst themselves deploying graphic symbols. The “semantic symbols (=yìfǔ意符)” and “phonetic symbols (=yīnfǔ音符)”, which are the core pillars of his theoretical framework of a “semanto-phonetic script”, are, of course, exponents of “graphic symbols”. They help characters fulfill their role as symbols representing language.
Raising some of Qiú’s examples may foster better understanding. Qiú exploits the character “huā花 (flower)” to illustrate the difference of the two levels of symbols.
“(T)he Chinese character “花” huā is a symbol of the Chinese word{花} "flower"; "艹" (the grass component, originally written “艸”, the old graph for căo “grass”) and “化” huà are the symbols used in writing the character “花”(a phonogram: “艹” is the signific and “化” the phonetic).”[60]"
In this case, the character is a “composite character (=hétǐ zì合體字)”. This means the character is composed of two components, each being understood as a graphic symbol. Hence, the distinction between the character as a symbol representing a Chinese word and the graphic symbols comprising this character is easy to comprehend.
The matter gets more intricate when the discussion turns to “non-composite characters (=dútǐ zì獨體字)”. Qiú presents the example of the character “rì日” , involving its very early writing form:
“(T)he ancient character
(日)when viewed as a symbol of the Chinese word{日} rì “sun” is a graph possessing both a meaning and a sound; looked at as a symbol for the character “日” rì, then it is merely a pictographic symbol with a shape resembling the sun.“[61]"
For Qiú, even “non-composite characters” are “composed” of graphic symbols. The difference between “composite characters” lies solely in the number of components[62]". The former has two or more components, whereas the latter has only one. In the case of the character “r ì日”,
or 日 is a pictographic symbol with a semantic load deriving from its sun-like shape[63]"; only when this graphic symbol composes the character “日” representing the word “rì”, it acquires a reading. We could also conclude that the word level includes both reading and meaning; the graphic symbol level often includes only one of both.
Interestingly, loangraphs (=jiǎjiè zì假借字) can make use of entire characters as phonetic symbols, no question whether the borrowed characters are “non-composite characters” or “composite characters”. With regard to the ancient character “
(= 其)” and the modern character “花“, Qiú explains as follows:
“(W)hen the semantograph
standing for {箕} jī "winnowing basket” was used to represent the modal particle{其} qí, the two words{箕} and {其} were not at all related semantically. Another example is the modern use of the phonogram{花} huā "flower" to represent the verb{花} huā "to spend (money)." While both of them are pronounced huā, they are totally unrelated semantically. Therefore, even though
was originally a semantograph and “花” was a phonogram, when they are borrowed to write the modal particle{其} and the verb {花}, they function purely as phonetic symbols.”[64]"
What is the point of taking the pain to discern these two levels of symbols? In my view, Qiú did not clearly specify the objective of this distinction, but his discussion on the concept of morphemic or morphographic script clarifies his stance.
Qiú’s argument about the relationship between Chinese characters and morphemes is twofolded. On the one hand, he admits the existence of morphemic features in the Chinese script, taking morphemic theories of the Chinese script as perspectives fundamentally different from his own approach, which is focused on “graphic symbols”. On the other hand, he argues that, even when changing the perspective, morphemic theories should be amended to “morphemo-syllabic script”, pointing to words that are comprised of two or more syllables and cannot be broken down into one-syllable morphemes. In these cases, no single character in the string of characters expressing the pertinent word can represent an entire morpheme; any single character merely refers to a phonetic fragment of the morpheme, a syllable. In other words, his second argument is centered on the phonetic function of loangraphs, whilst admitting that other types of characters, such as semantographs, etc., could be considered morphemic graphs. Let us first see how he puts this second argument:
“(N)on-composite, quasi-composite and composite semantographs, as well as sign graphs and semi-sign, semi-semantographs, can all be viewed as morphemic graphs. However, the phonetic symbols of the Chinese script, although they are all written with ready-made graphs which were originally morphemic symbols, ought to be viewed as symbols expressing syllables; ……In the case of those loangraphs which record transcribed foreign words of two or more syllables, it is perfectly clear that they express the nature of the syllabic structure of morphemes. For example, the four characters "達魯花赤" dálǔhuāchì which were borrowed in the Yuan dynasty to write the Mongolian title darugaci "governor, keeper of the seal," are clearly all used as syllabograms. Loangraphs used to write native Chinese binomes like "倉庚" cānggéng "name of a bird" and "猶豫" yóuyù "hesitate"(see Sec. 9. 3)also clearly are employed to express syllabic structure.”[65]
The first argument of difference in perspective highlights the objective behind his emphasize on graphic symbols:
“The terms morphemo-syllabic script and semanto-phonetic script(or semanto-phonetic-sign script)are names given to the Chinese script looked at from different points of view. The former focuses on the level of linguistic structure represented by graphic symbols; the latter focuses on the semantographic or phonographic functions of graphic symbols. These two terms can coexist.[66]"
In my view, the focus on the functions of graphic symbols that comprise the characters is nothing else than an inquiry into the question of how Chinese characters have been created or deployed for the representation of certain features of language. For instance, the rationale behind the character “日” representing the word “rì (= sun)” is explained as because the graphic symbol “
” is a pictorial representation of the sun, establishing a semantic link to the word “rì” but no phonetic link. In a similar manner, the character “花” can represent the word “huā (=flower)” because of the semantic link by the signific “艹/艸” and the phonetic link by the phonetic “化”. The same character “花” can represent the word “huā (=to spend)” because of the phonetic link established by the loan of the entire character as a single phonetic symbol.
These inquiries bear some resemblance to traditional Chinese philology, yet Qiú’s unwavering focus on “words” and “morphemes” strikes me as remarkably innovative. To be precise, Qiú‘s methodology fully integrates the morphemic theory, while also attempting to address traditional inquiries. Within this two-fold approach, the focus on the linguistic layer of “words” and “morphemes” clearly marks the linguistic turn in Chinese paleography, which had been carefully prepared by Táng Lán唐蘭 (1901-1979) and then accomplished by Qiú[67]". It marks a definite departure from the worn-out notion of ideographs that directly convey meaning, as broadly seen in the work of Shirakawa Shizuka白川静 (1910-2006) or still present even in Chén Mèngjiā陳夢家 (1911-1966)’s Zhōngguó wénzì xué中國文字學.
Having experienced the linguistic turn, the focus on “words” and “morphemes” might sound even banal to us. But regarding the Chinese script, this can be an extremely challenging journey as constantly changing shapes easily obscure the linguistic layer from our sight. Qiú‘s brilliant analysis on the semantic connection of the character “shǐ矢 (usually ‘arrow’)” in Ode 45.1 “之死矢靡它” to the well-known character “chén陳 (=to display, to state) demonstrates how subtle linguistic connections often go unnoticed[68]". At the same time, this linguistic layer, the unity of sound and meaning, is just what the character definition of the Unicode is lacking. This is also what we must focus on for a more historical comprehension of Chinese characters.
The greatest challenge for a paleographic understanding of characters is the reconciliation of the unity of sound and meaning, i.e. words or morphemes, with the shape of characters. Qiú (1988) devotes three entire chapters to explore the intricate relationship between words or morphemes and the character shapes: Chapter 10 “Allographs, Homographs, and Synonymic Interchange”, chapter 11 “Graphic Differentiation and Consolidation”, and, finally, chapter 12 “The Intricate Relationship Between Graphic Form and Sound and Meaning”. If we take the last chapter 13 “The Systematization and Simplification of Chinese Script” as a discussion on the endeavors in the People’s Republic of China to unravel this intricate relationship, four of the entire thirteen chapters are dedicated to this topic.
This, of course, is far too complex to unwind here. It must suffice to say that the bulk of discussions continues to oscillate between shape and morphemes; different shapes are brought together based on the unity of sound and meaning, and identical shapes are discerned based on the difference in the sound and/or meaning of the words they represent. Only the historical-linguistic context varies, and with it Qiú’s interpretation of the specific relationships. Qiú’s strong interest in how certain characters happen to represent certain words or morphemes further increases the complexity of the discussion as well as the possibilities of disagreement on certain interpretations.
Boiling it down to a phenomenological level, the shapes of Chinese characters change much faster in time and space than the morphemes they represent. The same morpheme can be represented by different shapes at different times or locations. A crucial role of paleography is to trace and keep records of these historical changes. For instance, Ōnishi Katsuya gathered the different writings of the word expressing the verb “to cast” on weaponry during the Warring States period (453/403-221 BCE). Except for the State of Yān燕 using a word cognate to the modern “wéi爲”, the others used cognates of the modern “zào造”. The various shapes of the character “zào造” are displayed in table 7[69]":
Qín秦 |
Qí齊 |
Chǔ楚 |
Wèi魏・Zhào趙 |
Hàn韓 |
Sòng宋 |
|
|---|---|---|---|---|---|---|
| 造 | ○ | ○ | ||||
| 艁, 鋯 | ○ | |||||
| 𢽍, 郜, 俈 | ○ | |||||
| 𬥘 | ○ | |||||
| 𫿜 | ○ | ○ | ||||
| 𪯓, 棗 | ○ |
These amount to ten allographs for a word that is still in use in modern Chinese. To my own surprise, I could find all ten being included within the Unicode character set, which, without a doubt, is a huge practical benefit. Given their distinct shapes, Unicode treats them as separate characters, which makes the Chinese script set significantly more complicated and convoluted than strictly needed from a linguistic perspective.
Shào Yǒnghǎi 邵永海 developed a concept that from the very beginning has been minted in order to decrease such complexities and intricacies. This concept is called “zìwèi字位” in Chinese and could be translated as “character position”. “Position” appears to refer to a fixed position within the morphemic space represented by the Chinese script.
Like Qiú Xīguī, Shào focuses on words or morphemes, but he stops at the stage of identifying the morphemes that are hidden behind the shapes and bundles identical morphemes into “character positions”, irrelevant to the variety of shapes that might represent them. In other words, Shào turns the perspective back to the morphemic writing aspect in order to reach a clearer picture of language features that are represented by characters but also somewhat concealed by the variety of their shapes.
Shào displays strong interest in the digital processing of characters in a linguistically meaningful manner. He points out repeatedly that different normative frameworks actually comprise unnecessary obstacles to this objective. For him, the impedimentary effect of normative frameworks can be witnessed in modern encoding of characters as well as in the historical changes of character shape throughout the three thousand and more years of Chinese writing history. One example of the impediments of modern character norms involves the character “guī龜 (tortoise, turtle)”, whose “formal/correct shape (=zhèngtǐ整體)” varies among East Asian states and regions:[70]"
![]() |
![]() |
![]() |
![]() |
![]() |
|---|---|---|---|---|
| China | Taiwan | Hong Kong | Japan | Korea |
Unicode reserves three code points for these five normative character shapes. “U+F907 /63751” for the Hong Kong version, “U+F908/ 63752” for the Korean version, and “U+9F9C/ 40860” for the other three shapes, which have been displayed in this paper with the help of different fonts. Similar discrepancies in normative requirements for character shapes are pointed out for the following string of allographs:
裏-裡;黃-黄;涙-泪;恥-耻;幫-幇;牀-床;昇-陞〔升〕;穽-阱;徑-逕;舉-擧;貍-狸;犁-犂;弔-吊;匯-滙;群-羣;剋-尅;妬-妒;荊-荆;够-夠;諮-咨;眾-衆;歎-嘆;啟-啓;髕-臏;腳-脚;汙-污
Shào refers to these examples and states that “character codes …… do not encompass information that connects allographs or other interchangeable graphs.” According to him, these “differences in character shape directly influence the stability and authority of digitized dictionary contents.”[71]" These problems are well-known, and many search engines are already capable of reconnecting allographs that have been separated by different code points within the Unicode standard. However, Shào claims the need for a linguistic solution to such problems that also includes the longer historical development of the Chinese script. In his view, the definition of “character positions” by common reading and meaning is a more reliable method of overcoming the intricacies caused by endless shape changes:
“After joining, every character position takes charge for multiple shapes, forming a character family. One character family holds the different shapes of one character position, such as Oracle bone inscription shapes, bronze inscription shapes, seal script shape, clerical script shapes, and all sorts of allographs and synonymic interchangeables. Since an effective connection is established within the character families, search results include all shapes of the same character position, no question which shape is used for the query.”[72]"
Relinking different character shapes through their representation of common morphemes appears to be the main objective behind Shào’s proposal of the concept of “character position”. In this regard, the concept could facilitate the handling of paleographic phenomena such as the various shapes of the character “zào造” gathered by Ōnishi.
However, the concept can also be utilized in the opposite direction, i.e. it can also help discern homographs and other shape overlapping phenomena. As Shào’s argument does not reach this far, the author will try to apply the concept to some overlapping phenomena frequently seen in paleographic material from the Qin and Han period.
The glyph “灋” is commonly known as the ancient form of the character “fǎ法 (=penalty, law)”. Xǔ Shèn許慎 (30/58 – 121/147 CE)’s Shuō wén jiě zì説文解字 records it as the seal script shape, stating that the current writing simplifies it to “法”. In legal literature from the late Warring States period, the same glyph represents two different words cognate to modern “fǎ法” and “fèi廢” respectively. The bamboo slips excavated from Tomb 11 in Shuihudi, Hubei province, entail an official notice by the prefect of the Nán province. On the slips 3 to 4 of the document, the same glyph is used for both “fǎ法” and “fèi廢”:
Now, the statutes on penalties are furnished. But the officials and the ordinary people do not adhere to it, …, that amounts to abandoning the clearly stipulated penal law.
今灋(法)律令已具矣,而吏民莫用,…,是卽灋(廢)主之明灋(法)殹(也)。(睡虎地11号秦墓竹簡『語書』簡003-004)
Be aware that the original character “廢” is, based on the semantic difference, regularized to “fǎ法” and “fèi廢” respectively.
A wooden tablet excavated from the Liye-site in Longshan, Hunan province, shows that, as a part of the character unification under the First Emperor of the Qin, the function of this glyph was split into two; the function of representing the word cognate to modern “fèi (=abandon, abolish)” was delegated to the newly created glyph “廢”. Figure 28 displays an image of the entire tablet. Here, we only cite three items:
After the script unification, “fǎ灋” and “fèi廢” distinctly refer to two separate words. Even before the unification, the semantic difference is significant[73]". It would hardly make sense to combine all occurrences of the glyph “灋” into a group of characters connected to the word “fǎ灋”. A more reasonable option would be to draw the line based on the semantic and phonetic difference that we asserted in the form of different regularizations in the transcription. That would mean that some instances of “灋” gather the group of characters connected to the modern character “fǎ法”, while others join the group of “fèi廢”. This grouping could be easily represented by Shào’s concept of “character position”. “Fǎ法” and “fèi廢” represent two different character positions. Characters that can be associated with either position are considered members of the same character family, regardless of which glyph was employed in the original text.
The next example is even more intricate. According to Shuō wén jiě zì, the character “zuì最 (=most, summary, etc.)” is composed of the graphic symbols “mào冃 (=cap)” and “qǔ取 (to take)”[74]". In popular writing, the two horizontal strokes inside the component “mào冃” are often abbreviated, reducing the entire character to the shape of “冣”. This makes this character a homograph to the character “jù冣 (to accumulate)”, which according to Shuō wén jiě zì is comprised of “mì冖 (cloth cover)” and “qǔ取”[75]". In his annotation to Shuō wén jiě zì, Duàn Yùcái段玉裁 (1735-1815) argues that the reading and meaning of “jù冣” are identical to “jù聚”. Hence, the character is generally considered an allograph to “jù聚”[76]".
The tendency of omitting the two horizontal strokes inside the “mào冃” component of “zuì最” has already been observed within the clerical Qin script at the end of the Warring States period, and has caused some scholars to argue that the characters “zuì最” and “jù冣” were merged during this time[77]". Indeed, the writing of these two characters is confusing, but they are not truly merged. They are distinguished in a relative sense from each other by writing different numbers of vertical strokes. Table 9 gathers all instances of “zuì最” and “jù冣” that can be found in the texts from Tomb M11 at the Shuihudi-Site in Yunmeng, Hubei. Comparing these instances, the shape “
” is used in the Yǔshū語書 and Qínlǜ shíbā zhǒng秦律十八種 to represent the character “zuì最”, and in Rìshū jiǎ zhǒng日書甲種, it represents the character “jù冣”. This phenomenon indeed gives the impression that the two characters are at least confused. However, there is one instance of the character “zuì最” in the Rìshū jiǎ zhǒng日書甲種, which is written as “
”, adding one more vertical stroke in comparison to the character “jù冣” in the shape of “
”. In other words, the scribe of the text clearly felt the necessity to distinguish these two characters, and did so by increasing or decreasing the number of vertical strokes.
Yǔshū語書 |
Qínlǜ shíbā zhǒng秦律十八種 |
Rìshū jiǎ zhǒng日書甲種 |
|
|---|---|---|---|
| Zuì最 |
013 |
013 014 |
056背參 |
| Jù冣 | - | - |
005正貳 / 015背壹 / 016背壹 |
In this case, Shào’s concept of “character position” could be helpful again. No question whether the shape is written as “冣”, “㝡”, “
”, or “
”, characters that can be identified as representing the same word “zuì最” should be associated with one character family, sharing the same character position within the Chinese writing system; characters representing “jù冣” belong to a distinct character family, occupying a completely different position within the writing system.
Partly overlapping in component shape is also observed quite frequently in the clerical script during the Qin and the Han. Well known exponents are the components “聿” and “隶”, “彖” and “录”, “辶” and “廴”. Some transcribers view these interchanges of components as mistakes. If the character “dài逮” is written with the components “聿” and “廴”, they transcribe “建〔逮〕”. But these writing habits apparently have not been treated as mistakes by contemporaries; it was their customary way of writing. They were aware of the words or morphemes their script represented. Hence, our task should be to identify these words or morphemes and not to correct their writing habits. This is possible by accumulating different shapes to character families based on their position in the writing system, or, plainly, the words they represent.
In the previous sections, the author tried to outline what it means to encode characters rather than glyphs. Now, we must find a way to express these ideas using TEI- elements. Actually, the <charDecl>, a member of the gaiji module, already distinguishes the elements <glyph> and <char>. Once again, the author is compelled to acknowledge that the TEI-guidelines were prepared well, yet we users easily overlooked what this difference means in the context of Chinese writing. Below, I will attempt to employ the paleographic examples provided in the previous section to delineate how to encode Chinese characters using TEI.
First, each character possesses an abstract shape. The <glyph> element can be utilized to document the shape. The basic syntax should be one <glyph> element for one shape, with an @xml:id attribute assigning an identifier to it. Within the <glyph> element, the <mapping> element could take charge of the task of defining shapes. This would look like this:
<glyph xml:id="glyph-ID"><mapping>shape</mapping></glyph>
Secondly, each character represents a word or morpheme. The word or morpheme can be identified through the synthesis of its sound and meaning. For the description of these two features, we primarily utilize the readings of characters as depicted in the Guǎngyùn廣韻, alongside with the gloss given in the same book. If a character is not documented in Guǎngyùn, we need to resort to the readings found in later dictionaries, but as a rule of thumb, a character not recorded in the Guǎngyùn should raise alarm within the context of paleography. With high probability, our interpretation of the character is inaccurate. Hence, this paper will not go into detail in this regard. For the glosses, reinforcement should be sought from early annotations and dictionaries. Glosses later than the 3rd century CE should be again handled with precaution. Judging from excavated language material, the Chinese language changed dramatically during the two Han dynasties. Accurate knowledge about early Chinese that still circulated in academic circles got lost during the 3rd century.
This character's linguistic representation aligns with the unit that combines sound and meaning. This unit gives the character a fixed position in the script system, a “character position” according to Shào’s theory. For illustrative purposes, we could give a modern representation of this position in the form of a character. The overall syntax for a character definition would then look like this:
<char xml:id="char-ID">
<mapping type="sound guǎngyùn">fǎnqiè反切</mapping>
<mapping type="gloss guǎngyùn ">a gloss</mapping>
<mapping type="early gloss">a gloss</mapping>
<mapping type="character position">a modern character</mapping>
<note>A note if necessary</note>
</char>
Lastly, we would need to search for a method to link characters within the main text to the glyph and character definitions. Within the <charDecl> element, each <glyph> and <char> element carries a unique ID, stored in the @xml:id attribute. These can serve as reference benchmarks. The characters in the main text are denoted by <g> elements, as outlined in chapter 3. <g> elements can carry @ref and @ana attributes. These facilitate links to the glyph and character definitions. While glyph references can be established automatically, as demonstrated previously, character references still need scholarly scrutiny. Hence, it appears to be reasonable to use the @ref attribute for glyph reference and the @ana attribute for character reference. A character within the main text would then look like this:
<g ref="#glyph-ID" ana="#char-ID">
For demonstration, we will encode the few paleographic instances mentioned in the previous section. The glyph “灋” that represents the word “fǎ法” could be distinguished from the same glyph representing the word “fèi廢” as follows.
Glyph definition:
<glyph xml:id="U28747">
<mapping>灋</mapping>
</glyph>
Character definitions:
<char xml:id="C0001">
<mapping type=" sound guǎngyùn">方乏切</mapping>
<mapping type="gloss guǎngyùn">則也</mapping>
<mapping type="gloss Shuō wén jiě zì">刑</mapping>
<mapping type="gloss Yùpiān(玉篇廌部)">刑</mapping>
<mapping type="gloss Yùpiān(玉篇水部)">法令</mapping>
<mapping type="position">法</mapping>
<note>……</note>
</char>
<char xml:id="C0002">
<mapping type=" sound guǎngyùn">方肺切</mapping>
<mapping type="gloss guǎngyùn">止也</mapping>
<mapping type="gloss Zhōulǐ Zhèngxuán Zhù周禮鄭玄注">猶退也</mapping>
<mapping type="gloss Yùpiān(玉篇广部)">退</mapping>
<mapping type="position">廢</mapping>
<note>……</note>
</char>
Character in main text:
“Fǎ灋(法)” = <g ref="#U28747" ana="#C0001">
“Fèi灋(廢)” = <g ref="#U28747" ana="#C0002">
Given that most glyph definitions have been standardized by Unicode, there's no need to reinvent the wheel. Therefore, the example used the decimal codepoint of “灋” (=28747) as glyph-ID. In fact, there is no necessity for glyph definition at all if the glyph is already part of the Unicode Standard. As every script can automatically encode and decode Unicode characters, the @ref attribute can use the code point directly as a reference. Only for non-Unicode characters, a glyph definition would indeed be necessary. Below follows an example of glyph definition for the non-Unicode shape of the characters “zuì最” and “jù冣” mentioned in the previous section:
:
<glyph xml:id="g0001">
<mapping type="IDS">⿳宀一取</mapping>
<graphic type="lìdìng隸定" url="gaiji0001.jpg"/>
</glyph>
:
<glyph xml:id="g0002">
<mapping type="IDS">⿳宀二取</mapping>
<graphic type="lìdìng隸定" url="gaiji0002.jpg"/>
</glyph>
Supposing that the character positions of “zuì最” and “jù冣” were assigned the character IDs “C0003” and “C0004” respectively, characters in the main text could reference this glyph definition as follows:
“zuì
(最)” = <g ref="#g0002" ana="#C0003">
“zuì
(最)” = <g ref="#g0001 " ana="#C0003">
“jù
(冣)” = <g ref="#g0001" ana="#C0004">
Unfortunately, the author could not yet implement the ideas presented in this final chapter yet. The character definitions as well as the @ana attributes need to be created manually. This will require a considerable amount of time. The author looks forward to being able to showcase concrete outcomes at an upcoming meeting. Suggestions and criticism that can guide the author towards achieving this goal are highly valued.