TEI2026 Presentation Material

The Implementation of the TEI-Guidelines on Early Chinese Excavated Administrative Documents: With a Focus on Paleographic Issues

by Arnd Hafner

Released on July 31, 2026,

under the CC BY-NC 4.0 license.

Last update on August 8, 2026.

For a full text version, click on docx or pdf.

4. The encoding and redefinition of characters

4.1. What is a Chinese character?

As a morphographic writing system, Chinese characters represent units of meaning, sometimes entire words, sometimes meaningful parts of words, in short morphemes. For a long time, this led to the misunderstanding that Chinese characters are ideographs, i.e. graphic characters that directly indicate or express meaning[45]". Actually, the morphemes and not the characters carry the meaning even in Chinese. In this regard, Chinese characters essentially function like any other script worldwide: they are a set of symbols that represent or record language. Hence, an accurate description of Chinese characters cannot avoid the question of the connection between characters and language.

Traditionally, this connection was conceived as the trinity of shape (xíng形), sound/reading (yīn音), and meaning (yì義). According to Morioka Tomohiko,

“Chinese characters (kanji/hànzì漢字) can be considered morphographic characters as one character corresponds to one morpheme in classical Chinese. Hence, they were traditionally analyzed and arranged based on the pairing of or the correspondence between various elements related to shape, sound and meaning."[46]"

It is noteworthy that Morioka limits his argument to morphemes in “classical Chinese古典中國語”. Even though the definition of “classical Chinese” is not readily apparent, a reference to language is indispensable for analyzing a set of symbols that is presumed to represent language. In my view, Morioka’s reference to “classical Chinese” could be seen as appealing to a common tradition shared by at least modern Japanese, Korean, and Chinese. In Japanese as well as in Korean, there still exists a set of readings for Chinese characters that can, just like modern Chinese readings, be more or less accurately traced back to the systematic and standardized description of medieval Chinese sounds in the Guǎngyùn廣韻, a comprehensive Chinese rhyme dictionary compiled in the early 11th century[47]". Most of the meanings of morphemes bound to these “common” readings in modern Japanese and Korean dictionaries derive from or are strongly influenced by the standardized understanding of Chinese characters displayed in the Kāngxī zìdiǎn康熙字典[48]", a standard dictionary of Chinese characters edited under the patronage of the fourth ruler of the Manchurian Qing empire.

To my understanding, this common tradition, though fuzzy as it might be, was the foundation that facilitated the integration of several national codes representing Chinese characters into one single character code, the Unicode Standard. The Unicode Consortium refers to this common ground as follows:

“For many centuries, written Chinese was the accepted written standard throughout East Asia. The influence of the Chinese language and its written form on the modern East Asian languages is similar to the influence of Latin on the vocabulary and written forms of languages in the West. …… The evolution of character shapes and semantic drift over the centuries has resulted in changes to the original forms and meanings. …… Still, the identical appearance and similarities in meaning are dramatic and more than justify the concept of a unified Han script that transcends language.”[49]"

Speaking of “transcending languages” is, of course, highly problematic. It contradicts both Morioka’s limitation to “classical Chinese” and the consortium’s own understanding of characters:

“The difference between identifying a character and rendering it on screen or paper is crucial to understanding the Unicode Standard’s role in text processing. The character identified by a Unicode code point is an abstract entity, such as “LATIN CAPITAL LETTER A” or “BENGALI DIGIT FIVE”. The mark made on screen or paper, called a glyph, is a visual representation of the character.”[50]"

Without question, the distinction between characters and glyphs does indeed clarify the dependency of characters on language. The expressions “CAPITAL LETTER A” and the “DIGIT FIVE” in the quotation would not make any clear sense without the modifications “LATIN” and “BENGALI”. The same glyph can represent a completely different character in a different language context.

The Unicode Consortium tries to define Chinese characters, or “unified Han ideographs”, by shape and meaning, excluding the reading element. Considering the shared phonetic origins, this exclusion is neither necessary nor fundamentally flawed. This phonetic feature is inherent to the common tradition that underlies the integration of various national standards. The defining readings of the characters are interconnected through their relation to the Guǎngyùn sounds. Omitting this inherent connection only makes the definition of a character overly simplistic.

The Unicode Three-Dimensional Conceptual Model
Figure 27 The Unicode Three-Dimensional Conceptual Model

Based on a three-dimensional conceptual model, the Unicode Consortium defines characters by shape and meaning. The model assigns “written elements” three primary attributes: semantic (meaning, function), abstract shape (general form), and actual shape (instantiated, typeface form), as depicted in figure 27[51]". Regarding the three dimensions, the following is said:

The semantic attribute (represented along the X axis) distinguishes characters by meaning and usage. Distinctions are made between entirely unrelated characters such as 澤(marsh) and 機(machine) as well as extensions or borrowings beyond the original semantic cluster such as 机1 (a phonetic borrowing used as a simplified form of 機) and 机2 (table, the original meaning).

The abstract shape attribute (the Y axis) distinguishes the variant forms of a single character with a single semantic attribute (that is, a character with a single position on the X axis).

……

Z-axis typeface and stylistic differences are generally ignored ……[52]"

The explanation about the semantic attribute sounds as if the Unicode Standard would distinguish 机1 and 机2, which, of course, is not true. The haracters 机1 and 机2 share the same code point “673A”. Taking practical considerations into account, this is reasonable for the Unicode Standard[53]", but it means that Unicode discerns “writing elements” merely based on abstract shapes as all characters that differ in meaning but overlap in shape are unified into one single code point. In other words, the Unicode Consortium merely provides a list of abstract shapes or glyphs, leaving the question of characters open.

4.2. Defining Chinese characters by sound and meaning

From a linguistic perspective, the focus of character analysis within the context of a morphographic script lies with the morphemes. These morphemes consist of sound and meaning, often concealed behind constantly shifting shapes. This perspective is literally the opposite of the shape-centered approach of the Unicode standard.

With regard to this perspective, Qiú Xīguī裘錫圭 (1935-2025)’s argument that the Chinese script ought not to be understood as a morphemic or morphographic script might be instructive. In the Chapter Two “The Nature of Chinese Characters” of his groundbreaking book “Chinese Writing”, Qiú strives to develop his own comprehensive perspective on the nature of Chinese characters. In his final argument, he offers additional evidence for its accuracy by comparing it with morphemic views on Chinese characters.

According to Qiú, the Chinese script is a set of characters that employ three different types of symbols: semantic symbols (=yìfǔ意符), phonetic symbols (=yīnfǔ音符), and signs (=jìhào記號)[54]". Regarding “signs”, Qiú argues that:

“(I)n the process of the development of the Chinese script, due to changes in graphic shape, phonology and semantics, many semantic and phonetic symbols lost their semantic and phonetic functions and became signs.”[55]"

He provides the definition of “signs” as “arbitrarily fixed symbols that had no inherent relationship to the object represented.”[56]" The wide use of signs is considered a rather late phenomenon. Hence, Qiú assumes the development of the Chinese script from a strictly “semanto-phonetic script” to a “late semanto-phonetic script” that incorporates considerable amounts of signs.[57]"

For an accurate understanding, it is indispensable to mention that Qiú distinguishes two different levels of symbols as follows:

“Characters are symbols representing language. But characters themselves taken as language representing symbols and the symbols used by the characters are concepts belonging to two different levels.”[58]"

The terms “character” and “symbol” are renderings of the Chinese words “wénzì文字” and “fúhào符號”, respectively. In this sentence, Qiú still uses the term “symbol” in a two-fold manner, but two paragraphs later he makes an important clarification: “For the sake of clarity, the symbols used by a character will be called graphic symbols (=zìfǔ字符).”[59]" Qiú dedicates himself throughout the book to adhere to this distinction. Characters are symbols representing language whilst themselves deploying graphic symbols. The “semantic symbols (=yìfǔ意符)” and “phonetic symbols (=yīnfǔ音符)”, which are the core pillars of his theoretical framework of a “semanto-phonetic script”, are, of course, exponents of “graphic symbols”. They help characters fulfill their role as symbols representing language.

Raising some of Qiú’s examples may foster better understanding. Qiú exploits the character “huā花 (flower)” to illustrate the difference of the two levels of symbols.

“(T)he Chinese character “花” huā is a symbol of the Chinese word{花} "flower"; "艹" (the grass component, originally written “艸”, the old graph for căo “grass”) and “化” huà are the symbols used in writing the character “花”(a phonogram: “艹” is the signific and “化” the phonetic).”[60]"

In this case, the character is a “composite character (=hétǐ zì合體字)”. This means the character is composed of two components, each being understood as a graphic symbol. Hence, the distinction between the character as a symbol representing a Chinese word and the graphic symbols comprising this character is easy to comprehend.

The matter gets more intricate when the discussion turns to “non-composite characters (=dútǐ zì獨體字)”. Qiú presents the example of the character “rì日” , involving its very early writing form:

“(T)he ancient character A non-Unicode character (日)when viewed as a symbol of the Chinese word{日} rì “sun” is a graph possessing both a meaning and a sound; looked at as a symbol for the character “日” rì, then it is merely a pictographic symbol with a shape resembling the sun.“[61]"

For Qiú, even “non-composite characters” are “composed” of graphic symbols. The difference between “composite characters” lies solely in the number of components[62]". The former has two or more components, whereas the latter has only one. In the case of the character “r ì日”, A non-Unicode character or 日 is a pictographic symbol with a semantic load deriving from its sun-like shape[63]"; only when this graphic symbol composes the character “日” representing the word “rì”, it acquires a reading. We could also conclude that the word level includes both reading and meaning; the graphic symbol level often includes only one of both.

Interestingly, loangraphs (=jiǎjiè zì假借字) can make use of entire characters as phonetic symbols, no question whether the borrowed characters are “non-composite characters” or “composite characters”. With regard to the ancient character “A non-unicode character(= 其)” and the modern character “花“, Qiú explains as follows:

“(W)hen the semantograph A non-Unicode character standing for {箕} jī "winnowing basket” was used to represent the modal particle{其} qí, the two words{箕} and {其} were not at all related semantically. Another example is the modern use of the phonogram{花} huā "flower" to represent the verb{花} huā "to spend (money)." While both of them are pronounced huā, they are totally unrelated semantically. Therefore, even though A non-Unicode character was originally a semantograph and “花” was a phonogram, when they are borrowed to write the modal particle{其} and the verb {花}, they function purely as phonetic symbols.”[64]"

What is the point of taking the pain to discern these two levels of symbols? In my view, Qiú did not clearly specify the objective of this distinction, but his discussion on the concept of morphemic or morphographic script clarifies his stance.

Qiú’s argument about the relationship between Chinese characters and morphemes is twofolded. On the one hand, he admits the existence of morphemic features in the Chinese script, taking morphemic theories of the Chinese script as perspectives fundamentally different from his own approach, which is focused on “graphic symbols”. On the other hand, he argues that, even when changing the perspective, morphemic theories should be amended to “morphemo-syllabic script”, pointing to words that are comprised of two or more syllables and cannot be broken down into one-syllable morphemes. In these cases, no single character in the string of characters expressing the pertinent word can represent an entire morpheme; any single character merely refers to a phonetic fragment of the morpheme, a syllable. In other words, his second argument is centered on the phonetic function of loangraphs, whilst admitting that other types of characters, such as semantographs, etc., could be considered morphemic graphs. Let us first see how he puts this second argument:

“(N)on-composite, quasi-composite and composite semantographs, as well as sign graphs and semi-sign, semi-semantographs, can all be viewed as morphemic graphs. However, the phonetic symbols of the Chinese script, although they are all written with ready-made graphs which were originally morphemic symbols, ought to be viewed as symbols expressing syllables; ……In the case of those loangraphs which record transcribed foreign words of two or more syllables, it is perfectly clear that they express the nature of the syllabic structure of morphemes. For example, the four characters "達魯花赤" dálǔhuāchì which were borrowed in the Yuan dynasty to write the Mongolian title darugaci "governor, keeper of the seal," are clearly all used as syllabograms. Loangraphs used to write native Chinese binomes like "倉庚" cānggéng "name of a bird" and "猶豫" yóuyù "hesitate"(see Sec. 9. 3)also clearly are employed to express syllabic structure.”[65]

The first argument of difference in perspective highlights the objective behind his emphasize on graphic symbols:

“The terms morphemo-syllabic script and semanto-phonetic script(or semanto-phonetic-sign script)are names given to the Chinese script looked at from different points of view. The former focuses on the level of linguistic structure represented by graphic symbols; the latter focuses on the semantographic or phonographic functions of graphic symbols. These two terms can coexist.[66]"

In my view, the focus on the functions of graphic symbols that comprise the characters is nothing else than an inquiry into the question of how Chinese characters have been created or deployed for the representation of certain features of language. For instance, the rationale behind the character “日” representing the word “rì (= sun)” is explained as because the graphic symbol “A non-Unicode character” is a pictorial representation of the sun, establishing a semantic link to the word “rì” but no phonetic link. In a similar manner, the character “花” can represent the word “huā (=flower)” because of the semantic link by the signific “艹/艸” and the phonetic link by the phonetic “化”. The same character “花” can represent the word “huā (=to spend)” because of the phonetic link established by the loan of the entire character as a single phonetic symbol.

These inquiries bear some resemblance to traditional Chinese philology, yet Qiú’s unwavering focus on “words” and “morphemes” strikes me as remarkably innovative. To be precise, Qiú‘s methodology fully integrates the morphemic theory, while also attempting to address traditional inquiries. Within this two-fold approach, the focus on the linguistic layer of “words” and “morphemes” clearly marks the linguistic turn in Chinese paleography, which had been carefully prepared by Táng Lán唐蘭 (1901-1979) and then accomplished by Qiú[67]". It marks a definite departure from the worn-out notion of ideographs that directly convey meaning, as broadly seen in the work of Shirakawa Shizuka白川静 (1910-2006) or still present even in Chén Mèngjiā陳夢家 (1911-1966)’s Zhōngguó wénzì xué中國文字學.

Having experienced the linguistic turn, the focus on “words” and “morphemes” might sound even banal to us. But regarding the Chinese script, this can be an extremely challenging journey as constantly changing shapes easily obscure the linguistic layer from our sight. Qiú‘s brilliant analysis on the semantic connection of the character “shǐ矢 (usually ‘arrow’)” in Ode 45.1 “之死矢靡它” to the well-known character “chén陳 (=to display, to state) demonstrates how subtle linguistic connections often go unnoticed[68]". At the same time, this linguistic layer, the unity of sound and meaning, is just what the character definition of the Unicode is lacking. This is also what we must focus on for a more historical comprehension of Chinese characters.

4.3. The challenges of paleographic understanding of characters

The greatest challenge for a paleographic understanding of characters is the reconciliation of the unity of sound and meaning, i.e. words or morphemes, with the shape of characters. Qiú (1988) devotes three entire chapters to explore the intricate relationship between words or morphemes and the character shapes: Chapter 10 “Allographs, Homographs, and Synonymic Interchange”, chapter 11 “Graphic Differentiation and Consolidation”, and, finally, chapter 12 “The Intricate Relationship Between Graphic Form and Sound and Meaning”. If we take the last chapter 13 “The Systematization and Simplification of Chinese Script” as a discussion on the endeavors in the People’s Republic of China to unravel this intricate relationship, four of the entire thirteen chapters are dedicated to this topic.

This, of course, is far too complex to unwind here. It must suffice to say that the bulk of discussions continues to oscillate between shape and morphemes; different shapes are brought together based on the unity of sound and meaning, and identical shapes are discerned based on the difference in the sound and/or meaning of the words they represent. Only the historical-linguistic context varies, and with it Qiú’s interpretation of the specific relationships. Qiú’s strong interest in how certain characters happen to represent certain words or morphemes further increases the complexity of the discussion as well as the possibilities of disagreement on certain interpretations.

Boiling it down to a phenomenological level, the shapes of Chinese characters change much faster in time and space than the morphemes they represent. The same morpheme can be represented by different shapes at different times or locations. A crucial role of paleography is to trace and keep records of these historical changes. For instance, Ōnishi Katsuya gathered the different writings of the word expressing the verb “to cast” on weaponry during the Warring States period (453/403-221 BCE). Except for the State of Yān燕 using a word cognate to the modern “wéi爲”, the others used cognates of the modern “zào造”. The various shapes of the character “zào造” are displayed in table 7[69]":

Table 7 Warring State Period shapes of the character“zào造”

Qín秦

Qí齊

Chǔ楚

Wèi魏・Zhào趙

Hàn韓

Sòng宋

艁, 鋯
𢽍, 郜, 俈
𬥘
𫿜
𪯓, 棗

These amount to ten allographs for a word that is still in use in modern Chinese. To my own surprise, I could find all ten being included within the Unicode character set, which, without a doubt, is a huge practical benefit. Given their distinct shapes, Unicode treats them as separate characters, which makes the Chinese script set significantly more complicated and convoluted than strictly needed from a linguistic perspective.

Shào Yǒnghǎi 邵永海 developed a concept that from the very beginning has been minted in order to decrease such complexities and intricacies. This concept is called “zìwèi字位” in Chinese and could be translated as “character position”. “Position” appears to refer to a fixed position within the morphemic space represented by the Chinese script.

Like Qiú Xīguī, Shào focuses on words or morphemes, but he stops at the stage of identifying the morphemes that are hidden behind the shapes and bundles identical morphemes into “character positions”, irrelevant to the variety of shapes that might represent them. In other words, Shào turns the perspective back to the morphemic writing aspect in order to reach a clearer picture of language features that are represented by characters but also somewhat concealed by the variety of their shapes.

Shào displays strong interest in the digital processing of characters in a linguistically meaningful manner. He points out repeatedly that different normative frameworks actually comprise unnecessary obstacles to this objective. For him, the impedimentary effect of normative frameworks can be witnessed in modern encoding of characters as well as in the historical changes of character shape throughout the three thousand and more years of Chinese writing history. One example of the impediments of modern character norms involves the character “guī龜 (tortoise, turtle)”, whose “formal/correct shape (=zhèngtǐ整體)” varies among East Asian states and regions:[70]"

Table 8 Various normative shapes of the character guī龜
Chinese gui Taiwanese gui Hong Kongite gui Japanese gui Korean gui
China Taiwan Hong Kong Japan Korea

Unicode reserves three code points for these five normative character shapes. “U+F907 /63751” for the Hong Kong version, “U+F908/ 63752” for the Korean version, and “U+9F9C/ 40860” for the other three shapes, which have been displayed in this paper with the help of different fonts. Similar discrepancies in normative requirements for character shapes are pointed out for the following string of allographs:

裏-裡;黃-黄;涙-泪;恥-耻;幫-幇;牀-床;昇-陞〔升〕;穽-阱;徑-逕;舉-擧;貍-狸;犁-犂;弔-吊;匯-滙;群-羣;剋-尅;妬-妒;荊-荆;够-夠;諮-咨;眾-衆;歎-嘆;啟-啓;髕-臏;腳-脚;汙-污

Shào refers to these examples and states that “character codes …… do not encompass information that connects allographs or other interchangeable graphs.” According to him, these “differences in character shape directly influence the stability and authority of digitized dictionary contents.”[71]" These problems are well-known, and many search engines are already capable of reconnecting allographs that have been separated by different code points within the Unicode standard. However, Shào claims the need for a linguistic solution to such problems that also includes the longer historical development of the Chinese script. In his view, the definition of “character positions” by common reading and meaning is a more reliable method of overcoming the intricacies caused by endless shape changes:

“After joining, every character position takes charge for multiple shapes, forming a character family. One character family holds the different shapes of one character position, such as Oracle bone inscription shapes, bronze inscription shapes, seal script shape, clerical script shapes, and all sorts of allographs and synonymic interchangeables. Since an effective connection is established within the character families, search results include all shapes of the same character position, no question which shape is used for the query.”[72]"

Relinking different character shapes through their representation of common morphemes appears to be the main objective behind Shào’s proposal of the concept of “character position”. In this regard, the concept could facilitate the handling of paleographic phenomena such as the various shapes of the character “zào造” gathered by Ōnishi.

However, the concept can also be utilized in the opposite direction, i.e. it can also help discern homographs and other shape overlapping phenomena. As Shào’s argument does not reach this far, the author will try to apply the concept to some overlapping phenomena frequently seen in paleographic material from the Qin and Han period.

The glyph “灋” is commonly known as the ancient form of the character “fǎ法 (=penalty, law)”. Xǔ Shèn許慎 (30/58 – 121/147 CE)’s Shuō wén jiě zì説文解字 records it as the seal script shape, stating that the current writing simplifies it to “法”. In legal literature from the late Warring States period, the same glyph represents two different words cognate to modern “fǎ法” and “fèi廢” respectively. The bamboo slips excavated from Tomb 11 in Shuihudi, Hubei province, entail an official notice by the prefect of the Nán province. On the slips 3 to 4 of the document, the same glyph is used for both “fǎ法” and “fèi廢”:

Now, the statutes on penalties are furnished. But the officials and the ordinary people do not adhere to it, …, that amounts to abandoning the clearly stipulated penal law.

灋(法)律令已具矣,而吏民莫用,…,是卽灋(廢)主之明灋(法)殹(也)。(睡虎地11号秦墓竹簡『語書』簡003-004)

Be aware that the original character “廢” is, based on the semantic difference, regularized to “fǎ法” and “fèi廢” respectively.

A wooden tablet excavated from the Liye-site in Longshan, Hunan province, shows that, as a part of the character unification under the First Emperor of the Qin, the function of this glyph was split into two; the function of representing the word cognate to modern “fèi (=abandon, abolish)” was delegated to the newly created glyph “廢”. Figure 28 displays an image of the entire tablet. Here, we only cite three items:

“Yǒu酉 (=tenth of the twelve Earthly Branches)” remains as was, “jiu (=alcoholic drink) is changed to “酒”.
酉如故更酒。
“Fǎ灋” remains as was, “fèi guān (=dismissal from office)” is changed to “廢官”.
灋如故更廢官。
“Shǔ鼠 (=rat) remains as was, “yǔ rén(=to give sb, to bestow sb)” is changed to “予人”.
鼠如故更予人。
Figure 28 Wooden tablet from the Liye-site connected to the script unification of the Qin
Figure 28 Wooden tablet from the Liye-site connected to the script unification of the Qin

After the script unification, “fǎ灋” and “fèi廢” distinctly refer to two separate words. Even before the unification, the semantic difference is significant[73]". It would hardly make sense to combine all occurrences of the glyph “灋” into a group of characters connected to the word “fǎ灋”. A more reasonable option would be to draw the line based on the semantic and phonetic difference that we asserted in the form of different regularizations in the transcription. That would mean that some instances of “灋” gather the group of characters connected to the modern character “fǎ法”, while others join the group of “fèi廢”. This grouping could be easily represented by Shào’s concept of “character position”. “Fǎ法” and “fèi廢” represent two different character positions. Characters that can be associated with either position are considered members of the same character family, regardless of which glyph was employed in the original text.

The next example is even more intricate. According to Shuō wén jiě zì, the character “zuì最 (=most, summary, etc.)” is composed of the graphic symbols “mào冃 (=cap)” and “qǔ取 (to take)”[74]". In popular writing, the two horizontal strokes inside the component “mào冃” are often abbreviated, reducing the entire character to the shape of “冣”. This makes this character a homograph to the character “jù冣 (to accumulate)”, which according to Shuō wén jiě zì is comprised of “mì冖 (cloth cover)” and “qǔ取”[75]". In his annotation to Shuō wén jiě zì, Duàn Yùcái段玉裁 (1735-1815) argues that the reading and meaning of “jù冣” are identical to “jù聚”. Hence, the character is generally considered an allograph to “jù聚”[76]".

The tendency of omitting the two horizontal strokes inside the “mào冃” component of “zuì最” has already been observed within the clerical Qin script at the end of the Warring States period, and has caused some scholars to argue that the characters “zuì最” and “jù冣” were merged during this time[77]". Indeed, the writing of these two characters is confusing, but they are not truly merged. They are distinguished in a relative sense from each other by writing different numbers of vertical strokes. Table 9 gathers all instances of “zuì最” and “jù冣” that can be found in the texts from Tomb M11 at the Shuihudi-Site in Yunmeng, Hubei. Comparing these instances, the shape “Example of a non-Unicode character” is used in the Yǔshū語書 and Qínlǜ shíbā zhǒng秦律十八種 to represent the character “zuì最”, and in Rìshū jiǎ zhǒng日書甲種, it represents the character “jù冣”. This phenomenon indeed gives the impression that the two characters are at least confused. However, there is one instance of the character “zuì最” in the Rìshū jiǎ zhǒng日書甲種, which is written as “Example of a non-Unicode character”, adding one more vertical stroke in comparison to the character “jù冣” in the shape of “Example of a non-Unicode character”. In other words, the scribe of the text clearly felt the necessity to distinguish these two characters, and did so by increasing or decreasing the number of vertical strokes.

Table 9 The relative distinction between “zuì最” and “jù冣” in the Shuihudi tomb 11

Yǔshū語書

Qínlǜ shíbā zhǒng秦律十八種

Rìshū jiǎ zhǒng日書甲種

Zuì最

One Occurence of zuiExample of a non-Unicode character

013

One Occurence of zuiOne Occurence of zuiExample of a non-Unicode character

013     014

One Occurence of zuiExample of a non-Unicode character

056背參

Jù冣

One Occurence of zuiOne Occurence of zuiOne Occurence of zuiExample of a non-Unicode character

005正貳  /  015背壹  /  016背壹

In this case, Shào’s concept of “character position” could be helpful again. No question whether the shape is written as “冣”, “㝡”, “Example of a non-Unicode character”, or “Example of a non-Unicode character”, characters that can be identified as representing the same word “zuì最” should be associated with one character family, sharing the same character position within the Chinese writing system; characters representing “jù冣” belong to a distinct character family, occupying a completely different position within the writing system.

Partly overlapping in component shape is also observed quite frequently in the clerical script during the Qin and the Han. Well known exponents are the components “聿” and “隶”, “彖” and “录”, “辶” and “廴”. Some transcribers view these interchanges of components as mistakes. If the character “dài逮” is written with the components “聿” and “廴”, they transcribe “建〔逮〕”. But these writing habits apparently have not been treated as mistakes by contemporaries; it was their customary way of writing. They were aware of the words or morphemes their script represented. Hence, our task should be to identify these words or morphemes and not to correct their writing habits. This is possible by accumulating different shapes to character families based on their position in the writing system, or, plainly, the words they represent.

4.4. A new way of TEI-encoding of characters

In the previous sections, the author tried to outline what it means to encode characters rather than glyphs. Now, we must find a way to express these ideas using TEI- elements. Actually, the <charDecl>, a member of the gaiji module, already distinguishes the elements <glyph> and <char>. Once again, the author is compelled to acknowledge that the TEI-guidelines were prepared well, yet we users easily overlooked what this difference means in the context of Chinese writing. Below, I will attempt to employ the paleographic examples provided in the previous section to delineate how to encode Chinese characters using TEI.

First, each character possesses an abstract shape. The <glyph> element can be utilized to document the shape. The basic syntax should be one <glyph> element for one shape, with an @xml:id attribute assigning an identifier to it. Within the <glyph> element, the <mapping> element could take charge of the task of defining shapes. This would look like this:

<glyph xml:id="glyph-ID"><mapping>shape</mapping></glyph>

Secondly, each character represents a word or morpheme. The word or morpheme can be identified through the synthesis of its sound and meaning. For the description of these two features, we primarily utilize the readings of characters as depicted in the Guǎngyùn廣韻, alongside with the gloss given in the same book. If a character is not documented in Guǎngyùn, we need to resort to the readings found in later dictionaries, but as a rule of thumb, a character not recorded in the Guǎngyùn should raise alarm within the context of paleography. With high probability, our interpretation of the character is inaccurate. Hence, this paper will not go into detail in this regard. For the glosses, reinforcement should be sought from early annotations and dictionaries. Glosses later than the 3rd century CE should be again handled with precaution. Judging from excavated language material, the Chinese language changed dramatically during the two Han dynasties. Accurate knowledge about early Chinese that still circulated in academic circles got lost during the 3rd century.

This character's linguistic representation aligns with the unit that combines sound and meaning. This unit gives the character a fixed position in the script system, a “character position” according to Shào’s theory. For illustrative purposes, we could give a modern representation of this position in the form of a character. The overall syntax for a character definition would then look like this:

<char xml:id="char-ID">

  <mapping type="sound guǎngyùn">fǎnqiè反切</mapping>

  <mapping type="gloss guǎngyùn ">a gloss</mapping>

  <mapping type="early gloss">a gloss</mapping>

  <mapping type="character position">a modern character</mapping>

  <note>A note if necessary</note>

</char>

Lastly, we would need to search for a method to link characters within the main text to the glyph and character definitions. Within the <charDecl> element, each <glyph> and <char> element carries a unique ID, stored in the @xml:id attribute. These can serve as reference benchmarks. The characters in the main text are denoted by <g> elements, as outlined in chapter 3. <g> elements can carry @ref and @ana attributes. These facilitate links to the glyph and character definitions. While glyph references can be established automatically, as demonstrated previously, character references still need scholarly scrutiny. Hence, it appears to be reasonable to use the @ref attribute for glyph reference and the @ana attribute for character reference. A character within the main text would then look like this:

<g ref="#glyph-ID" ana="#char-ID">

For demonstration, we will encode the few paleographic instances mentioned in the previous section. The glyph “灋” that represents the word “fǎ法” could be distinguished from the same glyph representing the word “fèi廢” as follows.

Glyph definition:

<glyph xml:id="U28747">

<mapping>灋</mapping>

</glyph>

Character definitions:

<char xml:id="C0001">

  <mapping type=" sound guǎngyùn">方乏切</mapping>

  <mapping type="gloss guǎngyùn">則也</mapping>

    <mapping type="gloss Shuō wén jiě zì">刑</mapping>

    <mapping type="gloss Yùpiān玉篇廌部)">刑</mapping>

    <mapping type="gloss Yùpiān玉篇水部)">法令</mapping>

    <mapping type="position">法</mapping>

    <note>……</note>

</char>

<char xml:id="C0002">

  <mapping type=" sound guǎngyùn">方肺切</mapping>

  <mapping type="gloss guǎngyùn">止也</mapping>

    <mapping type="gloss Zhōulǐ Zhèngxuán Zhù周禮鄭玄注">猶退也</mapping>

    <mapping type="gloss Yùpiān玉篇广部)">退</mapping>

    <mapping type="position">廢</mapping>

    <note>……</note>

</char>

Character in main text:

“Fǎ灋(法)” = <g ref="#U28747" ana="#C0001">

“Fèi灋(廢)” = <g ref="#U28747" ana="#C0002">

Given that most glyph definitions have been standardized by Unicode, there's no need to reinvent the wheel. Therefore, the example used the decimal codepoint of “灋” (=28747) as glyph-ID. In fact, there is no necessity for glyph definition at all if the glyph is already part of the Unicode Standard. As every script can automatically encode and decode Unicode characters, the @ref attribute can use the code point directly as a reference. Only for non-Unicode characters, a glyph definition would indeed be necessary. Below follows an example of glyph definition for the non-Unicode shape of the characters “zuì最” and “jù冣” mentioned in the previous section:

Example of a non-Unicode character:

<glyph xml:id="g0001">

  <mapping type="IDS">⿳宀一取</mapping>

  <graphic type="lìdìng隸定" url="gaiji0001.jpg"/>

</glyph>

Example of a non-Unicode character:

<glyph xml:id="g0002">

  <mapping type="IDS">⿳宀二取</mapping>

  <graphic type="lìdìng隸定" url="gaiji0002.jpg"/>

</glyph>

Supposing that the character positions of “zuì最” and “jù冣” were assigned the character IDs “C0003” and “C0004” respectively, characters in the main text could reference this glyph definition as follows:

“zuìExample of a non-Unicode character(最)” = <g ref="#g0002" ana="#C0003">

“zuìExample of a non-Unicode character(最)” = <g ref="#g0001 " ana="#C0003">

“jùExample of a non-Unicode character(冣)” = <g ref="#g0001" ana="#C0004">

Unfortunately, the author could not yet implement the ideas presented in this final chapter yet. The character definitions as well as the @ana attributes need to be created manually. This will require a considerable amount of time. The author looks forward to being able to showcase concrete outcomes at an upcoming meeting. Suggestions and criticism that can guide the author towards achieving this goal are highly valued.

Notes
[45] The Unicode Consortium, whose understanding and representation of Chinese characters will be discussed below, uses the term “ideograph”, but explicitly delineates the distinction between the representation of a morpheme (=logograph) and the representation of ideas or concepts:
"Notionally, a logograph (or logogram) is a unit of writing which represents a word or morpheme, whereas an ideograph (or ideogram) is a unit of writing which represents an idea or concept. However, the lines between these terms are often unclear, and usage varies widely. The Unicode Standard makes no principled distinction between these terms, but rather follows the customary usage associated with a given script or writing system. For the Han script, the term CJK ideograph (or Han ideograph) is used." (Unicode Consortium (2025, p.329 /2014, p.258).
In addition, specifically concerning Chinese characters:
The term ‘Han ideographic characters’ is used within the Unicode Standard as a common term traditionally used in Western texts, although ‘sinogram’ is preferred by professional linguists. Taken literally, the word ‘ideograph’ applies only to some of the ancient original character forms, which indeed arose as ideographic depictions. The vast majority of Han characters were developed later via composition, borrowing, and other non-ideographic principles, but the term ‘Han ideographs’ remains in English usage as a conventional cover term for the script as a whole.” (Unicode Consortium (2025, p.930/2006, p.409)
The basic stance on Chinese characters of the Unicode consortium apparently was basically consolidated by version 3.0 in 1998, with minor modifications in later versions. Here and below, quotations are therefore made from the latest version together with a reference to the earliest mention found by the author, separated by a slash.
[46] Morioka (2023) p.1:
「漢字は1字が古典中国語の1形態素に対応する表語文字と考えることができ、伝統的に形音義にかかわる諸要素の組合せや対応関係によって分析・整理されてきた。」
With regard to the trinity of shape, sound, and meaning, the reader may prefer to refer to the classical work of Wáng Lì on this topic, see Wang (1953).
[47] The Japanese readings appear to have departed furthest from the Guǎngyùn standard. Still, there are traces of Guǎngyùn sounds in Japanese usage of Chinese characters that have been lost in modern Chinese. For instance, the syllable “tsuつ” in the reading of “しつ室” derives from an ending consonant in the Guǎngyùn reading “式質”, which has been lost in the modern Chinese reading “shì”. The Korean readings appear to adhere most closely to the Guǎngyùn standard, though such a statement might collide with either Korean or Chinese modern national narratives.
[48] The divergence from this standard apparently is more significant in modern Chinese dictionaries as they do not consciously discern between meanings deriving from this traditional common ground and modern creations.
[49] The Unicode Consortium(2025, p47 /1999, p.5).
[50] The Unicode Consortium(2025, p935-936 /1999, p. 260-261).
[51] The Unicode Consortium(2025, p940 /1999, p.263).
[52] The Unicode Consortium(2025, p940 /1999, p.263).
[53] The Unicode Consortium accurately identifies the practicality issue regarding “Distinguishing Han Character Usage Between Languages”:
“There is some concern that unifying the Han characters may lead to confusion because they are sometimes used differently by the various East Asian languages. Computationally, Han character unification presents no more difficulty than employing a single Latin character set that is used to write languages as different as English and French. Programmers do not expect the characters ‘c’, ‘h’, ‘a’, and ‘t’ alone to tell us whether chat is a French word for cat or an English word meaning ‘informal talk.’ Likewise, we depend on context to identify the American hood (of a car) with the British bonnet. Few computer users are confused by the fact that ASCII can also be used to represent such words as the Welsh word ynghyd, which are strange looking to English eyes. Although it would be convenient to identify words by language for programs such as spell-checkers, it is neither practical nor productive to encode a separate Latin character set for every language that uses it. …… The Unicode Standard leaves the issues of language tagging and word recognition up to a higher level of software and does not attempt to encode the language of the Han characters.” (The Unicode Consortium(2025, p. 936-937/1999, p.261)
This is correct, but regarding a morphographic script that is supposed to represent morphemes, this does not alter the fact that the Unicode Standard encodes glyphs rather than characters. Within the traditional trinity of shape, sound, meaning, Unicode code points merely correlate to shape. That is sufficient for practical purposes such as the Unicode Consortium pursues, but a linguistically reliable description of Chinese writing as a morphologic script needs further exploration.
[54] Qiū (1988, p. 15/22). Page numbers before the slash refer to the revised Chinese version from 2013, while the numbers after the slash correspond to the English translation by Mattos and Norman.
[55] Ibid., p. 12/18.
[56] Ibid., p.3/5.
[57] That Qiū nevertheless sticks to the denomination “semanto-phonetic script” is based on the facts that “signs of the later stages are almost completely derived from semantic and phonetic symbols,” and that “a majority of characters are still formed from semantic and phonetic symbols.” (Ibid., p.15/22)
[58] Ibid., p.9/13. Mattos and Norman render the original term “wénzì文字” as “writing”, a choice I believe dilutes the emphasis of this passage.
[59] Ibid., p.10/14. Here, Mattos and Norman translate “wénzì文字” as “writing system”.
[60] Ibid., p. 9/13. Be aware that the second “symbol” in this passage refers to “graphic symbol”. This sentence precedes the definition of “graphic symbols”. Hence, this passage still uses the general term “symbol (=符號)” in an ambiguous manner.
[61] Ibid., p. 10/14.
[62] In this regard, the translation of “hétǐ zì合體字” and “dútǐ zì獨體字” as “composite character” and “non-composite character” is at least misleading. “dútǐ獨體” actually means “stand-alone body/item”, and “hétǐ合體” means “conjunctive body/item”. Hence, “one-component character” and “multi-component character” may come closer to Qiū’s original concept.
[63] According to Qiū, the standard script shape of “日” does not resemble the sun anymore. Hence, it has stopped being a pictographic symbol and turned into a “sign”. See Ibid., p. 12/14.
[64] Ibid., p. 11/16.
[65] Ibid., p. 16-17/24.
[66] Ibid., p. 18/26. The second sentence of this passage was omitted in Mattos’ and Norman’s English translation and was reinstated by the author of this article. The original text reads “前者着眼於字符所能表示的語言結構的層次,後者着眼於字符的表意、表音等作用.” Strictly speaking, Qiū failed to comply with his own distinction between the two levels of symbols in this sentence. Following the distinction, the first “graphic symbol (=字符)” should be changed to “characters (=wénzì文字)” or “Chinese characters (=hànzì漢字)” as language is represented by characters and not by graphic symbols. Morphemic theory puts the focus of interest on the relationship between characters and language, while Qiū’s theory focuses on how graphic symbols comprise characters as symbols representing language.
[67] In comparison to the morphemic perspective, one could argue that Qiū’s understanding of linguistics is still an old-fashioned rule-based language model. While the morphemic perspective focuses on phenomenologically describing how certain characters represent certain language features, Qiū relentlessly asks for the rules that make these representations possible. This very much resembles the ruled-based language models of traditional European grammarians, which apparently have been proven inaccurate by deep learning language processing models and, recently, by large language models.
[68] Ibid., pp. 261-263/397-398. Takashima , Ken-ichi (髙嶋謙一) already notd this highlight on page xxiv in his foreword for the English translation of Qiū (1988).
[69] Ōnishi and Miyamoto (2009).
[70] Shào (2015), p.73-74.
[71] Ibid., p. 73-74.
[72] Ibid., p. 76.
“經過歸並之後,每個字位下領屬若干形體,構成一個字族;同一字族中,包含一個字位各種不同的形體,如甲骨文,金文,小篆,隷書,各種異體和異寫形式等;由於内部建立了有效的連接,無論選擇哪個形體作爲檢索項,都可以獲得該字位的全部形體的查詢結果。”
[73] One might speculate an etymological connection between the two terms, as the punishments to which the word “fǎ灋” refers are mainly mutilative punishments. Mutilation might be associated with disability, disuse, waste, and so on. However, such connections cannot be proven by paleographic material of the said period of time.
[74] Duàn Yùcái段玉裁 (1735-1815) considers it a semantograph or syssemantograph (huìyì zì會意字), Xú Kǎi徐鍇 (920-974) a phonogram (xíngshēng zì形聲字).
[75] In this case, “qǔ取” is generally regarded as a phonetic symbol, turning the character “jù冣” into a phonogram.
[76] Furthermore, the component “mì冖” in both the character “jù聚” and the component “mào冃” within the character “zuì最” are often written in the form of “mián宀”, but this change will be ignored in the subsequent discussion.
[77] Ōkawa (2015).