TEI2026 Presentation Material

The Implementation of the TEI-Guidelines on Early Chinese Excavated Administrative Documents: With a Focus on Paleographic Issues

by Arnd Hafner

Released on July 31, 2026,

under the CC BY-NC 4.0 license.

Last update on August 8, 2026.

For a full text version, click on docx or pdf.

3. Automated text collation

3.1. Automatization enhancing syntax considerations 1

Traditional paleography has piled up huge amounts of research results since the first discovery of early Chinese wooden and bamboo texts at the dawn of the 20th century. Automated encoding as described in the previous chapter enables us to easily carry on these results into the digital age. There are many things that we can do with the encoded digital texts, such as further semantic and syntactical annotations, connecting different texts and text fragments based on these annotations, etc. Before getting indulged in such dreams, the author would like to automize another basic task of paleography or philology in general, i.e. text collation.

We possess various transcriptions of the same excavated text alongside different excavated witnesses of identical, similar, or related early texts. Collating these texts and text fragments is the first step toward in-depth comprehension of the text, but it can be painstaking, time-consuming, and sometimes even monotonous or dull. Automatizing this task appears to be a natural demand. This chapter will describe some syntax considerations that facilitated the development of an experimental version of automated text collation. The texts utilized for the experiment were two distinct transcriptions of the compilation, which has already extensively been employed in the previous chapter[29]".

No question whether automated or manual processing, the degree of complexity easily affects the accuracy of the collation. The greater the number of different layers of information processed at once, the higher the probability of misinterpretations and errors. In this regard, the inline encoding introduced in the previous chapter causes unnecessary complexity by scattering mutual unrelated annotations over the entire text in a highly arbitrary manner. Let us take a closer examination of the first two lines of text displayed in figures 7, 9, and 10:

得羣盜﹦(盜盜)殺人購┘(。)癸、行請告瑣等曰:瑣等弗能詣告,移鼠(予)癸﹦等﹦(癸等。癸等)詣州陵,盡鼠(予)瑣等(slip 006)

【死辠(罪)】四萬三百廿(二十)錢。癸【券付死辠(罪)購。瑣等利得死辠(罪)購,聽請相移。先以私錢二千】(slip 006(2))

On slip 6 and 6(2), we observe the following annotations:

6 character regularizations (3 of them within a supplement)

1 punctuation mark regularization

2 abbreviations and their expansions

2 supplements for damaged text fragments (both entailing character regularizations)

2 supplements for illegible characters

8 supplementary modern punctuation marks (3 of them within supplements)

Even in traditional analog collation, these annotations can easily cause confusion. It would be much easier if we initially separated the annotations from the main body of text[30]".

Main body of text:

得羣盜﹦殺人購┘癸行請告瑣等曰瑣等弗能詣告移鼠癸﹦等﹦詣州陵盡鼠瑣等(slip006)

購四萬三百廿錢癸等(slip006(2))

Character regularization:

Slip 6 position 23 鼠(予)

Slip 6 position 32 鼠(予)

Slip 6(2) position 6 廿(二十)

Slip 6(2) supplement 1 position 2 辠(罪)

Slip 6(2) supplement 2 position 4 辠(罪)

Slip 6(2) supplement 2 position 11 辠(罪)

Punctuation regularization:

Slip 6 between position 8 ┘(。)

Expansion of abbreviations:

Slip 6 position 3-4 盜﹦(盜盜)

Slip 6 position 24-27 癸﹦等﹦(癸等。癸等)

Supplements for damaged text fragments:

Slip 6(2) before position 1 【死辠】

Slip 6(2) after position 9 【券付死辠購瑣等利得死辠購聽請相移先以私錢二千】

Supplements for illegible characters:

Slip 6(2) position 1

Slip 6(2) position 1

Supplementary modern punctuation marks (omitted)

To achieve the same segregation and systematic organization in our encoding, we must transform the inline annotations into stand-off apparatuses categorized based on the annotation's content. The double-end-point-attached method allows us to do so. A comparison will illuminate the decrease in complexity. Figure 13 shows the inline version created by the script described in the previous chapter:

A picture of inline annotations
Figure 13 Example of inline encoding

By contrast, the stand-off version looks like figure 14:

A picture of stand-off annotations
Figure 14 Example well-sorted stand-off encoding

It should be self-evident that the comparison of well-sorted stand-off encodings needs far less complex comparisons. Different layers of information are pre-configured; all comparisons can be limited to a range of basically homogeneous sequences, i.e. simple strings in the main body, Arabic numbers in the links of the apparatuses, and XML elements with completely identical structure in the apparatuses. Since even the comparison of xml-elements with identical structure amounts to nothing more than the comparison of text-nodes entailed, the subject of the entire comparison process gets reduced to linear sequences of either characters or numbers.

Furthermore, the separation and classification of annotations in the stand-off version are also advantageous for adding multiple layers of annotation to the text. As mentioned before, additional semantic and syntactic annotations are highly desirable. These annotations might well involve intricate levels of nesting and overlap. Given that overlaps among xml-elements are not permitted, the complexity of the annotation layers would be extremely constrained in the scenario of inline annotation. Using separated and categorized stand-off annotations avoids this kind of limitation, since overlapping of annotations merely results in overlaps in link ranges that are nothing more than abstract numbers and not in overlaps of XML elements.

3.2. Automatization enhancing syntax considerations 2

The text examples given in the previous section did not contain any non-Unicode characters. As a matter of fact, the proportion of non-Unicode characters in the compilation as well as other written material of the Qin and Han periods is notably smaller compared to those found in bronze inscriptions of earlier times or in bamboo texts excavated from tombs within the area of the Warring States Period state of the Chu. This is because even modern Chinese script still remains significantly influenced by the script unification conducted during the reign of the First Emperor of the Qin, and the script unification itself was highly influenced by the actual character usage within the realm of the Warring States Period state of the Qin. The genealogical proximity to the modern Chinese script helped to ensure that a huge part of the Qin script got included in the Unicode character set. By contrast, many earlier characters and characters from other regions were weeded out during the Qin’s character unification. Many were accurately recognized only recently; it will still take some time before they get reflected in Unicode.

However, a notable presence of non-Unicode characters persists in excavated literature and documents from the Qin and the following Han. Considerations are required with regard to their influence on the homogeneity of the character sequences that will be subject to comparison during automated collation.

In a legal case that is recorded on slips 171 to 188 of the compilation, there are 63 appearances of the character “Example of a non-Unicode character”. It is used as the name of the victim of a sexual assault by her former husband in this case. The same name is frequently encountered in other excavated literature and documents as well as in personal seals. It appears to have been quite a common name, but, unfortunately, there are no occurrences of this character other than in personal names. Thus, it is impossible to exactly trace it down to any known later character.

In an article co-authored by the author and Chén Jiàn陳劍, we argued, based on suggestions by Shī Xièjié施謝捷, that the character could be broken down into the two components “交” and “于”[31]", the corresponding Ideographic Description Sequence would be “⿴交于”. We also advocated that it could be a variant of the character “yū㝼 (=crooked thighs)”. Fāng Yǒng further developed the argument by suggesting that the component “交” actually is a variant of “黃“, which he interprets as the original shape of the character “wāng尪” (also “wāng𡯁”, the name of a disease inflicting spine deformations) based on a theory of Táng Lán唐蘭[32]". According to Fáng’s comprehension, the character should be described as “⿴黃于” or “⿴𡯁于”, the latter neatly fitting into the argument that the character is just a variant of “yū㝼”.

What impact could these theories about this uncertain character have on automated collation? Let us review the initial instance of the character in the compilation. In my transcription the pertinent slip reads:

【……當陽隸臣得之气(乞)鞫曰:……】□,不與棄妻Example of a non-Unicode character奸,未𧐂(蝕)。當陽論【得之爲】(171)

According to the previous arguments, merely the reduced transcription of this text fragment, which would be stored in the main text body of the TEI-conform xml-file, could have the following variants:

1: □不強與棄妻⿴交于未𧐂當陽論耐

2: □不強與棄妻㝼未𧐂當陽論耐

3: □不強與棄妻⿴黃于未𧐂當陽論耐

4: □不強與棄妻⿴𡯁于未𧐂當陽論耐

The main obstacle that these variants pose is the fact that the number of characters varies although all four transcriptions suppose the same number of characters in the original text. In other words, the order of the characters loses the direct connection to the original text as understood by the transcriber. For instance, the glyph “𧐂”, which is regularized to the character “shí蝕” and, accordingly needs a link to an annotation, is the eleventh character in variation 1 and 3-4, but the ninth in variation 2. Corresponding links would easily point to different, unintended characters.

At the same time, comparison modules such as diff-lib are unable to distinguish between a character and a component of an Ideographic Description Sequence. This can easily result in misinterpretations if components of Ideographic Description Sequences appear as individual characters shortly before and after. For example, the results of automated comparison could be different if the character “wèi未” would have been understood as a “yú于” by the transcriber in variation 2, i.e. the character “yú于” would be considered an exact match to the IDS component “于”.

The TEI-guidelines provide the element <g>(gaiji) for the description of non-Unicode characters. If we substitute this element for Ideographic Description Sequences, we could prevent confusion between sequence components and independent characters. The unintended side effect would be that the reduced transcription would entail text nodes and element nodes in an arbitrary order. This again would be a challenge to the automated interpretation of the output of comparison modules like diff-lib.

To bypass this hurdle, I changed the semantic meaning of the element <g> to “g” meaning “glyph”. In my view, the paleographic interpretation of characters diverges significantly far from that of the Unicode consortium. Turning all Unicode and non-Unicode characters into homogeneous <g> elements opens up space for detailed character description in the <charDecl> element, while simultaneously maintaining the homogeneity of the comparison subjects. If all characters of a text are converted into <g> elements, text comparison essentially reduces to the comparison of the IDs assigned to the <g> elements. In general, these IDs consist of simple ASCII strings.

Applying this idea to the first variant of the text fragment on slip 171 would yield the <charDecl> and text body as depicted in figure 15:

A picture of <g> elements representing glyphs
Figure 15 Example of <g> elements representing glyphs

The encoding of the second variant would require the declaration of a new glyph, for instance “<glyph xml:id="_13"><mapping type="unicode">㝼</mapping></glyph>”. The output of an automated comparison would be that the difference between variation 1 and variation 2 is the replacement of the ID “_6” by “_13” at position 6[33]".

At first glance, the conversion of all characters into glyph-elements (<g>) appears to be counterproductive. For the human eye, the main text body is already incomprehensible. As figure 16 demostrates, only about ten lines of an XSLT script are needed to convert the file into a readable format for humans. I am confident that the same transformation can be effortlessly through in a viewer used by experienced TEI users via Java script. Thus, human readability hardly qualifies as a major hurdle.

Example of script converting <g> elements into human readable glyphs
Figure 16 Example of script converting <g> elements into human readable glyphs

The use of <g> elements also allows for the incorporation of the two traditional markups typically used in "restrained" transcriptions, i.e. boxes for supplement transcriptions of illegible characters based on context and question marks in parentheses for uncertain transcriptions. A @cert attribute could switch between the values “high” and “low” to describe the certainty or uncertainty of transcriptions; a @type attribute could carry the values “original” and “context” to distinguish transcriptions directly based on the original glyph from those that were based on context information. Using the glyph IDs of the previous example, “” would result in ‘<g cert="high" ref="#_2" type="context">’; “強(?)“ in ‘<g cert="low" ref="#_2" type="original">’, and “(?)” in ‘<g cert="low" ref="#_2" type="context">’.

3.3. A prototype of automated text collation

The author carried out the ideas outlined in the preceding two sections using a Python script called "compare.py"[34]". This script takes multiple text files as witnesses to one original text. The script hovers over an ordered list of file names, reads the corresponding files in the listed order, and treats the first as the main witness against whom all other witnesses are collated[35]". The main witness is assigned the ID “01”, while the others receive IDs “02”, “03” and so forth. After encoding all witnesses as outlined in section 3.1. and 3.2., the witnesses with the ID “02” and above are then compared to the main witness. Comparison is conducted information layer by information layer; each information layer comparison is processed witness by witness.

As a preparatory step, the witness ID of the main witness is stored in the @source attribute of the corresponding element for each information layer[36]". A sequence of strings or numbers representing the stored information in the main witness is also generated. Thereafter, corresponding sequences are generated for all other witnesses and examined to see if they are identical to or divergent from the main witness. This comparison is conducted by one witness after the other. If identical, the witness ID is added to the @source attribute of the corresponding element; if divergent, an apparatus containing the diverging reading is created and stored in the corresponding <div> element in the <back>.

In summary, the entire collation process is comprised of the encoding of all witnesses, the preparatory processing of the main witness for each information layer, and the comparison of all other witnesses with the main witness, also for each information layer separately. Figure 17 displays the flow chart:

Flow chart of automated text collation
Figure 17 Flow chart of automated text collation

In addition to the annotation layers mentioned in 3.1., two further <div> elements are created, each with a @type attribute referring to divergences in the slip arrangement and variations in the main text. Currently, the automated comparison of the experimental script focuses exclusively on these two annotation layers; the remaining layers will be incorporated in future updates of compare.py.

The limitation on these two layers of information in the experiment should not undermine its reliability or effectiveness. As all information is orderly stored in separate element layers within the main text body or in separate annotation layers in the back matter, collation can also be conducted in an orderly way layer by layer. Increasing or decreasing the number of layers handled does not substantially change the processing complexity.

Let us inspect what the collation script is actually doing with the different witnesses.

Considering that the physical appearance of all witnesses in this experiment is determined by modern transcribers rearranging the bamboo slips, the initial task is to compare the different slip arrangements. Slips are identified by ID numbers[37]". The script converts these numbers into @xml:id attributes[38]" and attaches them to the <p> element storing the text fragment found on the specific slip surfaces. The collation of slip arrangements gathers all @xml:id attributes for each text, converts them into string sequences, and then compares these sequences.

Sequence comparison typically yields four different results: 1) equal, 2) delete, 3) replace, 4) insert. For the collation of slip arrangements that means the following: 1) slip IDs found in the main witness are identical to slip IDs in another witness; 2) slip IDs found in the main witness are deleted in another witness; 3) slip IDs found in the main witness are replaced by another slip ID in another witness; 4) slip ID not found in the main witness is inserted by another witness.

We need to keep in mind that the results of comparison modules merely refer to specific positions within the main sequence. This can cause quite counter-intuitive results in the current experiment. For instance, if the comparison module states that a slip ID found in the main witness has been deleted in the other witness as in 2), that doesn’t mean that the said ID cannot be found at another position in that sequence. For our collation purposes, it would be necessary to conduct further inquiry into whether the deleted slip ID is found at a different position or is completely absent in the other witness. We could label the former result as 2a) and the latter as 2b). The same applies to result 4). For 3), there can be found even four sub-categories: 3a) The two slip IDs (or the two sub-sequences of slip IDs) are found in both witnesses, but at different positions; 3b) the two slip IDs are exclusively found in only one of the witnesses each; 3c) the slip IDs found in the main witness are exclusive, while the other slip IDs can be found in both witnesses at different positions; 3d) the slip IDs found in the other witness are exclusive, while the slip IDs of the main witness are also found in the other witness at another position[39]".

This rather unexpected outcome of the comparison modules stems from the fact that they compare “sequences”, whereas the slip arrangement is a list of ordered elements (i.e. slip IDs) without repetition. In mathematics, a “sequence” constitutes an ordered collection of elements that allows repetition of values. By contrast, slip IDs constitute a “set”, i.e. a collection of different things. In plain words, each slip ID is unique. Reconstructing the order of slips, or the text, can be understood as selecting a subset of slip IDs from the larger set of all known slip IDs and arranging the selected subset orderly without repetition. This is also known as a variation without repetition. Two different slip arrangements represent two variations without internal repetition.

What we actually want to know when comparing two variations without repetition taken from the same superset of slip IDs is the following: I) Are there identical IDs at the same positions? II) Are there identical slip IDs at different positions? III) Are there slip IDs present in the main witness but not in the other witness? IV) Are there slip IDs present in the other witness but not in the main witness?

Unfortunately, the author could not find a module that conducts comparisons of variations without repetition. Thus, the script needs to convert the outcomes of the comparison module by further inquiring about the presence or absence of inserted, deleted, replaced, and replacing slip IDs. The correspondence between the comparison module’s output and the desired output can be as outlined in table 5:

Table 5 Correlation of comparison module output and desired output

Desired Output

Comparison Module Output

I) Identical slip ID at identical position 1)
II) Identical slip ID at different position 2a), 3a), 3c), 3d), 4a)
III) Slip ID found in main witness only 2b), 3b), 3c)
IV) Slip ID found in other witness only 3b), 3d), 4b)

The script handles the comparison results as follows: For I) through III), the witness ID of the main witness is added to the @source attribute of the corresponding <p> element within the main text body. For I), the witness ID of the other witness is also included in the same @source attribute. For III), no further action is required. For II) and IV), a new apparatus containing a <rdg> element is created and stored in the <div> element with the @type attribute referring to the divergences in the slip arrangement. The witness ID is recorded in the @source attribute of the <rdg> element. Finally, a <p> element containing the slip text is inserted into the <rdg> element. For II), the slip text is replaced by a @copyOf attribute that references the same slip within the main body; for IV), the slip text is inserted into the <p> element in the same manner as slip texts of the main witness in the main text body[40]".

The following figures provide some examples. Figures 18 and 19 display the text fragments from slip 091 to 098 for each witness. When publishing the first edition of the compilation in 2013, the slips 091 to 098 were arranged in ascending order, knowing that two to three slips were missing. It was understood that slip 094 was the last slip of the preceding case, while the first slip of the following case could not be located at that point in time. It could only be partly reconstructed based on mirror-inverted imprints on the verso of slip 104, being assigned the ID “缺08 (=missing 08)”. Guessing from context data, two other missing slips were reconstructed after slip 095 and slip 097, the assigned IDs labeled “缺09 (=missing 09)” and “缺10 (=missing 10)” respectively.

In the revised edition of the compilation from 2021, the slip missing behind slip 094, “缺08 (=missing 08)”, had been located and assigned the ID “094(2)”. The newly discovered slip’s text indicated that slip 095 did not fit into the context behind slip 094(2). After further inquiries, slip 095 was finally inserted between slip 097 and 098, replacing the reconstruction of the missing slip “缺10”.

The text of the first edition is contained in the text file “input02.txt” and put in the second place of the ordered list of file names to be collated by the script. (Figure 18).

Figure 18 Collation script input file example (first edition of the comparison)
Figure 18 Collation script input file example (first edition of the comparison)

The text of the revised edition is stored in the text file “input01.txt”, which is put on first place of the ordered list of file names and handled as the main witness (figure 19).

Figure 19 Collation script input file example (revised edition of the comparison, main witness)
Figure 19 Collation script input file example (revised edition of the comparison, main witness)

For each input file, the script generates a sequence of slip numbers, reflecting the different slip arrangements. After comparing the sequences, the comparison module returns an output that states “equal” for the sub-sequences of “091” to “094”, “缺09” to “097”, and for “098”; “replaced by 缺08 and 095” for (the subsequence) “094(2)”[41]", and “replaced by 缺10” for “095”. For the slips 091 to 094, slips 缺09 to 097, and slip 098 (along with several subsequent slips), identical slip IDs are observed at identical positions (I). Hence, the IDs of both witnesses are added to the @source attribute of the <p> element containing the pertinent slip text within the main text body.

The slip ID “095” is observed in both witnesses, but at different positions (II). Only the witness ID of the main witness is added to the @source attribute of the pertinent <p> element within the main text body. The witness ID of the second witness is appended to an <rdg> element in an apparatus that is created to record its replacement by the slips 缺08 and 095 and linked to slip 094(2). As a slip text for the slip 095 is already given in the main text body, its content is merely referred to through a @copyOf attribute[42]".

The slip IDs “094(2)”, “缺08” and “缺10” are exclusively found in only one of the witnesses each (3b). For the slip ID “094(2)” that means the slip ID is found in the main witness only (III); the slip IDs “缺08” and “缺10” are found only in the second witness (IV). Hence, the witness ID of the main witness is stored in the @source attribute of the

element containing the slip text of “094(2)” within the main text body; the existence of the slips “缺08” and “缺10” in the second witness is recorded in the apparatuses within the <back> element, linked to slips “094(2)” and “095” respectively.

The main text body part is displayed in figure 20.

Figure 20 Example main text body after automated collation
Figure 20 Example main text body after automated collation

The back matter part is shown in figure 21[43]".

Figure 21 Example back matter after automated collation
Figure 21 Example back matter after automated collation

The second part of the current experiment aims at comparing the main bodies of text, slip by slip. The slip texts were not compared according to the order of appearance in the respective witnesses, but for each pair of slips with identical IDs, notwithstanding the position of the specific slips within the entire slip arrangement. The slip texts were also stripped off of most of their traditional markups. As stated earlier, the goal is to simplify the individual comparison processes by preselecting comparison objects according to their basic information load or the layer of information they belong to. For instance, differences have been detected for slips 032, 042, and 044. Stripping the slip texts off their markups, the text to be compared looks like the following table 6.

Table 6 Comparison example main text body without any markup
Slip 032
Input01好□□□部中即令覆□驩求盜尸等十六人追┘尸等產捕詣秦男子治等
Input02好□□□部中即令獄史驩求盜尸等十六人追┘尸等產捕詣秦男子治等
Slip 042
Input01□□□□□□攻盜┘京州降爲秦乃殺好等疑尺等購●𤅊固有審矣治等審秦人殹尸
Input02□□□□□□攻盜京州降爲秦乃殺好等疑尺等購●𤅊固有審矣治等審秦人殹尸
Slip 044
Input01●廿三年四月江陵丞文敢𤅊之廿三年九月庚子令下劾𢮑江陵獄上造敞士五
Input02●廿三年四月江陵丞文敢𤅊之廿三年九月庚子令下劾𢮑江陵獄上造敞┘士五

Even for the human eye, this comparison is an easy game. With regard to slip 032, the ninth and tenth characters “覆□” found in input01.txt are replaced by “獄史” in input02.txt; the original punctuation mark “┘” that is found on slip 042 at position 9 in “input01.txt” is deleted in “input02.txt”; finally, for slip 044, the same punctuation mark is added between the thirtieth and the thirty- first character of “input01.txt” by “input02.txt”. The same comparison result is documented in “output(collated).xml” as in figure 22:

Figure 22 Example result of main text body comparison
Figure 22 Example result of main text body comparison

The “<g cert="low" ref="#177" type="original"/>” and “<g cert="low" ref="#135" type="original"/>” in the first apparatus stand for the characters “獄” and “史” that replaced the 9th and 10th characters of slip 032 (= “覆□”)[44]"; the empty element “<rdg source="02"/>” in the second apparatus indicates that the corresponding character in input01 (“┘”) has been deleted in source 2; the link to the position instead of a subsequence of characters in the third apparatus signals that the following reading found in input02.txt is not a replacement of characters but an insertion at the specified position.

As text is a sequence selected from the set of characters allowing repetition, it frequently proves challenging to discern the editorial intervention that caused the differences observed between different texts. Certain comparison modules compare the shortest path of editorial changes from one text to the other; some modules search recursively for the longest overlapping and interpret the characters in between as changes. Whatever the rationale may be, automated comparison results occasionally fail to align with our human understanding. Though the author did not encounter such a scenario during the current experiment, encoding based on automated text comparison should always be prepared to check and reinterpret carefully the given comparison results.

Correspondingly, the final stage of the current experiment was the automated conversion of the collation results into annotated texts in HTML format, allowing for straightforward inquiry into the reliability of the collation process. Conversion was again conducted by XSL transformation. Figure 23 shows the main text fragment of the slips 094(2) to 098. The first line assigns the case that starts from slip 094(2) with a title, which is provided by the editor. The slip arrangement and the transcription follow the main witness. The transcription covers only two sorts of traditional markups: the boxes indicating supplementary transcription based on context and the question marks in parentheses expressing uncertainty.

Figure 23 XSL-transformation example of collated xml-file (main text body)
Figure 23 XSL-transformation example of collated xml-file (main text body)

Numbers in square brackets are annotation numbers. Links guide the reader to the annotations stored at the end of the case. Figure 24 shows the Japanese version:

Figure 24 XSL-transformation example of collated xml-file (annotations, Japanese version)
Figure 24 XSL-transformation example of collated xml-file (annotations, Japanese version)

Except for a language switch to Chinese, figure 25 displays the same contents:

Figure 25 XSL-transformation example of collated xml-file (annotations, Chinese version)
Figure 25 XSL-transformation example of collated xml-file (annotations, Chinese version)

Figure 26, in addition to displaying the same content, features the notes of an extra case documented on slips 137 to 141. These are in English now.

Figure 26 XSL-transformation example of collated xml-file (annotations, English version)
Figure 26 XSL-transformation example of collated xml-file (annotations, English version)
Notes
[29] Now, three comprehensive versions are available for the compilation. Hafner (2013) is the original version published in cooperation with the editorial committee of the Yuèlù-Academy; Wēn Jùnpíng et al. (2018) includes a revision done by Ōu Yáng歐揚; Hafner (2021) is an updated version by the author.
[30] In my view, this should be the main reason that traditional editions of excavated written materials often provide “enriched” and “restrained” transcriptions as shown in figures 9 and 10 of the previous chapter.
[31] See Hafner and Chén (2011) p. 413.
[32] Nowadays, the glyph “黃“ denotes the character “huáng黃(=yellow)”. According to Táng Lán唐蘭, the initial shape of this character was an ideograph describing the meaning of “wāng尪 (=crooked spine)”, which would mean that the character “huáng黃(=yellow)” is a phonetical loan from “wāng尪”.
[33] At this phase of the experiment, the glyph IDs are assigned in the order of appearance of the glyphs within the encoded text. As this could cause unnecessary confusion in case of comparing different texts, the method of glyph ID assignment will be discussed further in section 4 of chapter 4.
[34] The script is available at https://www.aa.tufs.ac.jp/~Ejina/DBexp/09/compare.py. The text comparison module applied in this script has been developed by Dr. Wáng Xún王荀. The module is available at https://www.aa.tufs.ac.jp/~Ejina/DBexp/09/newdiff.py, English documentation is included in the file, a Chinese explanation is found in the supplement to Hafner (2025b).
[35] It might be questionable whether the concept of “witness” is appropriate here. In diplomatic terms, the compilation as it was buried in a tomb is a “witness” of a certain ancient text. It is a more or less accurate copy of a certain text, which might have changed during its circulation among contemporaries. However, every modern transcription of the compilation arranges the excavated slips in different orders, again creating different witnesses of the same text, which was once copied, buried and eventually survived up to today. In this sense, this paper regards different transcriptions as different witnesses.
[36] Initially, the author intended to use the @wit attribute. However, the @wit attribute is not permitted in <p> elements. “compare.py” uses the @source attribute instead, which is allowed in <p> elements as well as <lem> and <rdg> elements. It might well be that the author’s misunderstanding about the semantic meaning of the two attributes caused this permission conflict.
[37] To avoid unnecessary confusion, we treat the slip ID numbers as unquestioned premises in our collation experiment. In truth, the identification and numbering of excavated writing material can be quite intricate and challenging. As for the case of the Yuelu-Academy bamboo and wood text collection, I was forced to write an entire book about this topic after my involvement in the first edition of the compilation, which was a part of this collection. See Hafner (2016).
[38] In the initial text edition, the recognized slips were numbered from 1 to 253; missing slips were numbered from 1 to 15, adding the suffix “quē缺 (=missing)”. When the revised edition included newly recognized slips, their IDs were decided depending on their relative position within the initial slip arrangement. For instance, a slip inserted behind slip 244 was assigned the number 244(2). As xml:id attributes don’t permit integer only IDs, all IDs were suffixed with the character “J”. The digit was also adjusted by adding “0”s between the character “J” and the original ID. Furthermore, brackets were replaced by the characters “S” (=“(”) and “E” (“)”). As a result, the conversion of slip numbers into @xml:id attributes produced three slightly different formats: 1) “J\d\d\d”, ranging from “J001” to “J253”; 2) “J缺\d\d”, ranging from “J缺01” to “J缺15”; 3) “J\d\d\dS\dE”, with 6 unordered occurrences.
[39] As comparison modules return output not term by term but in units of sub-sequences, 3c and 3d often overlap. For instance, taking the revised edition from 2021 as main witness and the initial edition from 2013 as second witness, one of the comparison outputs states that the sub-sequence of slip 244(2) and 244(3) in the main witness is replaced by the sub-sequence of slip 245, 246, 247, and 248. While slip 245 and 246 appear in both witnesses, slip 244(2) and 244(3) are unique to the main witness, in contrast to slip 247 and 248 which exclusively appear in the second witness. That means that with regard to the slips 244(2) and 244(3), we observe the situation 3c), and with regard to the slips 247 and 248, we are confronted with situation 3d).
[40] This method of processing the results of type IV) is prone to causing some trouble in the future and, therefore, needs further refinement. If slip IDs that are not found in the main witness appear at two different positions in two other witnesses, the current processing method leads to repeated recording of the same slip. This is a redundancy that will easily result in further difficulties if, for instance, the transcriptions of these witnesses differ. One solution would be to create a separate <div> element that stores in <p> elements the slip text of slips not entailed in the main witness. Taking the witness in which the slip first occurs as a main witness for this particular slip, the handling of differences between the following witnesses could be handled in analogy to the slip texts stored in the main text body. This processing method can be viewed as a form of recursive creation of main witnesses and their comparison with other witnesses.
[41] Be aware that the ID “094(2)” is replaced by “094S2E” in the XML file “output(collated).xml” due to restrictions on XML identifiers.
[42] This does not mean that the text on slip 095 needs to be identical in both witnesses. The differences between the slip text are recorded in a second collation process by creating an apparatus for each differing character separately.
[43] With regard to the text of the slips 缺08 and 缺10, it needs to be mentioned that the difference in their respective transcription methods stems from the difference in their recognition. Though slip缺08 was reconstructed as a missing slip, the transcription of its slip text was directly based on imprints on the verso of slip 104. The slip text was handled as existing within the witness itself and not as a supplement by the transcriber. By contrast, the text of slip缺10 is a pure supplement by the transcriber. In a traditional analog transcription, two different types of transcription could be provided for such a case. One merely records the loss of the slip by the markup symbol “〼” (see legend in 2.2). This markup is converted into a <g> element with the @ref attribute “gap03”. It is one of three elements predefined for the description of gaps. The other type of transcription puts the entire reconstruction of the slip text into black lenticular brackets (“【】”). For instance, the text of slip 缺10 was reconstructed as “【某曰:……】”. Since this is a supplement for a loss by damage, this text will be stored in a <div> element within the back matter extraordinarily dedicated to such supplements and linked to the simplified slip text “〼”.
[44] “#Range” refers to a sequence of one or more elements, while “#left” refers to the position before a certain element. See TEI Consortium, eds. “17.2.4 TEI XPointer Schemes”, Guidelines for Electronic Text Encoding and Interchange P5 Version 4.11.0. Last updated on 18th February 2026, https://tei-c.org/release/doc/tei-p5-doc/en/html/SA.html#SATS (accessed on July 6th,2026).