WH/T 90-2020 in English
VALIDUnity description for Chinese character identification
- Issued on:2020-09-01
- Implemented on:2021-01-01
- File Format:PDF
- Delivery:Via email within 1~3 business days
$146.00
| Standard No: | WH/T 90-2020 |
| Document status: | VALID |
| Title in English: | Unity description for Chinese character identification |
| Title in Chinese: | 汉文古籍文字认同描述规范 |
| Language: | English |
| File Format: | Electronic (PDF) |
| Delivery: | Via email within 1~3 business days |
| Issued on: | 2020-09-01 |
| Implemented on: | 2021-01-01 |
| ICS Classification: | 01.140-Information sciences. Publishing |
| Chinese Classification: | A14-Library, Archives, Literature and Information |
| Professional Classification: | WH-Culture |
| Related Keywords: | general standard chinese character table
standard technological breakthrough character character identity private character chinese character identification introduction analysis |
| Related Topics: | Word
ancient books text and data recognize describe the result word description How to describe in the paper Test English description English description |
本标准规定了汉文古籍文字认同描述的元数据、文字认同规则描述以及文字认同实例描述的内容、结构及各要素的描述规则。
本标准适用于图书馆及相关机构开展汉文古籍数字化工作中对文字认同过程和结果进行描述。民国时期文献的文字认同可参考执行。
Introduction
Analysis of the core content of the standard
| Element category | Traditional processing method | Requirements of this standard | Technological breakthrough |
|---|---|---|---|
| Character set specification | Use private character set | Mandatory use of Unicode basic set | Achieve cross-system data exchange |
| Variant character processing | Manual experience judgment | Based on the "General Standard Chinese Character Table" and other standards | Establish traceable conversion rules |
| Metadata records | Unstructured records | 8 major categories of standard metadata fields | Support machine learning training |
Key technical points
The standard innovatively proposes a three-factor model for character identity: character shape features (structure, strokes), character pronunciation association (ancient and modern sound changes), and semantic context (contextual relationship). In the Yongle Encyclopedia digitization project, this model was used to successfully solve the standardization of 1,243 difficult variant characters.
Implementation Suggestions
- Establish a project-level character identity rule library and regularly maintain the version update mechanism
- For large series such as the Siku Quanshu, it is recommended to formulate differentiated rules according to the division of classics, history, and collections
- Introduce IDS (Ideographic Description Sequence) to describe characters outside the collection
Historical Document Processing Case
In the digitization of the Song Dynasty edition of Wenxuan, it was found that the character "峕" has three variants:
1. 峕 (U+5CD5)
2. 𡵂 (U+21D42 in the extended area)
3. 时 (vulgar writing)
According to clause 4.3.2 of the standard, it is uniformly recognized as "时" (U+65F6), and the original glyph features are recorded in the metadata.

Loading PDF document...
Error loading PDF. Please make sure the file is valid and try again.
We also recommend
-

WH/T 101-2024 in English
Library Volunteer Service Management Guide
2024-07-22