Sign In |Help & Support
ALL SECTORS
  • ALL SECTORS
  • GB(National Standard)
  • CB(Shipping)
  • CECS(Engineering Construction)
  • CJ(Urban Construction)
  • CY(News and Publication)
  • DB(Provincial Standard)
  • DL(Electricity & Power)
  • DZ(Geology & Mineralogy)
  • FZ(Spinning & Textile)
  • GA(Public Security)
  • HB(Aviation)
  • HG(Chemical Industry)
  • HJ(Environmental Protection)
  • JB(Machinery)
  • JC(Building Materials)
  • JG(Building & Construction)
  • JJ(Metering)
  • JT(Highway & Transportation)
  • LY(Forestry)
  • MT(Coal)
  • NB(Energy)
  • NY(Agriculture)
  • QB(Light Industry)
  • QC(Automobile & Vehicle)
  • QJ(Aerospace)
  • SH(Petrochemical)
  • SJ(Electronics)
  • SL(Water Resources)
  • SN(Commodity Inspection)
  • SY(Oil & Gas)
  • TB(Railway & Train)
  • YB(Ferrous Metallurgy)
  • YC(Tobacco)
  • YD(Telecommunication)
  • YY(Medical Device)
Database: 365,228(8 Aug 2026)
core word segmentation rules compound word processing tibetan segmentation morphological analysis special processing case auxiliary word agglutination processing international language processing word segmentation granularity tibetan case grammar sanskrit transcription processing volume numbers braille publications environment protect requirements department store customer relationship management scopethis standard
GB/T 36452-2018 in English

GB/T 36452-2018 in English

VALID

Specification on Tibetan segmentation for information processing

  • Issued on:2018-06-07
  • Implemented on:2019-01-01
  • File Format:PDF
  • Delivery:Via email within 1~3 business days
Price(USD): $250.00
$243.00
Standard No: GB/T 36452-2018
Document status: VALID
Title in English: Specification on Tibetan segmentation for information processing
Title in Chinese: 信息处理用藏文分词规范
Language: English
File Format: Electronic (PDF)
Delivery: Via email within 1~3 business days
Issued on: 2018-06-07
Implemented on: 2019-01-01
ICS Classification: 35.240.01-Applications of information technology in general
Chinese Classification: L70-Information processing technology in general
Professional Classification: GB-National Standard
Related Keywords: core word segmentation rules compound word processing
tibetan segmentation
morphological analysis special processing case auxiliary word agglutination processing
international language processing word segmentation granularity
tibetan case grammar sanskrit transcription processing
Related Topics: information office
GBT36452
GB/T 36452-2018
Tibetan
Information detection and processing

《GB/T 36452-2018信息处理用藏文分词规范》由TC28(全国信息技术标准化技术委员会)归口,主管部门为国家标准化管理委员会。


Introduction

Analysis of the Standard Technical Framework

Dimensions Standardized Features of Tibetan Comparison with Chinese Word Segmentation Differences in International Language Processing
Word Segmentation Granularity 1-5 syllable compound words are segmented as a whole 2-4 word words are the main ones Agglutinative languages require morphological analysis
Special Processing Case Auxiliary Word Agglutination Processing (4.25-4.29) Function Words are Segmented Independently Inflectional Languages Require Stem Extraction
Tag system GB/T 36337 part-of-speech tags ICTCLAS system Universal POS tags

Detailed explanation of core word segmentation rules

Compound word processing (Clause 4.2)

Tibetan noun compound words adopt the "overall segmentation" principle:

  • Monosyllabic noun + adjective → 1 participle unit
  • Disyllabic noun + adjective → 1 participle unit
  • The word-forming suffix "" is not segmented separately

Proper noun processing (Clause 4.3-4.7)

Typical case: "གངས་རིན་པོ་ཆེ" (Kang Rinpoche) as a whole word segmentation unit reflects the word formation characteristics of Tibetan place names


Technical Evolution Analysis

This standard innovatively introduces:

  1. Case particle combination markers (4.24-4.29): solve the agglutinative characteristics of Tibetan case grammar
  2. Sanskrit transcription processing (4.36): retain the ability to process religious documents
  3. Cross-language symbols (4.35): compatible with multi-language mixed texts

Implementation Suggestions

System Development Specifications

  • Use "/" as the word boundary character
  • Establish a mapping table between Tibetan word classes and general tag sets
  • Special character processing module needs to support Unicode Tibetan segment

Corpus Construction

The recommended annotation system includes:

Basic layerSyllable boundary annotation
Syntactic layerCase particle attachment marker
Semantic layerProper noun entity classification

Sample only — not a preview of GB/T 36452-2018
Page: 1 / 0
100%

Loading PDF document...

Error loading PDF. Please make sure the file is valid and try again.

We also recommend