LLMpediaThe first transparent, open encyclopedia generated by LLMs

List of Commonly Used Characters in Modern Chinese

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Chinese characters Hop 6 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

List of Commonly Used Characters in Modern Chinese
NameCommonly Used Chinese Characters
LanguageChinese
ScriptsSimplified Chinese, Traditional Chinese
RegionsMainland China, Taiwan, Hong Kong, Macau, Singapore
Established20th century

List of Commonly Used Characters in Modern Chinese

The list of commonly used characters is a standardized inventory that informs literacy, pedagogy, publishing, and computing across Greater China and the global Sinophone world. It intersects policy decisions from institutions such as the Ministry of Education (People's Republic of China), curriculum frameworks used in Beijing, assessment regimes like the Hanyu Shuiping Kaoshi, and corpus projects connected to universities in Taipei and Hong Kong. The compilation draws on historical reform initiatives linked to figures such as Cai Yuanpei and movements like the New Culture Movement, while also informing contemporary digital platforms developed by entities including Baidu and Tencent.

Overview and Criteria for Inclusion

In setting criteria, policymakers and scholars compare standards promulgated by the State Council (People's Republic of China), advisories from the Ministry of Education (Republic of China), and guidelines used by publishers in Shanghai and Guangzhou. Inclusion typically depends on frequency evidence from corpora curated at institutions such as Peking University, National Taiwan University, and The Chinese University of Hong Kong, on character presence in canonical works like the Analects, the Book of Songs, and modern media produced by organizations like Xinhua News Agency and China Central Television. Practical criteria also reflect input from examination boards for tests administered by bodies linked to Confucius Institute Headquarters (Hanban), standards committees in Singapore, and international education groups in New York and London.

Frequency Categories and Ranking Methodologies

Ranking methods compare token counts from corpora assembled at research centers such as Tsinghua University, University of Oxford, and Harvard University. Approaches include corpus-based frequency lists used by projects like the Lancaster Corpus analogues, statistical models deployed by teams at Stanford University and Massachusetts Institute of Technology, and annotation conventions influenced by the editorial practices of publishers such as Commercial Press. Researchers may apply Zipfian modelling as discussed in works from Princeton University and frequency-normalization techniques used in datasets managed by Google and Microsoft Research.

Lists by Frequency and Usage (Top 100, 500, 1000)

Published lists vary: official inventories issued by the State Council (People's Republic of China) inform many Top 100 and Top 500 compilations, while academic lists from Peking University and National Taiwan University often extend to Top 1000. Media organizations like People's Daily and tech companies including Alibaba use tailored shorter lists for UI design, whereas language assessment creators at Beijing Language and Culture University and test developers associated with the Hanyu Shuiping Kaoshi assemble lists aligned with examination needs. Lexicographers at institutions like the Academia Sinica produce concordances that feed into educational textbooks from houses such as the Commercial Press and international publishers in Cambridge and New York.

Characters by Part of Speech and Function

Analyses categorize characters according to roles informed by studies at departments within Peking University, Fudan University, and The Chinese University of Hong Kong. Function tagging schemes are applied in projects linked to research groups at Stanford University, Tsinghua University, and Columbia University to differentiate pronouns, particles, numerals, and classifiers as seen in corpora reflecting discourse from institutions including Xinhua News Agency, the BBC Chinese Service, and scholarly journals published by Springer and Oxford University Press.

Regional and Script Variants (Mainland, Taiwan, Hong Kong, Simplified vs Traditional)

Regional standards differ: Mainland lists based on simplified characters are promulgated by the Ministry of Education (People's Republic of China), while Taiwan's traditional-character lists come from the Ministry of Education (Republic of China), and Hong Kong conventions are influenced by the Education Bureau (Hong Kong). Script reform debates echo historical policy choices debated during the era of the Republic of China (1912–1949) and later administrative decisions by the People's Republic of China. Publishing houses in Taipei, printers in Hong Kong, and digital platforms run by Apple Inc. and Google must reconcile variant glyphs for cross‑regional interoperability.

Educational and Pedagogical Uses (Primary, Secondary, HSK)

Curricula developed for primary and secondary schooling reference lists adopted by the Ministry of Education (People's Republic of China), teacher-training programs at institutions like Beijing Normal University and National Taiwan Normal University integrate graded character lists, and proficiency exams such as the Hanyu Shuiping Kaoshi and curricula promoted by Confucius Institutes employ frequency-driven syllabuses. International textbook series published by houses in Cambridge and Oxford adapt lists for learners in settings overseen by consulates or cultural institutes in cities like Sydney, Toronto, and Berlin.

Computational and Corpus-Based Resources

Digital resources include large corpora maintained by research centers at Peking University, language technology datasets released by Microsoft Research Asia and Google Research, and annotation projects coordinated with universities such as Tsinghua University and Stanford University. Software tools for segmentation, input method editors by companies like Sogou and Baidu, and open-source initiatives hosted by communities connected to GitHub rely on these lists for tokenization, indexing, and OCR work used in collaborations with libraries such as the National Library of China and archives in Taipei.

Category:Chinese language