LLMpediaThe first transparent, open encyclopedia generated by LLMs

Hong Kong Bilingual Corpus

⚠Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Linguistic Society of Hong Kong Hop 6 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Hong Kong Bilingual Corpus
NameHong Kong Bilingual Corpus
TypeCorpus
LocationHong Kong
LanguagesCantonese; English
CreatorsLinguists; Computational linguists; Corpus builders
Established2000s

Hong Kong Bilingual Corpus The Hong Kong Bilingual Corpus is a multilingual linguistic resource created to support research in Sociolinguistics, Computational linguistics, Corpus linguistics, and applied studies involving Cantonese and English language use in Hong Kong. The corpus aggregates spoken and written materials drawn from institutions such as University of Hong Kong, Chinese University of Hong Kong, Hong Kong Polytechnic University, and media outlets like South China Morning Post and RTHK, enabling comparative analyses across domains like media coverage, political discourse, and literary texts.

Introduction

The corpus was developed by collaborative teams including researchers from University of Hong Kong, Chinese University of Hong Kong, City University of Hong Kong, and international partners such as Stanford University, University of Cambridge, Massachusetts Institute of Technology, and Harvard University, reflecting influences from projects like the British National Corpus, the Corpus of Contemporary American English, and the Lancaster-Oslo/Bergen Corpus. It serves scholars studying code-switching phenomena between Cantonese and English language across contexts like legislative council proceedings, broadcasting interviews, and literary awards submissions.

Corpus Composition and Data Sources

The collection includes transcripts of spoken materials from sources such as Legislative Council of Hong Kong, RTHK, and university lecture recordings, alongside written texts drawn from newspapers including South China Morning Post, magazines like Time (magazine), and literary texts recognized by the Hong Kong Arts Development Council and the Bologna Prize for Children's Literature. It integrates material from public institutions such as Hong Kong Public Libraries, NGO reports by groups like Amnesty International and Human Rights Watch, and multilingual signage corpora sampled from districts like Central and Western District, Kowloon, and New Territories.

Annotation and Linguistic Features

Annotations include orthographic transcription of Cantonese syllables using schemes related to Jyutping, part-of-speech tags influenced by conventions from Penn Treebank, and morphosyntactic labels comparable to those used in the Universal Dependencies project. Pragmatic and discourse annotations draw on frameworks employed in studies by scholars associated with Princeton University, University of Oxford, and University of Toronto, marking features such as code-switching points, politeness strategies observable in interactions involving figures akin to Carrie Lam and coverage of events like the 2019–20 Hong Kong protests.

Creation and Processing Methodology

Data collection protocols were modeled after established corpora initiatives including the British National Corpus and the Corpus of Contemporary American English, with recording standards referencing practices used by archives like the British Library and the Library of Congress. Processing pipelines employed tools and toolkits produced by groups at Stanford University (for example, taggers and parsers), innovations from Google Research in speech recognition, and methods disseminated by the ACL and EMNLP communities to handle noise, transcription alignment, and anonymization.

Applications and Research Uses

Researchers have used the corpus to investigate code-switching patterns similar to studies at Columbia University, sociophonetic variation akin to work at University College London, and language contact phenomena studied by scholars from University of California, Berkeley. Applications span development of speech recognition models comparable to those by Apple and Microsoft, machine translation efforts in the vein of Google Translate, lexicography projects influenced by Oxford University Press dictionaries, and policy analyses addressing issues debated in bodies like the Legislative Council of Hong Kong.

Accessibility and Licensing

Access policies mirror norms from digital archives such as the British Library Sound Archive and institutional repositories at University of Hong Kong Libraries, combining open-access subsets with restricted-use data under licenses similar to those used by Creative Commons and university data-sharing agreements. Distribution mechanisms have involved partnerships with platforms used by ELRA and LDC to provide authenticated downloads and APIs for registered researchers.

Limitations and Future Development

Limitations include uneven representation across registers comparable to gaps noted in the British National Corpus, difficulties in automatic transcription for Cantonese tones as reported in studies from MIT and Carnegie Mellon University, and legal constraints tied to media rights held by organizations like TVB and South China Morning Post. Future development plans propose expanded sampling from community archives such as Hong Kong Heritage Museum, enhanced annotation layers inspired by the Universal Dependencies and OntoNotes initiatives, and collaborations with technology partners including Google, Microsoft Research, and regional institutions like Hong Kong Science and Technology Parks Corporation.

Category:Corpora