Through the 1990s and beyond, computers in Taiwan and Hong Kong used the Big5 encoding for traditional Chinese. But for Cantonese speakers, the standard had a large gap. Eleven of the characters written Cantonese uses most could not be typed: 嘅 咗 哋 喺 啲 嘢 嚟 嘞 攰 揾 瞓. The missing characters included the possessive 嘅, the completed-action marker 咗, and the location word 喺—the grammar that makes a sentence Cantonese rather than Mandarin. Hong Kong eventually filled the gap with its own supplementary character set.
We checked 22 of the most common written-Cantonese characters against the major encodings to see which were missing from Big5, and what was done about it. For how Big5 and the mainland's GB encodings differ in general, see our guide to why a character that works in GBK can break in Big5. To turn Mandarin text into written Cantonese, there is our Cantonese converter.
Twenty-two characters, checked
Written Cantonese is the everyday spoken language of Hong Kong put on paper. It appears in advertising, comics, forums, messaging and some newspaper columns; formal writing in Hong Kong usually uses standard written Chinese, read aloud in Cantonese. The characters below are the ones that mark written Cantonese as Cantonese.
| Character | Role | In Big5? |
|---|---|---|
| 嘅 | possessive, like Mandarin 的 | no |
| 咗 | completed action, like 了 | no |
| 哋 | plural for pronouns (我哋, "we") | no |
| 喺 | "at, in", like 在 | no |
| 啲 | "some"; plural marker | no |
| 嘢 | "thing, stuff" | no |
| 嚟 | "to come", like 来 | no |
| 嘞 | sentence-final particle | no |
| 攰 | "tired" | no |
| 揾 | "to look for", like 找 | no |
| 瞓 | "to sleep", like 睡 | no |
| 冇 係 佢 唔 乜 咁 睇 諗 咪 喎 噉 | "not have", "to be", "he/she", "not", "what", "so", "to look", "to think", and particles | yes |
Eleven of twenty-two. The eleven that Big5 lacks are, for the most part, characters that written Cantonese needs and standard written Chinese has little use for.
We ran the same check on a larger set: the 103 characters that appear only on the Cantonese side of our converter's dictionary. Big5 can write 90 of them. The thirteen it cannot write include ten of the eleven above — the dictionary does not use 嘞 — along with 啱, 嗰 and 攞.
Why Big5 did not have them
Big5 was drawn up in Taiwan in 1984, by an industry group rather than a government, to let computers write traditional Chinese. Its 13,053 characters were chosen for the written standard Chinese used in Taiwan. Nothing in that brief called for the characters that exist to write Cantonese grammar. They were left out not because of a decision against Cantonese, but because colloquial language fell outside the project's scope.
For Hong Kong that mattered. Taiwan and Hong Kong shared Big5 because both use traditional characters, but Hong Kong also writes personal names, place names and colloquial Cantonese that Taiwan's set never included.
Hong Kong's supplement
The Hong Kong government's answer began as a supplement to Big5 and became the Hong Kong Supplementary Character Set, or HKSCS. Its 2016 edition holds 5,033 characters (4,591 of them Chinese) to cover what Hong Kong needed to write, including personal names, place names, and colloquial Cantonese. The government does not break the set down by purpose, noting that one character can serve several.
All eleven missing characters have HKSCS code points. Unicode's database reflects this. It gives each of the eleven an HKSCS code and no Big5 code, while the other eleven characters in the table carry ordinary Big5 codes. In a system that knew Big5 but not HKSCS, 嘅 or 咗 had nowhere to go, and a Cantonese sentence came out with boxes or question marks in its most important places.
Why the mainland's encoding could write them
The mainland's GBK encoding, by contrast, can write all 22 characters, and all 103 in the larger set. That was not a decision to serve Cantonese. GBK was built in the 1990s to include all 20,902 Chinese characters in the Unicode standard of the time. Those 20,902 characters, which Unicode had consolidated from earlier national standards, already included the ones needed for Cantonese.
The older mainland standard, GB2312, writes only 5 of the 22. It was designed for simplified characters in 1980 and has no more room for Cantonese than Big5 does.
| Encoding | Of 22 core characters | Of 103 Cantonese-only characters |
|---|---|---|
| GB2312 (mainland, 1980) | 5 | 35 |
| Big5 (Taiwan, 1984) | 11 | 90 |
| GBK (mainland, 1990s) | 22 | 103 |
| Big5 with HKSCS | 22 | 103 |
| UTF-8 | 22 | 103 |
What it means for text today
Text written in UTF-8, the standard for nearly all modern software, has none of these problems because every character in the table has a single, universal code. The gap shows up in older files and systems. A Hong Kong document saved in plain Big5, or opened by software that assumes plain Big5, can lose exactly the characters that carry its Cantonese grammar. If a Cantonese file shows boxes in place of 嘅 or 咗, opening it as Big5-HKSCS usually fixes it.
Our converter had its own problem, fixed on 21 September 2026. It used to replace single characters wherever they appeared, so 大家 came out as 大屋企 and 小说 as 小講. It now protects common words before converting.
Where the figures come from
We tested each character in each encoding on 21 September 2026, using the standard libraries of a widely used programming language. The Unicode relationships are from the Unicode Consortium's Unihan database, where the Big5 field is empty for exactly the eleven characters above. The database's Hong Kong field, carried until recent versions, gives each of them an HKSCS code. The HKSCS figures are from the Hong Kong government's page on the set. The 103-character set is the Cantonese side of our converter's dictionary, and describes that dictionary rather than written Cantonese as a whole.