Explainer Developer Tools 5 min read

Why common Cantonese characters like 嘅 and 咗 could not be typed in Big5

Eleven of the most-used written Cantonese characters, including 嘅, 咗 and 喺, could not be typed in Big5, the traditional Chinese encoding used in Taiwan and Hong Kong. Here are all eleven, what Hong Kong's supplement did about them, and why the mainland's GBK could write them anyway.

Kenji Tanaka
Developer Tools & Cloud Analyst
Published 22 Sep 2026, 9:15 AM (SGT)
Share:
Shop signs in traditional Chinese characters lit up along a Hong Kong street at night Shop signs in traditional Chinese characters lit up along a Hong Kong street at night Photo by Tomás Monteiro on Pexels
Advertisement

Through the 1990s and beyond, computers in Taiwan and Hong Kong used the Big5 encoding for traditional Chinese. But for Cantonese speakers, the standard had a large gap. Eleven of the characters written Cantonese uses most could not be typed: 嘅 咗 哋 喺 啲 嘢 嚟 嘞 攰 揾 瞓. The missing characters included the possessive 嘅, the completed-action marker 咗, and the location word 喺—the grammar that makes a sentence Cantonese rather than Mandarin. Hong Kong eventually filled the gap with its own supplementary character set.

We checked 22 of the most common written-Cantonese characters against the major encodings to see which were missing from Big5, and what was done about it. For how Big5 and the mainland's GB encodings differ in general, see our guide to why a character that works in GBK can break in Big5. To turn Mandarin text into written Cantonese, there is our Cantonese converter.

Twenty-two characters, checked

Written Cantonese is the everyday spoken language of Hong Kong put on paper. It appears in advertising, comics, forums, messaging and some newspaper columns; formal writing in Hong Kong usually uses standard written Chinese, read aloud in Cantonese. The characters below are the ones that mark written Cantonese as Cantonese.

CharacterRoleIn Big5?
possessive, like Mandarin 的no
completed action, like 了no
plural for pronouns (我哋, "we")no
"at, in", like 在no
"some"; plural markerno
"thing, stuff"no
"to come", like 来no
sentence-final particleno
"tired"no
"to look for", like 找no
"to sleep", like 睡no
冇 係 佢 唔 乜 咁 睇 諗 咪 喎 噉"not have", "to be", "he/she", "not", "what", "so", "to look", "to think", and particlesyes

Eleven of twenty-two. The eleven that Big5 lacks are, for the most part, characters that written Cantonese needs and standard written Chinese has little use for.

We ran the same check on a larger set: the 103 characters that appear only on the Cantonese side of our converter's dictionary. Big5 can write 90 of them. The thirteen it cannot write include ten of the eleven above — the dictionary does not use 嘞 — along with 啱, 嗰 and 攞.

Why Big5 did not have them

Big5 was drawn up in Taiwan in 1984, by an industry group rather than a government, to let computers write traditional Chinese. Its 13,053 characters were chosen for the written standard Chinese used in Taiwan. Nothing in that brief called for the characters that exist to write Cantonese grammar. They were left out not because of a decision against Cantonese, but because colloquial language fell outside the project's scope.

For Hong Kong that mattered. Taiwan and Hong Kong shared Big5 because both use traditional characters, but Hong Kong also writes personal names, place names and colloquial Cantonese that Taiwan's set never included.

Advertisement

Hong Kong's supplement

The Hong Kong government's answer began as a supplement to Big5 and became the Hong Kong Supplementary Character Set, or HKSCS. Its 2016 edition holds 5,033 characters (4,591 of them Chinese) to cover what Hong Kong needed to write, including personal names, place names, and colloquial Cantonese. The government does not break the set down by purpose, noting that one character can serve several.

All eleven missing characters have HKSCS code points. Unicode's database reflects this. It gives each of the eleven an HKSCS code and no Big5 code, while the other eleven characters in the table carry ordinary Big5 codes. In a system that knew Big5 but not HKSCS, 嘅 or 咗 had nowhere to go, and a Cantonese sentence came out with boxes or question marks in its most important places.

Why the mainland's encoding could write them

The mainland's GBK encoding, by contrast, can write all 22 characters, and all 103 in the larger set. That was not a decision to serve Cantonese. GBK was built in the 1990s to include all 20,902 Chinese characters in the Unicode standard of the time. Those 20,902 characters, which Unicode had consolidated from earlier national standards, already included the ones needed for Cantonese.

The older mainland standard, GB2312, writes only 5 of the 22. It was designed for simplified characters in 1980 and has no more room for Cantonese than Big5 does.

EncodingOf 22 core charactersOf 103 Cantonese-only characters
GB2312 (mainland, 1980)535
Big5 (Taiwan, 1984)1190
GBK (mainland, 1990s)22103
Big5 with HKSCS22103
UTF-822103

What it means for text today

Text written in UTF-8, the standard for nearly all modern software, has none of these problems because every character in the table has a single, universal code. The gap shows up in older files and systems. A Hong Kong document saved in plain Big5, or opened by software that assumes plain Big5, can lose exactly the characters that carry its Cantonese grammar. If a Cantonese file shows boxes in place of 嘅 or 咗, opening it as Big5-HKSCS usually fixes it.

Our converter had its own problem, fixed on 21 September 2026. It used to replace single characters wherever they appeared, so 大家 came out as 大屋企 and 小说 as 小講. It now protects common words before converting.

Where the figures come from

We tested each character in each encoding on 21 September 2026, using the standard libraries of a widely used programming language. The Unicode relationships are from the Unicode Consortium's Unihan database, where the Big5 field is empty for exactly the eleven characters above. The database's Hong Kong field, carried until recent versions, gives each of them an HKSCS code. The HKSCS figures are from the Hong Kong government's page on the set. The 103-character set is the Cantonese side of our converter's dictionary, and describes that dictionary rather than written Cantonese as a whole.

Advertisement
Kenji Tanaka
Developer Tools & Cloud Analyst

Kenji Tanaka covers developer tools, cloud platforms, DevOps, CI/CD, and software supply-chain topics for RECATOOLS.

View author profile → · Editorial policy

About this byline Kenji Tanaka is a RECATOOLS editorial persona for developer tools, cloud, DevOps, and software supply-chain coverage. Articles are produced and reviewed under RECATOOLS editorial supervision.

Corrections policy

Advertisement