Typing 顯 in Cangjie takes five keystrokes: A F M B C. None of them is a sound, none is a meaning, and the first two are not even the whole of the left-hand side. They are what is left after a rule deliberately throws most of it away.
This is how a character gets from its printed shape to those five letters, worked through on real characters, and where the rules stop being able to tell you.
The keys are shapes, and mostly not the shapes they look like
Twenty-four letters carry the method, grouped into four families. The first family is philosophical — 日 月 金 木 水 火 土 on A to G. The last is bodily — 人 心 手 口 on O, P, Q, R.
The two families in between are the ones that mislead. Their real names are abstract: 斜, 點, 交, 叉, 縱, 橫, 鉤 for the stroke family, and 側, 並, 仰, 紐, 方, 卜 for the shape family. The characters printed on the keycaps — 竹 戈 十 大 中 一 弓, then 尸 廿 山 女 田 卜 — are stand-ins, picked because they are common characters that happen to carry the right shape. H is not "bamboo". H is "slanting", and 竹 is merely the advertisement for it.
This distinction allows the method to work on any character, including unfamiliar ones, because it is a system for reading shapes, not for recognising known components.
Three characters, taken apart
Every character splits into a head and a body — 字首 and 字身 — and the split is positional, not semantic. The head is the leftmost part, or the topmost. It is emphatically not the dictionary radical.
The simplest case has one of each.
| Character | Head | Body | Code |
|---|---|---|---|
| 明 | 日 (A) | 月 (B) | AB |
| 草 | 廿 (T) | 早 → 日 十 (A J) | TAJ |
| 荷 | 廿 (T) | 何 → 亻 (O), then 可 as 一 … 口 (M R) | TOMR |
荷 is where the shape of the system appears. The body is itself a head-and-body pair: 亻 is a single letter, so it costs one code, and 可 then gets two — its first shape and its last, with the 丁 in the middle simply not taken.
Reading order is fixed: left before right, top before bottom, outside before inside.
Why the answer is never longer than five letters
The cap is five, and it is derived rather than chosen. Working over a corpus of sixty thousand characters, the method's author counted the distinct parts: 594 different heads, and 9,897 different bodies.
Twenty-five keys give 625 possible one-or-two-letter sequences, which covers 594 heads. They give 15,625 three-letter sequences, which covers 9,897 bodies. So a head takes one or two codes, a body takes one to three, and the total lands between one and five. Fewer would not separate the parts; more would be waste.
A key consequence, and one that makes the method learnable, is that a given component always produces the same code. 其 is T M C in 騏, in 棋 and in 箕. Learn a part once.
What gets thrown away
When a part is too big for its allowance, the rule is not to compress it but to drop the middle. A head longer than two codes keeps its first and its last. A body that is a single connected shape keeps its first, second and last.
Two examples show how aggressively the rules discard information.
顯 has a left-hand side that would code in full as 日 く 厶 く 厶 灬 — six shapes. It keeps the first and the last: A F. The right-hand 頁 then codes normally as 一 月 金, and the whole character is AFMBC.
鬱 is worse. Its top block, the 木 缶 木 arrangement, would run 木 ⺊ 十 凵 木. It keeps 木 and 木 — D D — and the rest of the character supplies B U H. Five letters for one of the densest characters in common use.
There is a second kind of dropping, and it is the more elegant one. When the shape you are taking as a last code is an enclosure — 口, 冂, コ, 匚, 廿, 刀, 几, 瓦, 勹, 乃 — you take the enclosure and discard whatever sits inside it. 粵 codes as HWMVS: the 米 inside the mouth is simply not there. The manual's own phrasing is to take the mouth and not the rice.
The characters that do not obey
About ninety-five per cent of Chinese characters code by the rules above. The remaining five per cent are handled by lists, and the lists are closed — the manual says so in as many words, and warns that a reader who invents his own exceptions will not produce correct codes.
There are five kinds. Eight compound heads and seven compound shapes always take just their first and last codes wherever they appear. A set of shapes counts as difficult and takes the 難 key: 身 is H X H, 龜 is H X U, 舅 is H X W K S. Five shapes — 木, 火, 戈, 大, 七 — are special when another stroke runs through them, so 東 is DW, taking 木 whole and then reading the remainder as 田 rather than 日. And then there are the duplicates.
When two characters want the same code
Nothing in the system prevents a collision, so collisions are resolved by fiat: the commoner character keeps the plain code, and the other takes a 重 prefix.
The canonical demonstration is three characters that differ by almost nothing. In our own lookup table, 巳 returns RU and 己 returns SU — a mouth versus a flank, which is a real visual difference. 已 sits between them and returns SU as well, with XSU alongside it. It collides with 己, and only the prefix separates them.
The same table shows 九 and 夷 both returning KN, with 夷 carrying XKN as an alternative. Collisions between visually unrelated characters are not a defect, but an expected outcome of mapping tens of thousands of characters onto only twenty-five keys.
The 九 case has a history. Some users argued it should be K U on the grounds of how the character is written. The method's author refused, on the grounds that the rules never mention writing order at all — and pointed out that if they did, the first code of 申 would have to be 田.
Which version you are typing
There are two Cangjie generations in live use and they disagree, so the same character can have two correct answers depending on what your keyboard is running.
Two characters tell them apart, because both changed a 卜 into a 尸:
| Character | 3rd generation | 5th generation |
|---|---|---|
| 面 | MWYL | MWSL |
| 非 | LMYYY | LMSY |
Type 面 and count the keystrokes. Four with an S in third place is the fifth generation; four ending Y L is the third.
The disagreement is narrow but real. Taiwan's government character database publishes an official Cangjie table of 27,452 characters, and 1,061 of them — 3.9% — carry two accepted codes rather than one. That is where the generations part company. Our own lookup holds 21,469 characters and is the fifth generation throughout, which is why it answers MWSL for 面.
The Cangjie field in Unicode's own character database should be treated with caution, as it is not internally consistent. It gives fifth-generation codes for some characters and third-generation codes for others, including a handful matching neither published column. It is widely described as being the third generation. It is not, and a tool built on that assumption will be quietly wrong on a scattering of characters.
The faster, worse version
速成 — Quick — is not a different decomposition. It uses the same shapes, the same rules and the same table, and then types only the first and last code of whatever Cangjie would have produced.
So 粵 becomes H S, 鬱 becomes D H, 顯 becomes A C, and 明 stays A B because it only had two codes to begin with. Nothing longer than two keystrokes.
The trade-off is simple: far fewer keystrokes, but a candidate list to select from for almost every input. Two codes over twenty-five keys is roughly six hundred possible inputs for tens of thousands of characters, so almost every Quick input produces a candidate list to choose from. It is much easier to learn and much slower to type without looking.
What this is worth knowing for
If you are learning the method, the head-and-body split is where the errors live, not the key mapping. The keys take an afternoon. Deciding what counts as the head — and remembering that it is a position rather than a radical — is the part that takes practice.
If you are building anything on a Cangjie table, check which generation it is before you trust it, and do not assume Unicode's field is consistent.
If a character's code seems wrong, the cause is less likely a misremembered key than one of the five per cent of characters handled by an exception list.
Sources, and what could not be pinned down
The rules, the code limits, the exception lists and the head-and-body definitions are from the fifth-generation manual written by the method's author, read from the scanned edition published page by page online. The derivation of the five-code limit — 594 heads and 9,897 bodies measured over a sixty-thousand character corpus — is the manual's own arithmetic. The 27,452-character official table with its 1,061 dual-coded characters is Taiwan's national character database, published as open government data. Every character code quoted here was checked against the table this site's own lookup uses, which reconciles exactly against its upstream source. All were read on 20 September 2026.
Three points could not be definitively settled. The number of auxiliary shapes is usually given as seventy-seven on top of the twenty-four letters; the manual caps letters and auxiliaries together at about a hundred, and a direct count off the scans did not reproduce the usual figure, so it is reported here as a secondary claim rather than a verified one. Which generation any given operating system ships today could not be confirmed by inspection. And the method's licensing is often described as public domain, which traces to a newspaper announcement renouncing the patent in 1982 rather than to a licence anyone has published — a description worth repeating with its provenance attached, not as a settled legal fact.
Usage figures deserve the same care. A 2011 Taiwanese survey put Cangjie at under ten per cent, behind both phonetic input and a rival shape-based method. The figure most often quoted for Hong Kong — that a clear majority of entrants used Cangjie — comes from an inter-school typing competition, which is a self-selected sample of people who type quickly for sport. It is evidence about competitors, not about a population.