← BlogSignage-style card reading: why we built suo for CJK

Why we built suo for CJK

2026-09-26

Type 清华 into most search boxes and, if the content only has 清华大学 indexed, you get nothing. Type 稍後閱 — partway toward 稍後閱讀 — and the same thing happens. Not because the content isn't there, but because of how the search engine split the text into words in the first place.

The problem is word segmentation, not translation

English tokenizes almost for free: whitespace tells you where one word ends and the next begins. Chinese and Japanese don't have that signal. A segmenter has to guess where the word boundaries are, using a dictionary and statistical models — and it's genuinely good at this. Intl.Segmenter, running in the same edge runtime our engine runs in, correctly splits 清华大学的自然语言 into 清华大学 | 的 | 自然 | 语言.

The problem shows up at the edges, in exactly the moment search matters most: while someone is still typing.

  • A partial query inside a longer indexed word. 清华 is not, on its own, a segmented word — it's the first half of 清华大学. A search engine that only indexes whole segmented words has nothing to match 清华 against, even though a human reading it knows exactly what's being searched for.
  • A half-typed query, mid-word. Typing toward 稍後閱讀 (read later), the moment you've typed 稍後閱, a segmenter looking at that fragment in isolation splits it as 稍後 | 閱 — and 閱 alone, one character, usually isn't indexed as a meaningful token. The half-typed query effectively disappears.

This isn't a bug in the segmenter. Segmentation is doing exactly what it's supposed to do: find real words. The problem is that search-as-you-type produces text that isn't a real word yet, by definition, on every keystroke until the word is finished.

Bigrams, underneath the words

suo indexes two token streams for every searchable attribute, not one:

  1. ICU words — the dictionary-segmented tokens, exactly as before. These are precise, and they drive exact-match ranking.
  2. Character bigrams and unigrams — every overlapping two-character (and single-character) slice of the same text. 清华大学 becomes 清华, 华大, 大学 (plus the four single characters) in addition to the whole-word token.

A query is checked against both streams. If it happens to be a complete, correctly segmented word, the ICU match wins and ranks highest — that's still the most precise signal available. But if it's a fragment — 清华, or a half-typed 稍後閱 — the bigram stream still has it, because 清华 is a genuine two-character slice of 清华大学, and 稍後 and 後閱 are genuine slices of 稍後閱讀. The query matches, ranked a tier below an exact word match, but it matches.

This is the same principle Meilisearch's charabia tokenizer uses via jieba's sub-word emission, and part of why Meilisearch has historically led on CJK relevance — it's the bar we measured our own engine against before shipping.

What this looks like while typing

| You type | Looking for | Whole-word only | suo (bigrams) | | --- | --- | --- | --- | | 清华 | 清华大学 (Tsinghua University) | No results — 清华大学 is one indexed word; 清华 isn't a match | Finds it — 清华 is a bigram of 清华大学 | | 稍後閱 | 稍後閱讀 (read later) | No results — segmenter splits 稍後 \| 閱, drops the fragment | Finds it — 稍後 and 後閱 are bigrams of 稍後閱讀 | | 東京 | 東京都の天気予報 (Tokyo's weather forecast) | No results — only a full 東京都 matches | Finds it — 東京 is a bigram inside 東京都 |

Highlighting without breaking the word apart

A bigram match can land in the middle of a word — so the highlighted fragment has to look like emphasis on part of a word, not two words with a gap between them. A highlight mark that starts or ends inside a CJK run gets no inline padding and no margin; padding only applies next to Latin text or whitespace. And highlights are always built from the original text by character offset, never through a generic full-text-search highlight() function, which tends to insert spaces into CJK text and drop punctuation.

One more fold: traditional and simplified

Bigrams solve recall inside a word. They don't, by themselves, solve a zh-CN reader typing 讨论 against content written in zh-TW as 討論. For that, normalization runs a deterministic traditional-to-simplified character fold before segmentation and indexing, so a query in either script matches content in both.

Where this leaves whole-word matching

Nothing here replaces ICU segmentation — it sits underneath it, as a recall floor. A complete, correctly segmented word is still the strongest, highest-ranked signal in the ranking cascade; bigrams exist specifically for the moments segmentation alone would return nothing at all: a query still being typed, or a query shorter than a full word. Getting both right — precision when the query is complete, recall while it isn't — is what "CJK done right" means to us, and it's the first thing we measured before anything else in suo shipped.

Try it yourself in the playground: type a partial CJK query against a real index and watch the match land inside the word before you've finished typing it.