Repository navigation
fix(docs): make the Chinese pages searchable by word - #518
Conversation
Chinese runs no spaces, so the indexed text carried no word boundary at all: the separator broke on whitespace and hyphens, which left each run as a single token, and a reader who typed a word rather than a whole heading found nothing. Two things had to line up. The build segments the text with jieba, which marks every boundary with a zero-width space, and the separator now breaks on that mark: the whitespace class the theme defaults to does not cover it, so the marks stayed buried inside the token. jieba was reaching the build only as a transitive dependency of something else, so it was absent wherever the docs group was installed on its own, which is every build that ships. Measured on a build made in a docs-only environment, searching the Chinese site: a word from the middle of a heading goes from 0 results to 36, and two single-word headings from 5 to 12 and from 5 to 10. A compound the segmenter splits goes the other way, 5 to 0, because the theme does not segment the query and a run that is not itself a word no longer matches. The guard for this passed while the site was broken. It took its query from the page's own heading, zero-width marks and all, which matches the stored token exactly and is a string no reader can type. It strips the marks now, and a second guard searches a word taken from the middle of a title, which can only be found if the run was indexed as words. Co-authored-by: Claude (claude-opus-5[1m]) <noreply@anthropic.com>
gloryfromca
left a comment
There was a problem hiding this comment.
No blockers; this can merge as far as I am concerned.
I reviewed the full diff, AGENTS.md constraints, the Material search plugin and shipped worker path, the preceding language-index split and relevant history, dependency/lock compatibility, test changes, backward compatibility, and architecture impact. The added dependency is in the existing docs group, the separator preserves the active English behavior while adding the Jieba boundary, and the new checks cover both halves of that contract without weakening the prior language-isolation assertion. No architecture constraint is implicated.
Verification: the search/index integration suite passed 3 tests; the docs and source-language unit suites passed 31 tests; a fresh strict build produced 149 unmarked English entries and 146/146 marked Chinese entries with the configured separator; the shipped Material worker found the source page when queried with a middle title word (8 results); the source-language and large-file scripts, uv lock --check, and git diff --check passed. The browser suite could not run in this container: it first skipped for absent Chrome, and bundled Chromium could not start because the host lacks libatk-1.0.so.0; I did not count that environment failure as a passing test. All GitHub checks are green.
ZuyiZhou
left a comment
There was a problem hiding this comment.
Approving. I checked whether the bug is real before checking whether the fix is, and it is: the currently published Chinese index at evermind-ai.github.io/Raven/zh/search/search_index.json has 146 entries, zero of them carrying a U+200B boundary, and its separator is still the whitespace-and-hyphen default. This is shipped and live, not theoretical.
Root cause confirmed in the vendored code rather than inferred. material/plugins/search/plugin.py lines 36-38 are try: import jieba / except: jieba = None -- a silent optional import, so a missing jieba produces no segmentation and no warning, which is exactly how this shipped unnoticed. And jieba genuinely was absent from the docs group: it reached the lock only as a dependency of everos.
Verified after the fix:
- All 146 Chinese entries carry boundaries; the English index carries none, which is correct; the separator is updated in the built index.
- The claim in the comment about the whitespace class is right, and worth saying because it is the whole reason the second half is needed. JavaScript's whitespace class runs from U+2000 to U+200A and stops there, so U+200B falls outside it and the separator has to name the character explicitly. Declaring jieba without this would have segmented the text and still found nothing.
- The new build guard fails with the separator reverted, reporting that it does not break on the word boundary. 8 tests pass with both halves in place.
The paragraph admitting the old guard passed while the site was broken is the most valuable part of this description, and the diagnosis is right: the query was lifted from the page's own heading with the boundary marks still in it, so it matched the stored token exactly, and no reader can type that string. Stripping the marks and searching a word taken from the middle of a title is the correct replacement, because that word can only be found if the run was indexed as words.
Two notes, neither blocking:
The separator assertion checks a regex that lunr will evaluate in JavaScript using Python's engine. It is sound for this separator, where U+200B is a literal member of the class, but the two engines do not agree in general -- the whitespace class being the very difference this fix turns on. The browser guard is what actually covers the behaviour.
Both guards live in tests/integration/, which norecursedirs excludes from every CI run. I raised this on #516 and this pull request is the argument for it: a live production bug in a tree whose guards run only on the author's machine. Neither file needs a browser for the build-level assertions, and the docs build job already has the docs group -- though note that if the search file moves into CI, the fixtures' pytest.skip on a failed build has to become a hard failure, or the job goes green on a site that did not build.
The disclosed regression is the right call to disclose: a compound the segmenter splits is no longer findable, because the theme does not segment the query. A custom dictionary is the fix for that, and rightly not in this change.
Summary
Chinese runs no spaces, so the indexed text carried no word boundary at all.
The search separator broke on whitespace and hyphens, which left each Chinese
run as a single token, and a reader who typed a word rather than a whole
heading found nothing.
Two things had to line up, and neither works without the other:
zero-width space. jieba was reaching the build only as a transitive
dependency of another group, so it was absent wherever the docs group was
installed on its own, which is every build that ships. It is declared now.
to does not cover it, so the marks stayed buried inside the token.
The guard for this passed while the site was broken, which is worth saying
plainly. It took its query from the page's own heading, boundary marks and
all, and a query carrying those marks matches the stored token exactly. No
reader can type that string. The query strips the marks now, and a second
guard searches a word taken from the middle of a title, which can only be
found if the run was indexed as words rather than kept whole.
Type
Verification
Measured on a site built inside an environment holding only the docs group,
served locally and driven in a real browser, searching the Chinese site.
Repeated three times with identical counts:
The same measurement against the published site reproduced the "before"
column exactly, which is what says the local environment matches the one
that builds it.
uv run pytest tests/integration/test_docs_site_search_e2e.py tests/integration/test_docs_site_controls_e2e.py tests/test_docs_site.py-> 34 passedbuild guard reports that the separator does not break on the boundary and
the browser guard finds nothing for its word
make lint-python,make lint-imports,make lint-deps,make lint-types-> exit 0, run against an environment synced the way CI installs, not the
working one
make check-commits,make check-source-language,make check-large-files,and
pre-commit run --from-ref origin/main --to-ref HEAD-> exit 0mkdocs build --strict-> built, 146 of 146 Chinese entries carry wordboundaries
Risk
User-visible: Chinese search returns results for ordinary words where it
returned none. One case gets worse: a compound the segmenter splits into
parts is no longer found, because the theme does not segment the query, so a
run that is not itself a word no longer matches any token. A custom
segmenter dictionary would recover those; it is not in this change.
Rollback: revert the commit. The two halves are one line of mkdocs.yml and
one dependency entry, and dropping either one restores the previous
behaviour on its own.
Related Issues
N/A