Address third round of cubic-dev-ai review feedback:
P3: The three binary format tests (docx, epub, odt) were nearly identical
in structure, differing only in the file extension. Consolidated into a
single test.each parameterized test. Also fixed inconsistent test name
(odt was missing the 'with -V adjacency' suffix). All three now use the
unified name pattern via %s substitution.
Address second round of cubic-dev-ai review feedback:
P1: detectCJKFont only scans raw file bytes for CJK characters, which
misses CJK content inside compressed/binary formats (docx, epub, odt,
pptx). For these formats, always pass a CJK font to avoid missing glyphs.
Text formats still use content-based detection.
P3: Test assertions only checked that CJKmainfont value exists in args
array but did not verify -V flag precedes it. Added assertCJKFontArg
helper that checks -V adjacency for all CJK font tests.
Changes:
- Add binaryFormats set for docx/epub/odt/pptx
- Binary formats always get Noto Sans CJK SC (cannot detect language
from raw bytes, SC is the safest default covering all CJK characters)
- Text formats continue using detectCJKFont for locale-aware detection
- Add assertCJKFontArg helper verifying -V precedes CJKmainfont value
- Add 3 new tests for binary formats (docx, epub, odt) with -V check
- 14 tests total, all passing
Address code review feedback from cubic-dev-ai:
P1: Unconditionally passing CJKmainfont broke non-CJK conversions in
environments without CJK fonts. Now the source file is scanned for CJK
characters and the font argument is only added when CJK content is
detected.
P2: Hardcoded Noto Sans CJK SC only covers Simplified Chinese. Now
detects the specific CJK language (Japanese Kana, Korean Hangul, or
Chinese Hanzi) and selects the appropriate font variant (JP/KR/SC).
Changes:
- Add detectCJKFont() that reads the source file and returns the
correct Noto Sans CJK variant, or null for non-CJK content
- CJK font argument is only added when detectCJKFont returns non-null
- Detection order: Japanese (Kana) > Korean (Hangul) > Chinese (Hanzi)
since Japanese also uses Kanji shared with Chinese
- Graceful fallback: if file can't be read, no CJK font is added
- Update tests to use temp files with real CJK content (11 tests total)
When converting documents containing Chinese, Japanese, or Korean
characters to PDF via pandoc/xelatex, the output was blank or missing
glyphs because no CJK-capable font was installed or configured.
Changes:
- Dockerfile: Add fonts-noto-cjk and texlive-lang-chinese packages
- pandoc.ts: Pass -V CJKmainfont=Noto Sans CJK SC to pandoc when
converting to pdf/latex, so xelatex uses a CJK-capable font
- pandoc.test.ts: Add 3 tests verifying CJK font argument is present
for pdf/latex output and absent for other formats
Fixes#547