Skip to main content

Prevent broken line wrapping in CJK and Thai app text

2026-10-07

Prevent broken line wrapping in CJK and Thai app text

CJK Thai line breaking can fail even when every container grows correctly. Japanese punctuation may land at an unreadable line edge, and Thai text may split inside a word because ordinary spaces do not mark each boundary. Hardcoded newlines hide the bug at one width. It returns on another device or at a larger text size. Keep translation text independent of screen width and preserve its language, then choose a line-break policy for each content role. Tests with real scripts at several widths catch the failures that an English screenshot misses.

Why flexible layouts do not choose good break points

Text expansion and line breaking solve different problems. Flexible constraints decide how much room a label may use. A line-breaking engine decides where a line may end inside that room. A view can have unlimited vertical space and still produce a bad break.

The Unicode Line Breaking Algorithm defines classes and rules for allowed, prohibited, and required break opportunities. It also identifies Southeast Asian scripts, including Thai, as needing language-specific contextual analysis. That matters because a renderer cannot always find a readable Thai word boundary by looking for a space.

Japanese has another set of constraints. Punctuation, small kana, opening brackets, closing brackets, and other character classes do not all behave like Latin words separated by spaces. W3C's Requirements for Japanese Text Layout documents line-start and line-end restrictions and several strictness levels used in Japanese composition.

Most wrapping bugs come from a mismatch between those rules and the renderer:

  1. The app treats every space as a possible break and every no-space run as one indivisible token.
  2. The renderer breaks anywhere that fits without applying language-specific punctuation or word-boundary rules.
  3. Translators insert newlines for a screenshot width, coupling the translation to one device and text scale.

Manual breaks often survive review because the screenshot looks correct. They fail later on a smaller phone, in landscape, with accessibility text enabled, or after a copy edit changes one word.

Keep width decisions out of translation resources

A translation resource should express the message, not the dimensions of the current component. Remove manual line breaks that exist only to make one screenshot look balanced. Retain hard breaks only when they carry meaning, such as a postal address, a poem, a code sample, or a product requirement for separate paragraphs.

Store language metadata beside remote or user-selectable content. A renderer that receives only a string must guess which dictionary and tailoring rules apply. The app's effective UI locale may be a reasonable default for interface copy, but remotely managed content may have its own resolved content language. Pass that resolved language through the rendering boundary rather than deriving it again from the device region.

Do not bulk-insert invisible separators into Thai or Japanese text as a universal repair. An invisible character becomes persistent content. It can affect search, selection, copy and paste, accessibility, analytics, and future rendering. If a specific system surface cannot perform adequate segmentation, treat separator insertion as a narrowly owned transformation with native review, a reversible source, and tests. It should not become part of every translation by default.

Select a policy by content role

One line-break setting is not appropriate for every component. Input fields prioritize stable cursor behavior and fast updates. Long display paragraphs can spend more work finding readable breaks. Headings have short lines and may accept looser rules to avoid an isolated character.

Android's guide to paragraph line breaking in Jetpack Compose makes this distinction explicit. It presents simple, heading, and paragraph policies, and notes that strictness and word-break controls are designed for CJK text.

Use a small policy table in the design system:

Content role Default policy Reason
Editable field Simple and stable Avoid moving cursor positions while the user types
Body copy Paragraph quality Prefer readable multi-line composition
Short heading Heading policy Balance short lines without treating them as body text
Button or tab Product-specific Wrapping may damage navigation, so redesign before shrinking text
System notification Platform-owned The app may not control the final width or renderer

A Stack Overflow report about Thai line breaks on iOS and Android describes a push notification whose no-space Thai sentence split inside a word. This is practitioner evidence of a real boundary, not proof that every platform version fails. The operating system owns the notification layout, so the app cannot assume that an in-app text setting will carry over.

For a system-owned surface, verify the actual release build on supported operating systems. If the result remains unreadable, prefer shorter translator-approved notification copy with a link to the full in-app message. Do not manufacture a device-specific set of line breaks on the server.

Implement the policy at one rendering boundary

Centralize line-break selection instead of configuring individual screens ad hoc. A design-system text component can accept a semantic role and resolve the platform setting in one place.

A Compose implementation can start with a mapping like this:

import androidx.compose.ui.text.style.LineBreak

enum class LocalizedTextRole {
    Input,
    Heading,
    Paragraph
}

fun lineBreakFor(role: LocalizedTextRole): LineBreak = when (role) {
    LocalizedTextRole.Input -> LineBreak.Simple
    LocalizedTextRole.Heading -> LineBreak.Heading
    LocalizedTextRole.Paragraph -> LineBreak.Paragraph
}

Apply the result through the shared text style, then test it under Japanese and Thai locales. Keep the role semantic. A screen should ask for paragraph behavior rather than naming a low-level algorithm that may change across framework versions.

On platforms without equivalent presets, preserve the same contract: input, heading, and paragraph are separate roles; the locale reaches the text engine; native shaping and segmentation stay enabled; and custom measurement code does not split strings by spaces or arbitrary code-unit offsets.

Avoid pre-wrapping text on the server. The server does not know the final font, width, text scale, operating-system renderer, or adjacent icon layout. Return message content and its language. Let the client choose line breaks at layout time.

Cache the source message and language, not a version with inserted newlines for a particular width. When the language, font, text scale, or container width changes, layout should run again without fetching or rewriting the translation.

Build fixtures that expose the real failures

English pseudo-localization is useful for expansion, but it does not reproduce Japanese punctuation rules or Thai segmentation. Add real-script fixture strings approved by native reviewers. Each fixture should target one behavior rather than trying to represent an entire language.

A compact fixture record can include:

{
  "id": "thai-no-space-sentence",
  "locale": "th-TH",
  "role": "paragraph",
  "text": "ให้แตะปุ่มเมนูที่ด้านซ้ายบนของหน้าจอตัวละครหลัก",
  "widths_dp": [120, 180, 320],
  "text_scales": [1.0, 1.3, 2.0],
  "review": "native-reader-required"
}

Create separate Japanese fixtures for opening punctuation, closing punctuation, small kana, numerals, Latin abbreviations, and mixed Japanese and URL text. Include Chinese and Korean fixtures used by the product rather than assuming that one CJK sample covers all three writing systems.

Run every fixture through these dimensions:

  • The narrowest supported component width.
  • A normal phone width and a wide tablet width.
  • Default and large accessibility text scales.
  • Every bundled font and fallback chain used by that locale.
  • Light and dark themes if font weight changes.
  • In-app labels, dialogs, widgets, and system notifications where the same copy appears.
  • Current and oldest supported operating-system versions.

Screenshot tests can detect clipping and layout movement, but they cannot decide whether a Thai word was split correctly. Pair visual regression with native review for the expected break opportunities. Record the accepted screenshots by fixture, locale, platform version, font, and width so later framework upgrades can be compared deliberately.

Handle failures without corrupting the source

When a fixture fails, first identify which layer owns the break:

  1. Confirm that the content has the correct language tag and no stale manual newline.
  2. Confirm that the chosen component role resolves to the intended line-break policy.
  3. Confirm that the expected font supports the script and that fallback did not change metrics unexpectedly.
  4. Reproduce the result in a minimal native text component outside the product layout.
  5. Compare operating-system versions before adding application-level transformations.

If the minimal native component succeeds, the app's layout or style wrapper is probably overriding the correct behavior. If it fails only on one supported platform version, isolate the compatibility path and retain a regression fixture. If it fails on every surface, ask a native reviewer whether the expected boundary is actually required before changing content.

URLs, account identifiers, code, and mixed-script product names create legitimate awkward cases. Do not pass machine tokens through a natural-language dictionary segmenter and assume every proposed break is safe. Keep tokens atomic where the product requires exact copying, allow wrapping around them where the platform supports it, and test overflow separately.

Avoid the fixes that create later bugs

Common shortcuts move the problem instead of solving it:

  • Adding \n to each translation couples content to one width.
  • Splitting on spaces assumes a writing system that marks every word with spaces.
  • Breaking by code unit can split a user-perceived character and repeats the problem solved by grapheme-safe truncation.
  • Shrinking the font hides wrapping defects and harms accessibility.
  • Copying a Japanese policy to Thai ignores different segmentation needs.
  • Trusting one screenshot ignores text scale, fallback fonts, and system-owned surfaces.

A useful code review question is not, "Does this string fit?" Ask, "Which component chooses the break, which locale does it use, and which fixture proves the result?" That question exposes ownership before a translation ships.

Verify the release path

Make line-breaking checks part of the locale release gate. The test should launch the real app with the target locale, render the fixture catalog through production components, and capture results at required widths and text scales. Review failures against an approved baseline and require native signoff when break semantics change.

Re-run the set after updates to the operating system, UI framework, font files, Unicode data, or text component library. Those changes can alter legal break opportunities or line metrics without any translation-file change.

Start with the most constrained Japanese or Thai screen in the app. Remove width-specific newlines, route it through a semantic text role, and add three widths plus two text scales to the fixture catalog. Then test the same message on any notification or widget surface that renders it. The work is complete when the source remains width-independent and every supported renderer produces a native-reviewed break.

References

  1. Unicode Standard Annex #14: Unicode Line Breaking Algorithm defines line-break classes, opportunities, prohibitions, and contextual handling for Southeast Asian scripts.
  2. W3C: Requirements for Japanese Text Layout documents Japanese line-start and line-end restrictions, character classes, and strictness levels.
  3. Android Developers: Paragraph line breaking in Jetpack Compose documents simple, heading, and paragraph policies plus CJK strictness controls.
  4. Stack Overflow: Thai line breaks on iOS and Android provides practitioner evidence of a Thai notification breaking inside a word.