Multilingual PDF generation fails in ways that a localized app screen does not reveal. A report can use the wrong language, render dates under the server locale, drop Korean glyphs, reverse Arabic text, or expose gibberish when a reader copies a line. These defects often appear after the file has left the app, where ordinary UI tests cannot catch them. The fix is to treat each PDF as a versioned localization artifact, not as a screenshot of translated strings. This guide defines one export contract, an implementation sequence, and release checks that keep templates, typed values, fonts, text direction, and language metadata aligned.
Why a localized screen can produce a broken PDF
App UI and document rendering usually follow separate code paths. A mobile client might resolve its effective locale through native resources while an export service loads an English HTML template. The screen might use a platform number formatter while the report interpolates a preformatted value returned by an API. The app font stack might fall back to system fonts, but the PDF library can use only the fonts explicitly embedded in the file.
That separation creates several common failure modes:
- Document headings and explanations live in a template outside the reviewed message catalog.
- The renderer guesses from a server process, request header, device region, or user profile instead of receiving one explicit content locale.
- Dates, numbers, currency, measurements, and names arrive as display strings produced under a different locale.
- The PDF library draws glyphs but lacks the shaping, bidirectional ordering, fallback, or Unicode mapping the script requires.
- Visible text looks correct, but the PDF does not declare its natural language or preserve searchable, extractable Unicode text.
The last two failures are easy to miss. A practitioner report about multilingual PDF rendering describes missing glyphs across several scripts and copied text that becomes gibberish. That report is one implementation incident, not proof that every PDF library behaves the same way. It shows why a visual spot check is not enough.
Define the export contract before choosing a library
A reliable export starts with an explicit request object. Do not let the PDF layer discover language state from global process settings. The caller should provide the content locale, time zone, template identity, template version, typed data, and any required rendering policy.
A practical contract looks like this:
type DocumentExportRequest = {
documentId: string
templateId: string
templateVersion: string
contentLocale: string
timeZone: string
currency?: string
data: Record<string, unknown>
generatedAt: string
}
type DocumentExportResult = {
documentId: string
contentLocale: string
templateVersion: string
rendererVersion: string
checksum: string
storageKey: string
}
contentLocale answers which language and regional conventions the document uses. It is not automatically the device locale, billing country, delivery country, or language of every user-provided field. timeZone is separate because a language tag does not identify a time zone. currency is explicit when business rules require a particular unit rather than a locale default.
Store the resolved values with the generated file. If support receives a bad export, the team must be able to reproduce the exact template, locale, and renderer combination. Without that record, regeneration can silently produce a different document from the one the user received.
Build one deterministic rendering sequence
The sequence below closes the gap between translated app resources and a durable document. Each step should finish before pagination because late text replacement can change line length, direction, and page count.
Resolve one content locale
Resolve the locale at the product boundary, then validate it against the templates and formatter data available to the renderer. Use a documented fallback chain, such as a specific language and region, then the base language, then the product default. Record both the requested locale and the locale actually used.
Do not hide fallback. If fr-CA resolves to fr, include that result in diagnostics. If it resolves to English because the French template is missing, decide whether the export may proceed. A receipt might require a usable fallback, while a regulated disclosure may require the requested version and should fail closed. That decision belongs in product policy, not in a generic translation helper.
Keep template copy in the localization workflow
Assign stable message IDs to headings, labels, footnotes, empty states, and error explanations. The document template should reference those IDs rather than contain freehand English. Reviewers need the document type, field meaning, space constraints, variables, and legal context alongside each message.
Version the template and its messages together. A translated string can be structurally valid yet describe an older calculation or policy. When source meaning changes, mark dependent translations stale and prevent the new template from claiming full locale coverage until review is complete.
User-provided text is different. Preserve it in its original language and script unless the product explicitly offers translation. Do not pass names, addresses, notes, or quoted messages through the UI catalog simply because they appear beside translated labels.
Format typed values at the render boundary
Pass timestamps, decimal values, currency codes, quantities, and identifiers as typed data. Format them only after the content locale and time zone are known. Unicode CLDR defines the locale patterns used for numbers and currencies, including decimal symbols, grouping, percent output, and currency forms.
Separate locale rules from business rules. The locale controls presentation, but it does not decide which currency an invoice uses or which time zone a transaction occurred in. A Canadian French report can contain US dollars, and an English report can show an event in Tokyo time. The renderer needs both pieces of information.
Keep machine identifiers invariant. Order IDs, hashes, SKU values, and protocol fields should not gain localized digits or grouping unless the product specification explicitly says they are reader-facing quantities. This distinction prevents a visually pleasant export from becoming impossible to reconcile with backend records.
Shape text before measuring and pagination
Right-to-left support is not a final alignment switch. Mixed Arabic, Hebrew, Latin, numbers, and punctuation require a text engine that applies bidirectional ordering and script shaping before the layout engine measures lines. Unicode Standard Annex #9 defines the bidirectional algorithm for text that combines left-to-right and right-to-left runs.
Use logical document structure. Paragraph direction should follow the paragraph content or an explicit template rule. Tables should define semantic column order rather than reverse an already positioned canvas. Isolate user-entered identifiers, phone numbers, and URLs when they appear inside right-to-left prose so punctuation does not jump between runs.
Measurement must use the shaped output and the final font choices. If the renderer measures unshaped characters, then substitutes contextual glyphs later, line breaks and page boundaries can change. Run shaping, fallback selection, line breaking, and pagination as one deterministic pipeline.
Embed fonts with Unicode mappings
There is no dependable single-font shortcut for every script and symbol. Build a deliberate fallback set based on the locales the product supports. Confirm that each font license permits embedding and that the renderer can preserve a Unicode mapping for copied and searched text.
Map coverage by script, not only by language name. A Japanese document may contain Latin product codes, emoji, and mathematical symbols. Arabic requires connected shaping. Devanagari uses combining behavior that a simple glyph lookup cannot reproduce. A font file containing code points does not prove that the rendering stack can shape them correctly.
Subset fonts only after the full text is known. A correct subset reduces file size while retaining every glyph used in the document. A bad subset can remove marks, ligatures, or fallback glyphs that appear only in customer data. Test dynamic fields, not just fixture headings.
Set language and accessibility metadata
A PDF should carry machine-readable language information in addition to visible translated text. W3C's PDF16 technique explains how a tagged PDF sets its default language through the document catalog so assistive technology and other user agents can choose suitable language behavior.
Set the document default to the resolved content locale. If substantial passages use another language, tag those passages when the chosen PDF tool supports it. Preserve headings, lists, tables, and reading order as structure rather than drawing everything as unrelated text fragments.
Language metadata does not fix incorrect text. It helps readers and software interpret correct text. Keep it in the same export contract so it cannot drift from the template locale.
Use a locale manifest instead of scattered checks
For each supported document locale, maintain a manifest that release tooling can inspect. It should include:
- resolved language tag and allowed fallback;
- template version and translation review state;
- required font files and expected script coverage;
- text direction and shaping capability;
- number, currency, date, and time-zone fixtures;
- document-language metadata value;
- golden test data for long text, mixed scripts, and page breaks;
- policy for missing translations or unsupported glyphs.
Generate this manifest from source-controlled configuration. Do not infer it from whichever files happen to exist in a build directory. CI can then fail when a new locale is enabled in the app but missing from the document pipeline.
A manifest also clarifies ownership. Localization owns reviewed template copy. Product and legal teams own the meaning and fallback policy. Engineering owns typed data, rendering, metadata, and deterministic verification. Font licensing and accessibility need named reviewers rather than assumptions buried in a PDF helper.
Handle failures without shipping a plausible wrong file
A PDF renderer should return typed failures before it stores or emails the artifact. Useful categories include missing template locale, stale translation state, unsupported script, missing glyph, shaping unavailable, formatter data unavailable, and metadata write failure.
Some failures can use a declared fallback. Others should stop generation. Define the decision per document type:
- A casual activity summary may fall back to the product default if the export states which locale was used.
- A customer invoice may allow language fallback but must preserve the transaction currency and tax data exactly.
- A consent record or regulated notice may require an approved locale-specific template and should not silently substitute another language.
- A user-generated message export should preserve original text even when the surrounding report falls back.
Never replace an unsupported character with a blank box and call the export successful. Record the code point, selected font chain, locale, and field identifier in internal diagnostics. Do not put sensitive field contents into logs.
Make retries idempotent. The same document ID, template version, locale, and data snapshot should produce the same logical artifact. If a renderer upgrade changes layout or font output, assign a new renderer version so support can distinguish regeneration from the original file.
Verify text, structure, and appearance
A complete test matrix checks more than screenshots. Use deterministic fixtures for every launch locale and include representative user data from each supported script.
Run these checks in CI:
- Assert the requested and resolved locales match the manifest policy.
- Confirm every template message exists and has the approved review state.
- Scan rendered text for missing-glyph markers and unexpected fallback language.
- Extract text from the PDF and compare key values with the expected Unicode strings.
- Search the PDF for names, identifiers, and translated headings.
- Inspect document language metadata and structural tags.
- Verify page count, overflow, clipped content, and reading order against stable expectations.
- Render selected pages to images for visual regression checks across LTR, RTL, CJK, and mixed-script fixtures.
- Open the file with at least two independent PDF readers when the document is customer-critical.
- Reproduce the file from its stored locale, template version, renderer version, and data snapshot.
Text extraction catches defects that screenshots miss. Visual comparison catches layout defects that extraction misses. Metadata inspection catches accessibility defects that both can miss. Keep all three layers.
Use native-language review for the final release fixtures. Automated checks can prove that a string exists and survives extraction, but they cannot confirm that a line break is readable, a translated label fits its business meaning, or an RTL table follows reader expectations.
Common implementation mistakes
Do not generate a PDF by taking localized HTML and assuming the browser result will transfer unchanged. The conversion engine may support a different font set, shaping engine, CSS subset, or accessibility model.
Do not use the server's default locale. Container images and worker processes often default to a locale unrelated to the user, and that state can change across environments.
Do not format values before enqueueing an export job. A queued job should carry typed values plus locale and time-zone context, not presentation strings that hide where formatting happened.
Do not validate fonts by file name. Test actual code points, shaping, extraction, and fallback order. A font branded as global can still omit scripts or symbols your users provide.
Do not treat a rendered page image as document accessibility. A PDF can look correct while lacking searchable text, reading order, or language metadata.
Do not reuse rank or locale as the only artifact identity. Store document ID, data snapshot, content locale, template version, and renderer version so the exact output can be traced.
The next action
Choose one customer-facing export and write its contract before changing the renderer. Record the content locale, time zone, template version, typed inputs, font fallback set, PDF language metadata, and fallback policy. Then add four fixtures: long German text, mixed Arabic and Latin content, CJK customer data, and a value set with locale-sensitive currency and dates. Require text extraction, metadata inspection, and rendered-page comparison to pass before that export joins the normal localization release gate.
References
- W3C PDF16 supports setting a tagged PDF's default language for assistive technology and other user agents.
- Unicode Standard Annex #9 defines bidirectional ordering for mixed left-to-right and right-to-left text.
- Unicode CLDR number and currency formatting supports locale-aware decimal, percent, compact-number, and currency presentation.
- Stack Overflow: multilingual PDF rendering question provides practitioner evidence of missing glyphs and broken copied text in one generated-PDF workflow.
