Skip to main content

Prevent locale-sensitive case conversion bugs in apps

2026-09-30

Prevent locale-sensitive case conversion bugs in apps

Locale sensitive case conversion breaks apps when one lowercase call handles unrelated data. A profile heading should follow the reader's language. An HTTP header name, cache key, database identifier, or file extension usually needs a stable rule that survives locale changes. When those paths share a helper, Turkish users can see misspelled words or failed lookups that never appeared in English testing. Fix the ambiguity at the call site: classify the value, choose an explicit casing or comparison policy, and test every boundary that stores, transmits, indexes, or reuses the result.

Why ordinary lowercasing is not one operation

Unicode casing is not a simple character substitution. The Unicode ICU case mapping guide distinguishes generic single-character mappings from full, language-specific mappings. Full mappings matter because a character can map differently under a particular language, and a conversion can produce more than one code point.

Turkish provides the familiar example. English pairs I with i. Turkish has two pairs: dotted İ with i, and dotless I with ı. A user-facing uppercase or lowercase conversion that ignores that distinction can misspell a name or heading. MDN demonstrates the difference directly in its toLocaleLowerCase examples: lowercasing İstanbul under English and Turkish does not produce the same sequence.

Problems start when code inherits a device or host locale without saying why. A key created under English may not equal the version reconstructed after the app switches to Turkish. A background worker can use a server default that differs from the mobile client. Data saved before a locale change can then become unreachable, even though each component called a documented API.

A Java practitioner question about choosing a case-conversion locale captures the confusion: should a developer choose English, the user's locale, or something else? The answer depends on the value's role. There is no single correct locale for every string.

Classify the value before changing its case

Assign every conversion one of four jobs. Put the choice in a shared utility or design note so call sites cannot guess.

  1. Linguistic presentation covers visible titles, labels, names, and sentence fragments. Use the content locale or effective app locale, then confirm that automatic casing suits the language and product.
  2. Machine normalization covers protocol tokens, enum names, cache keys, route segments, internal identifiers, and file extensions. Follow the specification that owns the value. An ASCII case-insensitive format needs an explicit ASCII or invariant operation, never a user locale.
  3. Caseless comparison covers login matching, deduplication, filtering, and some search behavior. Lowercasing both sides does not define equality by itself. Choose the platform's case folding or collation behavior for the domain.
  4. Search presentation combines matching with highlighting and result display. Keep the original text, build a separate normalized index, and preserve a mapping back to displayed ranges.

The classification blocks two opposite mistakes. Applying the user's locale to every string can destabilize machine values. Applying invariant casing everywhere can damage visible language.

Names deserve special care. An app should not uppercase a person's entered name merely to create a visual style unless product and linguistic review support it. CSS or a platform text style may still use casing rules that differ by environment, so test the rendered result. Preserve the original value even when a secondary display form is needed.

Build separate APIs for visible text and machine keys

Make the intent visible in code. A generic helper named lowercase() encourages accidental reuse. Prefer interfaces whose names expose the contract.

caseForDisplay(text, contentLocale):
    require contentLocale is explicit
    return languageAwareCaseMap(text, contentLocale)

normalizeProtocolKey(value, protocol):
    require protocol defines the accepted alphabet and comparison rule
    return protocolInvariantNormalize(value)

compareUserText(left, right, locale, sensitivity):
    return localeCollator(locale, sensitivity).equals(left, right)

indexSearchText(original, locale):
    normalized = searchNormalizer(original, locale)
    return { original, normalized, offsetMap }

These interfaces force each caller to provide the inputs its policy needs. The display function cannot silently read the operating system's default locale. The protocol function does not accept an arbitrary language tag. The search function keeps the original text and the range mapping needed for highlighting.

On Java or Kotlin, pass a deliberate locale to language-sensitive casing and use the runtime's invariant option only for values whose contract is not linguistic. On JavaScript, pass the content locale to toLocaleLowerCase or toLocaleUpperCase for presentation. If a protocol defines ASCII matching, implement that specification instead of treating a browser locale API as a protocol normalizer.

English is not a substitute for an invariant contract. English casing happens to produce the desired ASCII output for many identifiers, but the owning format sets the rule. Record whether a key permits only ASCII, accepts Unicode, uses Unicode case folding, or remains case-sensitive. Keep that decision in the protocol or storage adapter.

Keep casing out of persistent identity

Do not make a derived lowercase string the sole identity for user data. Preserve an immutable ID and the original display value. Store any normalized lookup field separately and attach a normalization-policy version.

Versioning matters because Unicode data and application rules can change. A database row created with one rule may not match a query normalized under another. A version lets a migration rebuild indexes deliberately instead of making old records disappear from search.

Apply the policy to caches too. Include its version and any locale that legitimately affects the derived value. A display-title cache keyed only by source text can return a Turkish result to an English screen. Leave the user's locale out of a machine-key cache, where it would create duplicate entries for one locale-independent identity.

Never recalculate a persistent machine key from translated display copy. Translators can change capitalization, punctuation, and wording. Routes, analytics dimensions, feature flags, and API parameters should use stable identifiers, while localized labels remain presentation data.

An OpenAI Java SDK issue reports a Turkish-locale header and key lookup failure. This is a practitioner report, not a platform guarantee. It points to a useful test boundary: every component handling a case-insensitive infrastructure value must agree on comparison and normalization.

Treat comparison and search as separate designs

Lowercasing both operands is easy to inspect, but it does not define an international comparison policy. The domain may need exact code-point equality, canonical equivalence, locale-aware collation, Unicode caseless matching, or its own search equivalence.

Account identifiers often need a server-owned policy that is stable across clients. User-facing contact search may need locale-aware matching. Product codes may be restricted to ASCII and compared under a published protocol. These cases should not share one helper merely because all three are described as case-insensitive.

For search, keep normalization symmetric. Index and query text must use the same policy and version. Preserve original text for display and maintain offsets if normalization changes code-point count. Without an offset map, highlighting can select the wrong characters after a full case mapping or combining sequence appears.

This article's scope ends before accent removal, tokenization, and ranking, but the boundary is deliberate. Accent removal is covered separately in building accent-insensitive search for multilingual apps. Casing policy feeds those layers. Define it first, then let the search pipeline decide whether diacritics, punctuation, or word boundaries are equivalent for the product's users.

Handle locale changes without corrupting state

When the app language changes, invalidate derived presentation and leave source data alone. Re-render visible casing with the new effective locale. Rebuild a locale-specific search index only when its contract uses that locale. Invariant identifiers and protocol keys must stay unchanged.

Track the source of the effective locale. A content item may declare Turkish while the interface is English. In that case, casing the content title with the interface locale may be wrong. Pass the content locale when the text has one. Use the app locale for interface-owned copy. Use no linguistic locale for a machine token governed by a protocol.

Background jobs need explicit inputs too. Push preparation, export generation, and server-side rendering should receive the locale associated with the text. They should not inherit the host's environment. If a locale is missing, apply a documented fallback and log the missing context. Do not silently pick whatever locale happens to be installed on that machine.

Offline caches require migration rules. When a user changes language, clear or version locale-derived presentation caches. Do not delete canonical data. If a search index depends on locale, mark it stale and rebuild it atomically so queries do not mix old and new normalization policies.

Test the boundaries that English fixtures miss

One Turkish word is not a test suite. Build fixtures around each contract, and assert the intermediate values along with the final screen output.

For linguistic casing, include Turkish dotted and dotless I, Lithuanian text with combining marks, Greek words where context affects the rendered lowercase form, and strings whose full mapping changes length. Include decomposed and composed input where the product accepts both. Verify that the content locale, not the test machine locale, controls the result.

For machine values, run the same header, enum, route, cache key, and file-extension cases under English, Turkish, and another non-English default locale. The outputs must remain identical when the governing protocol is locale-independent. Restart the app under each locale rather than changing only a mock formatter.

For persistence, create data under one locale, change the locale, restart, and retrieve the same record. Repeat the test across an app upgrade that changes the normalization-policy version. Confirm that migrations rebuild derived fields without changing immutable IDs or original display text.

For search and comparison, test both matching and highlighted ranges. Include values that expand during case mapping, canonically equivalent strings, mixed scripts, and a query entered under a different keyboard locale. Assert that index creation and query normalization use the same version.

Also test cross-platform exchanges. Create a key on iOS, send it through an API, cache it on the server, and read it on Android. A unit test inside one runtime will not catch a disagreement between language libraries or Unicode-data versions.

Add a casing review gate

Audit every lowercase, uppercase, title-case, case-insensitive comparator, and normalized-key call. Record the value type, owner, allowed alphabet, locale source, persistence lifetime, and comparison rule. A call site without those answers is unfinished.

Move conversions behind intent-specific APIs. Add Turkish-locale CI runs for machine boundaries and explicit-locale fixtures for visible text. Keep original data beside derived forms, version persistent normalization, and invalidate only caches that depend on locale.

Start with the riskiest shared helper. Trace every caller and split it into display casing, machine normalization, comparison, or search indexing. The work is complete when changing the app or host locale can alter reader-facing casing where expected without changing a stored identifier, breaking a lookup, or corrupting a highlighted result.

References