Skip to main content

Build Accent-Insensitive Search for Multilingual Apps

2026-09-28

Build Accent-Insensitive Search for Multilingual Apps

Accent insensitive search fails when a user types Jose and cannot find José, or types a visually identical name that uses a different Unicode representation. The usual patch is to lowercase text and strip every combining mark. That makes one demo pass, but it can damage meaningful distinctions, produce false matches, and leave database results inconsistent with client-side filtering.

Build a search-equivalence contract instead. Keep original display text, normalize comparable forms consistently, choose accent folding by field and locale, rank stronger matches first, and preserve a mapping for highlights. The same policy must control indexing, querying, ranking, and result rendering.

Why visually identical text can fail to match

Unicode can represent some visible text in more than one valid sequence. For example, an accented character may be stored as one precomposed code point or as a base letter followed by a combining mark. The strings can look identical while a byte comparison or basic substring test treats them as different.

Unicode Standard Annex #15 defines normalization forms and canonical equivalence. Normalization gives an application a stable representation for comparison. It does not mean that every accented letter should equal its unaccented base letter.

Each requirement expands equality differently:

  1. Canonical matching: é should match the equivalent sequence e plus a combining acute accent.
  2. Accent-insensitive matching: e may match é when the product deliberately treats that distinction as optional.
  3. Locale-sensitive matching: case and accent behavior may depend on the language associated with the content or user.
  4. Transliteration: text in one script may match a representation in another script. This requires a separate policy and is outside simple accent folding.

Confusing those layers causes unpredictable search. NFC normalization can solve canonical representation differences while retaining accents. Removing marks changes the search alphabet. Transliteration changes it further. Each step expands the result set and should have its own ranking weight.

A practitioner asking about diacritic-insensitive SQLite search on iOS gives a concrete example. An unaccented query such as thu is expected to find Vietnamese names with several marked forms, but ordinary SQL LIKE behavior does not provide that equivalence. Treat this as a reported implementation need, not as a universal rule for Vietnamese search.

Define the search-equivalence contract first

Write the behavior for each searchable field before selecting an API or tokenizer. A contact name, product code, city name, and full-text article do not need the same matching rules.

For every field, decide:

  • Which locale controls case and language-sensitive comparison.
  • Whether canonically equivalent sequences must match.
  • Whether accents are primary distinctions, ranking signals, or ignored.
  • Whether punctuation, spaces, width variants, and digits are significant.
  • Whether the query can cross scripts through transliteration.
  • Whether matching happens on a device, a server, or both.
  • How a result maps back to the original text for highlighting.

Start with the least destructive policy. Canonical normalization is a safe baseline when all systems apply the same form. Accent folding should be enabled only for fields where users benefit from forgiving input and where collisions are acceptable.

For a people directory, an unaccented query may reasonably find accented names, but an exact spelling should rank first. For a username, identifier, medication code, or legal record, removing distinctions can return the wrong entity. Search convenience does not override domain accuracy.

Do not infer the content locale from the device locale. A French contact name inside an English interface still belongs to the content, not the interface chrome. Store a language or locale with content when matching rules depend on it. If that metadata is unavailable, use a documented neutral policy rather than silently changing behavior by device.

Build one normalization pipeline

Use the same transformation code for indexing and querying. If the server folds text differently from the mobile client, the app can display a local result that disappears after synchronization or highlight a substring the server did not actually match.

Store each representation separately:

  • display_text: the original string used for rendering.
  • canonical_text: a consistent Unicode normalization form, commonly NFC.
  • folded_text: an optional search form created under a named field policy.
  • search_version: a version for the transformation and tokenizer rules.
  • offset_map: a mapping from folded search units back to display-text ranges when highlighting is required.

The following JavaScript-style pseudocode separates canonical normalization from optional folding. It is an implementation example, not a universal Unicode search algorithm.

function buildSearchForms(input, policy) {
  const canonical = input.normalize("NFC");

  if (!policy.ignoreAccents) {
    return {
      display: input,
      canonical,
      folded: canonical.toLocaleLowerCase(policy.locale)
    };
  }

  const decomposed = canonical.normalize("NFD");
  const withoutMarks = decomposed.replace(/\p{M}/gu, "");

  return {
    display: input,
    canonical,
    folded: withoutMarks.toLocaleLowerCase(policy.locale)
  };
}

The mark-removal branch is intentionally conditional. Combining marks are not decorative noise in every script or language. A product that enables this branch should do so for specific fields and test languages, rather than treating it as a global cleanup step.

Also version the policy. Unicode data, database tokenizers, platform libraries, and product decisions can change. A stored folded column created under version 1 must be rebuilt when version 2 changes normalization, case handling, or tokenizer settings. Without a version, stale and fresh records can produce different matches for the same query.

Choose a matching engine without splitting the policy

Different search engines expose different controls, but the application contract should remain stable.

Use collation-aware search when language rules matter

The ICU String Search guide describes language-sensitive searching driven by a locale or collator. Its comparison behavior can consider collation attributes, and normalization is handled within collation behavior. This is useful when matching needs language-aware equivalence rather than a hand-built sequence of lowercase and replacement operations.

A collator still needs explicit settings and fixtures. Do not assume that selecting a locale automatically matches the product's desired accent, case, punctuation, and substring behavior. Record the chosen strength and attributes as part of the search version.

Configure local database tokenization deliberately

For local full-text search, SQLite FTS5 can tokenize text through its unicode61 tokenizer. The SQLite FTS5 documentation documents remove_diacritics options and states that this removal applies to Latin-script characters. That is a specific engine behavior, not a complete multilingual equivalence system.

If the app uses FTS5, create the index with an explicit tokenizer configuration instead of accepting an undocumented default. Keep a migration that can rebuild the virtual table when the policy changes. Test the generated query against the installed SQLite versions you support because the local database is part of the release artifact.

Do not run FTS on the device and then apply a different accent-folded filter in memory. That can drop valid database results, add results the index cannot retrieve, and corrupt pagination counts. Use in-memory checks only as assertions during development or as part of a clearly defined second-stage ranker.

Keep server and device behavior compatible

A server-backed app should send the query and the effective search policy identifier, not a client-generated SQL fragment or arbitrary locale object. The server owns index lookup and ranking. The response should include stable record IDs, original display text, and either trusted match ranges or enough information for the client to reproduce highlighting safely.

When offline search is required, share fixture data and expected match tiers across implementations. Identical libraries are helpful but not mandatory. Compatible behavior is mandatory.

Rank exact matches above folded matches

Accent-insensitive matching should increase recall without making weak matches look exact. Use tiers instead of placing every equivalent form into one bucket.

Use this default ranking order:

  1. Exact original-text match.
  2. Case-insensitive match with accents preserved.
  3. Canonically equivalent match after normalization.
  4. Accent-folded match under the field policy.
  5. Optional fuzzy or transliterated match.

Within a tier, use product-specific signals such as prefix match, token coverage, recency, or popularity. Always add a stable record ID as the final tie-breaker so result order does not change between runs. The same tie-breaker discipline governs list order generally, covered in building locale-aware sorting for multilingual apps.

Log the tier that admitted each result. When a user reports an irrelevant match, the team can see whether accent folding, fuzzy matching, or transliteration caused it. Without that trace, tuning becomes guesswork.

Preserve correct highlighting

A folded string does not have a reliable one-to-one character relationship with the original. Normalization can combine or decompose code points, and removing marks changes offsets. Applying a folded-string index directly to the display string can highlight half a grapheme or the wrong range.

Build the offset map while producing the searchable form. Process display text in user-perceived character clusters, append each cluster's transformed search form, and record which display range produced each output range. When the engine returns a match, translate that range through the map before rendering.

For token-based server search, returning whole matched token IDs can be safer than returning raw code-unit offsets across platforms. The client can then find and highlight the corresponding original token using the same versioned segmentation fixtures.

If accurate ranges are unavailable, highlight the full field or omit highlighting. A correct unhighlighted result is better than marking unrelated letters and implying a match the user cannot understand.

Test collisions and missed matches

One successful Jose fixture is not enough. Build a matrix that proves both expected matches and expected non-matches.

Include these categories:

  • NFC and NFD representations of the same visible word.
  • Accented and unaccented Latin names.
  • Turkish dotted and dotless forms under Turkish and neutral policies.
  • Vietnamese words with several marks.
  • Arabic text with and without optional marks, only if the product has approved that behavior.
  • Mixed-script content that must not become equivalent by accident.
  • Exact identifiers that must retain every mark and character.
  • Prefix, substring, multi-token, and empty queries.
  • Duplicate folded forms with stable result ordering.
  • Highlights that cover complete displayed characters.

For each fixture, assert the returned IDs, match tier, order, and display range. Run the same fixture set against the server index, local database, and any fallback in-memory implementation.

Also test migrations. Populate an index under the previous search version, upgrade the app, interrupt the rebuild, and restart offline. The app should either keep using the complete old index or activate the new one only after a successful rebuild. A partially rebuilt multilingual index is a release failure, not a minor ranking issue.

Common implementation mistakes

  • Stripping all marks globally expands equality without considering language or field accuracy. Scope folding to an explicit policy.

  • Normalizing only the query leaves representation-dependent misses. Stored and incoming text must pass through compatible transformations.

  • Using the UI locale for every field ignores the difference between content and interface language. Resolve matching from field ownership and content metadata.

  • Letting each layer improvise creates mismatched results. Database tokenization, server ranking, client filtering, and highlighting must share a contract and version.

  • Returning folded text to the UI exposes internal index data. Always render the original reviewed content.

  • Changing rules without rebuilding indexes mixes old and new behavior. Persist the search version and make reindexing an atomic migration with rollback behavior.

Verify one field end to end

Choose one high-value field, such as contact names or product titles. Write its equivalence rules, generate canonical and folded forms, configure one index, rank exact matches above folded ones, and preserve highlight ranges. Then run the shared fixture matrix on every execution path.

Ship only after logs can identify the policy version and match tier for a result. Once that field behaves consistently online, offline, and after an index migration, extend the same contract to the next field instead of adding another ad hoc accent-removal helper.

References