Skip to main content

Count and truncate user text by grapheme cluster

2026-09-30

Count and truncate user text by grapheme cluster

A grapheme cluster character count prevents a common multilingual input bug: the counter says a visible name uses several characters, or truncation leaves a broken accent or half an emoji. This happens when the interface promises a limit in characters but the code enforces UTF-16 units, Unicode scalars, code points, or bytes. Define the unit first. Count extended grapheme clusters for editing, enforce storage bytes separately, and run the same fixtures on iOS, Android, web, and the server. The counter then matches what the user sees without hiding real backend and protocol limits.

Why one visible character has several lengths

Unicode text has several valid measurement layers. They are not interchangeable:

  1. A byte count measures an encoded payload, such as UTF-8 stored in a database or sent through an API.
  2. A code-unit count measures a runtime representation. JavaScript and many Android string operations expose UTF-16 indices.
  3. A code-point or scalar count measures assigned Unicode values.
  4. An extended grapheme cluster approximates one character as a user perceives it during editing.

Unicode Standard Annex #29 defines the default segmentation rules and describes grapheme clusters as user-perceived characters. The distinction matters because one displayed unit can contain several scalars. A decomposed e plus a combining acute accent displays like precomposed é. A skin-tone modifier combines with an emoji. Flags use pairs of regional indicators, and family emoji use joiners to form one displayed sequence.

The same visible input can therefore have different byte, code-point, and code-unit lengths. Here are measurements from one current JavaScript runtime using Intl.Segmenter:

Input Grapheme clusters Code points UTF-16 units UTF-8 bytes
é 1 1 1 2
e plus combining acute accent 1 2 2 3
Woman plus medium skin tone 1 2 4 8
Joined family emoji 1 7 11 25
Japan flag 1 2 4 8

A rule called maxLength: 20 is incomplete. It could mean twenty grapheme clusters in the editor, twenty code points in a protocol, twenty UTF-16 units in a legacy API, or twenty bytes in storage. Each interpretation rejects and truncates a different set of strings.

The long-running practitioner question about getting the length of a Swift String reflects this ambiguity. A method name does not settle which unit it returns or whether that unit matches the product promise.

Write the limit contract before the helper

Define the product boundary before choosing a string utility. For each constrained field, record:

  • The user-visible unit and limit.
  • The transport encoding and payload limit.
  • The storage encoding and indexed-column limit.
  • Whether the value may be truncated or must be rejected.
  • Which component owns final validation.
  • The Unicode or runtime versions covered by compatibility tests.

A profile display name might allow 40 grapheme clusters in the interface and 160 UTF-8 bytes at the API. A protocol identifier might instead define an ASCII byte limit, making grapheme counting irrelevant to acceptance. A push provider can impose a payload limit for the whole message rather than a character limit on one field.

Represent those decisions explicitly:

{
  "field": "display_name",
  "display_limit": {
    "unit": "extended_grapheme_cluster",
    "maximum": 40
  },
  "transport_limit": {
    "encoding": "UTF-8",
    "maximum_bytes": 160
  },
  "overflow_policy": "reject",
  "validator_version": "unicode-segmentation-v1"
}

Do not expose the transport maximum as if it were the user's character allowance. The counter can show 31 of 40 while validation separately checks the encoded payload. If a value fits the visible limit but exceeds the byte limit, return a specific error instead of silently removing the final cluster.

Choose truncation only for disposable previews, generated excerpts, or display-only labels where losing the suffix is acceptable. Reject overflow for names, messages, credentials, addresses, legal text, or any field where deleting user input changes meaning.

Implement grapheme-safe counting on every client

Use a platform segmentation API rather than maintaining an emoji table or splitting around combining marks yourself. Unicode segmentation rules cover interactions that a short regular expression will miss, and the rules evolve with Unicode.

Swift uses Character boundaries

The Swift book explains that every Swift Character represents one extended grapheme cluster. That makes ordinary String iteration appropriate for a user-visible count.

struct TextLimitResult {
    let graphemeCount: Int
    let value: String
    let exceeded: Bool
}

func applyDisplayLimit(_ text: String, maximum: Int) -> TextLimitResult {
    precondition(maximum >= 0)

    let count = text.count
    return TextLimitResult(
        graphemeCount: count,
        value: String(text.prefix(maximum)),
        exceeded: count > maximum
    )
}

Keep the original text when the policy is rejection. The prefix is useful for a preview or for showing what would fit, but it should not overwrite a submitted value without an explicit product decision.

Do not convert a Swift String to NSString and use its length for the visible counter. That exposes a UTF-16-oriented measurement rather than Swift Character boundaries. UTF-16 offsets are still valid where an Apple API specifically requires NSRange; they simply answer a different question.

Android uses a character BreakIterator

Android's BreakIterator.getCharacterInstance locates logical character boundaries. Its documentation calls out characters represented by multiple code points and explains that text editors should operate on the units users think of as characters.

import android.icu.text.BreakIterator
import android.icu.util.ULocale

data class TextLimitResult(
    val graphemeCount: Int,
    val preview: String,
    val exceeded: Boolean
)

fun applyDisplayLimit(text: String, maximum: Int, localeTag: String): TextLimitResult {
    require(maximum >= 0)

    val iterator = BreakIterator.getCharacterInstance(ULocale.forLanguageTag(localeTag))
    iterator.setText(text)

    var count = 0
    var previewEnd = iterator.first()
    var boundary = iterator.next()

    while (boundary != BreakIterator.DONE) {
        count += 1
        if (count <= maximum) previewEnd = boundary
        boundary = iterator.next()
    }

    return TextLimitResult(
        graphemeCount = count,
        preview = text.substring(0, previewEnd),
        exceeded = count > maximum
    )
}

The returned boundaries are valid indices for the string supplied to that iterator. Do not increment the index manually after obtaining a boundary. Recreate or reset the iterator when the text or locale changes, and keep input-method composition separate from committed validation.

Web clients use Intl.Segmenter

MDN documents Intl.Segmenter for locale-sensitive segmentation into graphemes, words, or sentences. With grapheme granularity, the segments can drive both a counter and a safe preview.

function applyDisplayLimit(text, maximum, locale) {
  if (!Number.isInteger(maximum) || maximum < 0) {
    throw new TypeError("maximum must be a non-negative integer");
  }

  const segmenter = new Intl.Segmenter(locale, { granularity: "grapheme" });
  const clusters = Array.from(segmenter.segment(text), part => part.segment);

  return {
    graphemeCount: clusters.length,
    preview: clusters.slice(0, maximum).join(""),
    exceeded: clusters.length > maximum,
  };
}

Check supported browser and embedded-webview versions before shipping. If the product must support a runtime without Intl.Segmenter, select a tested Unicode segmentation library and run it through the same fixtures. Do not fall back to text.length, spread syntax, or a homegrown emoji expression while continuing to label the result as a character count.

Keep the server authoritative without changing the unit

A client-side counter gives feedback. The server still enforces the contract because a client can be outdated, bypassed, or built with another Unicode data version.

The server should return structured diagnostics:

validateDisplayName(value):
    graphemes = segmentIntoExtendedGraphemeClusters(value)
    utf8Bytes = encodeUtf8(value).byteLength

    if graphemes.count > 40:
        return error("display_limit_exceeded", maximum=40, actual=graphemes.count)

    if utf8Bytes > 160:
        return error("storage_limit_exceeded", maximumBytes=160, actualBytes=utf8Bytes)

    return accepted(value)

Do not normalize or rewrite user text merely to make it pass a limit. Normalization can be a valid separate policy for identifiers, comparison, or indexing, but it must have its own owner and migration plan. Preserve the submitted display value unless the product contract says otherwise. Measure the actual representation that will be stored or transmitted for byte limits.

When runtimes disagree, log the client count, server count, app version, platform, and validator version without logging sensitive text. Reject the value when the authoritative server count exceeds the limit. Raise a compatibility alert instead of clipping a cluster and hiding the mismatch.

Handle editing and truncation failures deliberately

Text input has intermediate states. An input method can hold marked or composing text before committing the final sequence. Avoid destructive truncation during composition. Update the counter if the platform exposes a stable preview, but apply final acceptance after composition completes. Otherwise the app can delete a mark while the user is still building a character.

Pasted input needs the same path as typed input. Do not validate keyboard events and then trust paste, autofill, drag and drop, dictation, or restored drafts. Run the complete string through one segmentation and byte-validation function whenever committed text changes.

If a display preview must fit a cluster limit, segment first and join complete clusters. Never cut bytes and attempt to decode the result. Never take an arbitrary UTF-16 prefix. If an additional byte ceiling still fails after a grapheme-safe prefix, remove one complete trailing cluster at a time or reject, according to the field policy.

Rendering width remains a different constraint. Forty grapheme clusters can occupy very different widths across scripts, fonts, devices, and accessibility sizes. A grapheme limit prevents broken text boundaries; it does not replace flexible layout, wrapping, text expansion testing, or visual clipping checks. Those layout checks are covered in handling app translation text expansion and UI overflow.

Verify the full client-to-storage lifecycle

Create one shared fixture catalog and run it against every implementation. Include:

  1. Plain ASCII and ordinary whitespace.
  2. Precomposed and decomposed forms of the same accented letter.
  3. Emoji with skin-tone modifiers.
  4. Flags made from regional indicators.
  5. Joined family and profession emoji.
  6. Keycaps, variation selectors, and combining marks.
  7. Indic and other complex-script sequences used by supported locales.
  8. Empty input, line breaks, pasted text, and active composition.
  9. Values exactly at and one cluster beyond the display limit.
  10. Values within the display limit but beyond the UTF-8 byte limit.

Assert more than the final count. Verify that every preview ends on a segmenter boundary, rejoining all segments reproduces the original string, client and server return the same acceptance decision, and database round trips preserve the submitted value. Run the fixtures after operating-system, browser, language-runtime, or Unicode-library upgrades.

Add one end-to-end release test. Enter the same fixtures on iOS, Android, and web; submit them through the production-shaped API; retrieve them; and compare the rendered result. Include an older supported client against the current server. That is where a validator-version mismatch becomes visible.

Start with the field most likely to contain names, emoji, or multilingual free text. Write down its display and byte limits, replace the current length helper with platform segmentation, and add the shared fixtures before changing another field. The job is complete when every layer agrees on acceptance and no truncation path can split a user-perceived character.

References

  1. Unicode Standard Annex #29: Unicode Text Segmentation defines default extended grapheme cluster boundaries and the user-perceived character model.
  2. The Swift Programming Language: Strings and Characters documents Swift Character as an extended grapheme cluster.
  3. Android BreakIterator.getCharacterInstance documents logical character boundaries for text editing and multi-code-point characters.
  4. MDN Intl.Segmenter documents locale-sensitive grapheme segmentation for JavaScript.
  5. Stack Overflow: Get the length of a Swift String provides practitioner evidence of recurring ambiguity around string length and counting units.