Why identical-looking text can be different

The same visible character can be encoded as one Unicode code point or as a base character followed by a combining mark. For example, é can be U+00E9 or the sequence U+0065 U+0301. A direct comparison sees different sequences even when the text looks identical.

café / cafe\u0301

NFC and NFKC

Unicode Standard Annex #15 defines normalization. NFC composes canonical equivalents where possible. NFKC additionally applies compatibility mappings, which can be useful for search but may change semantics. Treat normalization as a deliberate transformation, not a cosmetic fix, and record which form your system expects.

Try the normalization forms locally

NFC and NFKC fixture

Use the fixture to compare a decomposed é and a circled digit. Both are followed by U+FE0F VARIATION SELECTOR-16; the output lists that selector so you can verify it remains in the sequence.

The local result appears here.

Compare before transforming

Use Compare two strings exactly to inspect raw differences and whether the strings become equal after NFC or NFKC. The tool reports equality; it does not decide which representation is correct for your application.