Normalize caracteres confundíveis em texto seguro
Normalizar confundíveis significa substituir cada imitação pelo caractere comum que ela imita, para que duas strings visualmente iguais também sejam iguais na comparação. Faça isso antes de comparar nomes de usuário, aplicar listas de bloqueio ou remover identificadores duplicados.
Exemplo resolvido
- Entrada
- Cоnfig file: аdmin2, Noёl, café
- Sistemas de escrita encontrados
- latim (Latn), cirílico (Cyrl)
- Caracteres sinalizados
- 7
| Posição | Caractere | Ponto de código | Sistema de escrita | Substituição | Regra |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | latim | C | Normalização NFKC |
| 1 | о | U+043E | cirílico | o | Mapa de confundíveis |
| 7 | fi | U+FB01 | latim | fi | Normalização NFKC |
| 10 | : | U+FF1A | comum | : | Mapa de confundíveis |
| 12 | а | U+0430 | cirílico | a | Mapa de confundíveis |
| 17 | 2 | U+FF12 | comum | 2 | Mapa de confundíveis |
| 22 | ё | U+0451 | cirílico | ë | Mapa de confundíveis |
- Preservar Unicode legível
- Config file: admin2, Noël, café
- Fallback ASCII estrito
- Config file: admin2, Noel, café
Como funciona
- Imitações conhecidas são substituídas primeiro pelo mapa embutido. Caracteres fora do mapa são normalizados com NFKC, que converte formas de largura total, ligaduras e outros caracteres de compatibilidade em seus equivalentes padrão.
- Preservar Unicode legível mantém uma letra acentuada quando o mapa define uma, como o ё cirílico → ë. Fallback ASCII estrito usa a letra ASCII simples (ё → e).
- Letras que não estão no mapa e não mudam com a NFKC são mantidas, então café conserva o acento nos dois modos. Texto normalizado é uma chave de comparação, não um veredito de segurança: guarde também o original e revise os caracteres sinalizados.
A conversão é o melhor esforço: confusão mapeada e dobramento NFKC são determinísticas, mas alguns Unicode legítimos não serão sinalizados.
Seu texto
Colar ou digitar – os resultados são atualizados conforme você digita (levemente rebatidos para entradas longas).
30 caracteres analisados
7 suspeitos
Fallback ASCII estrito
Original (caracteres suspeitos marcados)
Caracteres suspeitos na visualização original são sublinhados e rotulados como “susp.” além de realçar a cor.
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
Saída limpa
Análise de caracteres
| Índice (baseado em 0) | Original | Substituição | Ponto de código | Razão |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |