Skip to main content
When generating spelling suggestions, the library needs to know which characters sound alike, look alike, or differ only by tone. These tables power the phonetic hasher, visual confusion detection, and tonal variant generation used throughout the suggestion pipeline.

Overview

Phonetic Groups

Characters grouped by phonetic similarity (same sound category):

Consonant Groups

Nasal Groups

Approximants and Liquids

Medials

Vowels

Tone

Visual Similarity

Characters that look similar and are commonly confused. Accessed via VISUAL_SIMILAR dict. Most pairs are mapped bidirectionally.

Vowel and Medial Confusions

Consonant Confusions

Aspirated vs Unaspirated Pairs

Other Confusions

Tonal Groups

Characters that differ by tone, commonly confused in typing. Accessed via TONAL_GROUPS dict.

Colloquial Substitutions

Multi-character substitutions found in colloquial/social media text. The COLLOQUIAL_SUBSTITUTIONS dict maps colloquial forms to their standard equivalents (25 entries total).

Particles

Verb Endings

Pronouns

Common Words

Adverbs and Reduplication

Contractions and Texting

Reverse Mapping: STANDARD_TO_COLLOQUIAL

The STANDARD_TO_COLLOQUIAL dictionary is the inverse of COLLOQUIAL_SUBSTITUTIONS. It maps each standard form back to its set of colloquial variants. This is built automatically at module load time.

Helper Functions

Usage Examples

PhoneticHasher Integration

Visual Confusion Detection

Tonal Variant Generation

Data Constants

Available constants and functions:

Phoneme-Grapheme Notes

E vs AI Vowels

The module correctly distinguishes:
  • (U+1031) - E vowel, IPA /e/, prefix position
  • (U+1032) - AI vowel, IPA /ɛ/, suffix position
These are phonetically distinct and should NOT be treated as interchangeable.

Aspirated vs Voiced

Consonant groups contain both aspirated and voiced variants:
  • (unaspirated) vs (aspirated) vs (voiced)
  • These sound similar and are often confused

See Also