Korean Name Generator

THE NOTES BEHIND THE NAME · ENGINE 0.6.2

Every choice,
out in the open.

This studio proposes a Korean-style name. It first shows a Hangul representation of the original pronunciation, then compares that with edited Korean name forms, then assigns readable Hanja. The result is a creative adaptation with visible tradeoffs.

1. Resolve pronunciation before comparing names

The dictionary contains 523 normalized name spellings and 554 readings across six guides. The default English / international guide has 498 entries, including editorial Korean renderings of Arabic, South Asian, Chinese, Japanese and Vietnamese names. Those entries are not necessarily English pronunciations. French, Spanish, Italian, German and Dutch have separate, smaller guides. Each spelling has one selected reading per guide; regional and personal pronunciations can differ. Coverage varies by language and is not comprehensive. These are editorial pronunciations, not a complete implementation of the National Institute of Korean Language’s spelling rules. Jean resolves to 진 in English and 장 in French; Michael resolves to 마이클 in English and 미하엘 in German.

Case and accents are normalized for lookup. Hyphenated or multiword given names require every part to be present. There is no automatic English fallback for another language, and no letter-by-letter fallback for an unknown name. When every part of a name is known in another language, the error names that language and offers a one-click switch; the form never switches by itself. A supplied Hangul pronunciation overrides the dictionary. This follows the principle that the sound of a name matters, rather than asserting that spelling always determines pronunciation. See the National Institute of Korean Language’s language conventions.

2. A separate, explicit surname decision

The family-name field is independent of the given-name field. With no family name and no chosen Korean surname, the result has only a given name. A Korean surname typed in Latin letters is recognized when its spelling is listed: the displayed conventional form (Kim, Lee, Park, Cho, Yoon, Lim, Ahn, Moon, Noh and so on) and variants such as the Revised Romanization forms Jeong, Jo, Yun, Im, An, Mun and No, or Rhee. Ryu and Ryoo map to 류, separately from Yoo and Yu → 유. Name cards show that conventional surname spelling and the given name in Revised Romanization. Other known family names are compared with 44 single-syllable Korean surnames. A listed Korean surname typed in Hangul, including 류 and the compound surnames 남궁, 황보, 제갈, 선우, 독고, 사공 and 서문, is kept as written and never replaced by a sound match. The family-name pronunciation dictionary is independent of the selected given-name guide and uses one reading for each of 430 spellings: mostly English, plus frequent Spanish, German, French, Italian, Dutch, Chinese, Japanese, South Asian, Vietnamese, Filipino and Eastern European ones. Spacing and apostrophes are ignored when looking one up, so O'Brien and OBrien match. A family name that is not in the dictionary is never guessed from spelling; the result is shown without a surname instead of stopping with an error. A user can still supply a Hangul pronunciation or pick a Korean surname. Changing the entered family name clears its previous pronunciation override; changing the given name clears only its own override. Changing case alone does not clear a pronunciation.

Surname scoring uses 95% syllable similarity and a small 5% editorial familiarity prior (1.0 for the first ten listed surnames; 0.9 for the rest). The compared syllable is prepared first. A leading ㅡ syllable that only carries a consonant cluster is read with the next real vowel: Smith 스미스 compares 시 and becomes 신, Brown 브라운 compares 바 and becomes 박, Clark 클라크 compares 카 and becomes 강. For foreign-surname adaptation, this studio uses initial-sound forms (ㅇ before i and y vowels, ㄴ elsewhere): Lynch 린치 compares 인 and becomes 임. This is an editorial adaptation rule, not a rule that all Korean surnames must avoid ㄹ. Entered or selected 류 stays 류. Then the opening consonant comes first. If any surname shares it or a near consonant (ㄱ·ㅋ, ㄷ·ㅌ, ㅂ·ㅍ, ㅈ·ㅊ, ㅅ·ㅆ), only those surnames compete on vowel and final, so Foster 포스터 becomes 배 instead of a surname that matches the vowel but drops the consonant. Its dedicated syllable distance weights are initial consonant 50%, vowel 40%, final consonant 10%. Compressing a family name to one syllable should preserve the opening sound before its coda. For Johnson, 존 → 조 keeps ㅈ and ㅗ while dropping ㄴ; 존 → 전 changes the vowel. The former now wins without a Johnson-specific result mapping. This prior is not measured population frequency. Only the opening sound is considered. A chosen surname bypasses that comparison. Cantonese Leung uses 렁, while Mandarin Liang uses 량. Each surname has one example Hanja spelling. A spelling match cannot determine a real family’s Hanja, clan or ancestry. The surname’s meaning is kept out of the given-name interpretation.

3. Compare sound with ordered evidence

Each Hangul syllable is decomposed into its initial consonant, vowel, and final consonant. Substitution cost is 0.45 × initial + 0.35 × vowel + 0.20 × final. Identical features cost zero; nearby features cost less than unrelated ones. An initial ㅇ is a silent onset, so it is equally far from every consonant; the ㄴ·ㅁ·ㅇ nasal neighbourhood applies only to final consonants, where ㅇ is ng. The vowel and consonant neighborhoods are explicit editorial approximations, not trained phonetic distances.

A dynamic-programming alignment finds the least-cost sequence of syllable changes, omissions, and additions. Deletion usually costs 0.8. An open syllable with ㅡ costs 0.48 because that vowel often supports a foreign consonant cluster. The first-syllable deletion penalty is multiplied by 1.25 and the last by 1.05. Addition costs 0.85. A source shorter than two syllables must gain syllables on every path, so exactly that many additions are relabelled as extensions costing 0.4. Every candidate drops by the same amount, so the ranking is unchanged and only the loose-match judgement becomes fairer. ㅐ, which writes English /æ/, is treated as near ㅏ (distance 0.5 instead of 1), so Anderson becomes 안 and Campbell becomes 강. Costs are divided by the longer syllable count and converted to a 0–1 score. The alignment is used internally to score candidates; it is a comparison path, not a linguistic derivation of the recommended name.

Sound comparison combines 45% syllable alignment + 25% consonant-sequence similarity + 25% vowel-sequence similarity + 5% first-syllable similarity. Vowels are compared in order with insertion and deletion costs of 1 and the same explicit vowel substitution distances used in syllable comparison. Sequence costs are divided by the longer sequence length. This prevents matching consonants alone from hiding a change in the vowel pattern. The consonant sequence allows a consonant to remain recognizable after crossing a syllable boundary, such as ㅁ in 임 versus 민. We do not claim to model stress, vowel duration, all liaison rules, or the original language’s full phonology. Every successful generation immediately selects a Korean name, regardless of sound score. Scores of 0.62 or above use shared-sound guidance, scores from 0.50 to below 0.62 use sound-inspired guidance, and lower scores offer a Korean name while acknowledging the different sounds. These bands only select wording; they never block a recommendation. Three alternative names appear below the card, and saving and sharing are available immediately. Scores and the comparison process are collapsed under “See details.” These are editorial guidance bands, not validated accuracy thresholds.

4. Choose a category, rank by sound, then choose Hanja

The pool contains 248 edited two-syllable names: 96 masculine, 81 feminine and 71 unisex. A masculine preference admits masculine and unisex forms (167); a feminine preference admits feminine and unisex forms (152). Unisex admits only the 71 unisex forms. These are editorial usage categories, not identity or popularity claims. A surname excludes names repeating one of its syllables.

Names are ranked by sound score alone. The former fixed 15-point name-form bonus is removed. Impression only breaks exact sound-score ties; remaining ties use stable Hangul ordering. A meaning preference chooses Hanja for each name and cannot change the ordered name list. Thus changing Chloe from Balance to Light no longer promotes a different name based on its Hanja.

All current names have at least one Hanja pair whose readings match the name and the pinned Unicode data. Candidate expansion and optional Hanja-free naming are separate future work. The displayed score is the given-name sound comparison out of 100; it excludes the surname and is neither probability nor measured linguistic accuracy. Results are deterministic for an input and engine version. Version 0.6.2 share links restore an explicitly chosen candidate; older links require regeneration.

5. Real readings and restrained interpretations

The reference dictionary embeds 2,000 characters from Unicode 17.0.0. The eligibility filter uses the union of kKoreanEducationHanja and kKoreanName, with Hangul readings and a definition. Essential recommendation and surname characters come first, then educational characters, then an editorial list of characters common in modern Korean given names, then remaining eligible characters by kGradeLevel and code point. Outside the essential set only CJK Unified Ideographs are admitted, so rare Extension A variants that many phones cannot render are excluded. This is not a “top 2,000 popular naming characters” list.

A separately edited set has 221 character/reading entries, of which 214 can be recommended. Seven context-dependent senses are reference-only. Every edited reading is checked against kHangul during data generation. English definitions in the library come from kDefinition, which primarily describes modern written Chinese. They are not certified Korean naming definitions and can differ from Korean usage. A missing definition for 澯 uses a visibly marked editorial gloss. Korean glosses and chosen English naming senses are edited separately. These are selected senses, not every dictionary meaning of a character. For example, 材 is rendered as timber rather than the interpretive potential, 宰 as a minister, and 鎭 as calm. Source membership and an ordinary Korean reading do not establish current legal approval for that particular character/reading combination.

For each candidate, the engine enumerates compatible Hanja pairs. Pair score is 0.75 × theme fit + 0.15 × meaning diversity + 0.10 × editorial preference. Balance theme uses 0.85 fit for every pair; other themes give each character 1.0 when its tag matches and 0.3 otherwise. Diversity is zero only when the two selected English senses are identical. Editorial preference averages 1 / (1 + 0.3 × option index) for the two characters, using their visible order in the edited source list. Duplicate characters cannot form a pair.

For example, 才 means talent and 珉 means a jade-like stone. “A name bringing together talent and a jade-like stone” is our compositional interpretation of 才珉; it is not an attested dictionary phrase, and it is not the etymology of James. Likewise, 朴 is handled as the surname Park and is not joined into a story about a flower.

Why Hiddink is a useful counterexample

The nickname 희동구 for Hiddink was documented in MBC’s reporting in 2002. Contemporary Kyunghyang reporting includes playful Hanja versions, including 喜東球. That example uses joy, east, and ball as wordplay. It does not establish one official Hanja name.

The nickname follows the three-syllable rhythm well. But 희 is outside our everyday surname list, and Hiddink itself is a family name rather than a given name. In our isolated sound experiment it can be entered as a naming seed: the engine compares 히딩크 with two-syllable names such as 희동 (希東), with no invented surname. That adaptation means hope + east in the chosen characters, rather than the football pun.

Sources and data maintenance

NIKL name examples distinguish Leung → 렁 from the Mandarin form Liang → 량. Such examples support a chosen reference pronunciation, not every individual’s pronunciation.

Pronunciations and the Korean name pool remain editorial data, without comprehensive independent native-speaker review. Structural and source checks do not establish linguistic accuracy. The data verification command also checks the pinned ZIP and the three extracted source files, then compares the entire generated reference dataset without rewriting it.

한국어 검토 메모

1차 엔진은 ‘발음 사전 → 초·중·종성 정렬 → 검토한 한국 이름 후보 → 독음이 맞는 한자 → 명시적인 뜻풀이’의 순서입니다. 서양 이름에 한국 이름을 직접 배정한 정답표는 없습니다. 발음 사전과 한국 이름 후보 사전을 분리하고, 가운데 비교 알고리즘이 결과를 결정합니다.

이름 어원은 아직 데이터로 내장하지 않았습니다. Sophia의 역사적 어원을 사후에 끼워 넣는 대신, 사용자가 ‘지혜’를 고르면 書 같은 의미 태그가 맞는 글자를 우선합니다. 한자별로 선택한 뜻과 두 글자를 엮은 해석을 따로 표시합니다. 이름 순위는 발음으로 결정하고 의미 선호는 한자 선택에만 사용합니다. 뜻을 바꿔도 이름 순위가 뒤집히지 않습니다.

0.6.2에서는 고정 형태 가산점을 제거하고, 모음 순서를 별도로 비교합니다. 남성형·여성형에는 공용 이름을 포함합니다. 점수에 관계없이 한국 이름을 바로 추천합니다. 62점 이상은 닮은 소리, 50점 이상 62점 미만은 소리를 참고한 이름, 50점 미만은 발음이 다른 한국 이름 제안으로 안내합니다. 다른 후보 3개와 저장·공유를 바로 제공하고 점수·비교 과정은 접어 둡니다. 음절별 치환 경로를 작명 근거처럼 설명하지 않습니다. 50·62점은 검증된 정확도 경계가 아니라 문구를 나누는 편집 기준이며, 새 점수와 구버전 점수를 품질 향상률처럼 직접 비교할 수 없습니다.

2,000자는 참고용 검색 사전이며, 그 2,000자를 무차별로 조합해 작명하지 않습니다. 추천용은 221개 독음 항목 중 문맥 의존적인 7개를 제외한 214개입니다. 다음 품질 개선은 한자 수를 늘리기보다 발음 사전의 원어민 검수, 이름 형태의 한국어 화자 평가, 실제 후보 간 선호 데이터 수집이 우선입니다. 현 버전의 점수는 그런 실증 평가를 완료한 수치가 아닙니다.

← Back to the name studio / 이름 생성기로 돌아가기