--- title: "Same Meaning, Three Times the Bill — The Reality of Korean Token Inefficiency and How to Counter It" date: 2026-09-29 model: admin category: knowhow summary: "Korean consumes two to three times more tokens than English for the same meaning. We lay out tokenizer mechanics, per-language consumption, the cost, context, and speed consequences, and practical countermeasures." tags: tokens,tokenizer,Korean,cost-optimization,BPE --- Running an AI service in Korean means spending two to three times more tokens than in English for the same task, and the API bill follows. This is not a model performance problem but a structural consequence of a tokenizer designed around English, and it can be mitigated through model choice and prompt strategy. ## What a token is: not one character, not one word A token, the smallest unit an LLM processes, does not correspond one-to-one with a character or a word. BPE (Byte Pair Encoding), the mainstream method, merges frequently occurring string patterns into single tokens and splits unfamiliar words into smaller pieces. - `hamburger` → `[hamburger]`, 1 token (common enough to be registered whole) - `unbelievably` → `[un] [believ] [ably]`, 3 tokens (less common, so it is split) In other words, token count is determined not by text length but by **how well that language is registered in the tokenizer vocabulary**. ## Why English has an overwhelming advantage The tokenizer vocabulary is built from combinations that appear frequently in the training data. Since 60 to 80 percent of the training data is English web text (Common Crawl, Wikipedia, and so on), English words and phrases are registered as a single token, while non-English languages are split down to the byte level. In English, one token corresponds to roughly four characters (about 0.75 words). ## Token consumption by language (English = 1, estimated) | Language | Script | Multiple vs. English | Cause | | --- | --- | --- | --- | | English | Latin | 1.0x | The overwhelming majority of training data, word-level tokenization works well | | French/German | Latin | 1.2x ~ 1.5x | Accent marks, long German compounds split | | Russian | Cyrillic | 2.0x ~ 2.5x | Different script, frequent byte-level splitting | | Korean | Hangul | 2.5x ~ 3.5x | Agglutinative particle and ending combinations, plus jamo/byte splitting in older models | | Japanese/Chinese | Kanji/Kana | 2.5x ~ 4.0x | Tens of thousands of kanji are not all in the vocabulary, so they split | There is variance between tokenizers, so we recommend measuring it yourself with the code below. ```python # Measure directly with tiktoken (pip install tiktoken) import tiktoken enc = tiktoken.get_encoding("cl100k_base") for s in ["Hello, how are you?", "안녕하세요, 어떻게 지내세요?"]: print(s, "→", len(enc.encode(s)), "tokens") ``` ## Three concrete problems created by the token gap 1. **Cost asymmetry**: billing is based on token count, not character count. A Korean service pays two to three times English for the same work. 2. **A smaller effective context window**: with a 32K token limit, English documents admit dozens of pages while Korean documents admit roughly a third, which hurts long-document summarization and analysis. 3. **Slower generation**: models generate tokens sequentially. If three times as many tokens must be emitted, Korean responses are physically slower. ## Three practical countermeasures 1. **Choose models with multilingual tokenizers**: GPT-4o, Llama 3, and post-Claude 3.5 models expanded non-English vocabularies substantially. A fan of older models is a spender of old prices. 2. **Instructions in English, output in Korean**: write the system prompt and few-shot examples in English and specify only the output language in Korean, which saves a large share of input tokens. 3. **Split and summarize long documents**: when context runs short, do not push the whole document in; chunk it and use a map-reduce style summarization. Token efficiency rarely shows up on a model spec sheet, but it is a number that hits the operating bill directly. If you run a Korean service, start by looking at the tokenizer vocabulary.