Same Meaning, Three Times the Bill β The Reality of Korean Token Inefficiency and How to Counter It
Running an AI service in Korean means spending two to three times more tokens than in English for the same task, and the API bill follows. This is not a model performance problem but a structural consequence of a tokenizer designed around English, and it can be mitigated through model choice and prompt strategy.
What a token is: not one character, not one word
A token, the smallest unit an LLM processes, does not correspond one-to-one with a character or a word. BPE (Byte Pair Encoding), the mainstream method, merges frequently occurring string patterns into single tokens and splits unfamiliar words into smaller pieces.
hamburgerβ[hamburger], 1 token (common enough to be registered whole)unbelievablyβ[un] [believ] [ably], 3 tokens (less common, so it is split)
In other words, token count is determined not by text length but by how well that language is registered in the tokenizer vocabulary.
Why English has an overwhelming advantage
The tokenizer vocabulary is built from combinations that appear frequently in the training data. Since 60 to 80 percent of the training data is English web text (Common Crawl, Wikipedia, and so on), English words and phrases are registered as a single token, while non-English languages are split down to the byte level. In English, one token corresponds to roughly four characters (about 0.75 words).
Token consumption by language (English = 1, estimated)
| Language | Script | Multiple vs. English | Cause |
|---|---|---|---|
| English | Latin | 1.0x | The overwhelming majority of training data, word-level tokenization works well |
| French/German | Latin | 1.2x ~ 1.5x | Accent marks, long German compounds split |
| Russian | Cyrillic | 2.0x ~ 2.5x | Different script, frequent byte-level splitting |
| Korean | Hangul | 2.5x ~ 3.5x | Agglutinative particle and ending combinations, plus jamo/byte splitting in older models |
| Japanese/Chinese | Kanji/Kana | 2.5x ~ 4.0x | Tens of thousands of kanji are not all in the vocabulary, so they split |
There is variance between tokenizers, so we recommend measuring it yourself with the code below.
# Measure directly with tiktoken (pip install tiktoken)
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
for s in ["Hello, how are you?", "μλ
νμΈμ, μ΄λ»κ² μ§λ΄μΈμ?"]:
print(s, "β", len(enc.encode(s)), "tokens")
Three concrete problems created by the token gap
- Cost asymmetry: billing is based on token count, not character count. A Korean service pays two to three times English for the same work.
- A smaller effective context window: with a 32K token limit, English documents admit dozens of pages while Korean documents admit roughly a third, which hurts long-document summarization and analysis.
- Slower generation: models generate tokens sequentially. If three times as many tokens must be emitted, Korean responses are physically slower.
Three practical countermeasures
- Choose models with multilingual tokenizers: GPT-4o, Llama 3, and post-Claude 3.5 models expanded non-English vocabularies substantially. A fan of older models is a spender of old prices.
- Instructions in English, output in Korean: write the system prompt and few-shot examples in English and specify only the output language in Korean, which saves a large share of input tokens.
- Split and summarize long documents: when context runs short, do not push the whole document in; chunk it and use a map-reduce style summarization.
Token efficiency rarely shows up on a model spec sheet, but it is a number that hits the operating bill directly. If you run a Korean service, start by looking at the tokenizer vocabulary.
AI Knowledge Hub
Comments (1)
λ³Έλ¬Έμ "2~3λ°°"λ ν ν¬λμ΄μ μΈλμ λ°λΌ ν¬κ² λ¬λΌμ§λ€. κ°μ μλ―Έμ λ¬Έμ₯ 4μμ tiktoken 0.13.0μΌλ‘ μ§μ μ¬λ³΄λ©΄ cl100k_base(GPT-4 μ΄κΈ° μΈλ)μμλ 2.21λ°°μ§λ§, o200k_base(GPT-4o μΈλ)μμλ 1.36λ°°λ‘ μ€μ΄λ λ€. μ¦ νκ΅μ΄ λΉν¨μ¨μ μλΉ λΆλΆμ "μΈμ΄" λ¬Έμ κ° μλλΌ "ꡬν ν ν¬λμ΄μ " λ¬Έμ μ΄λ©°, λμ μ°μ μμλ μμ΄λ‘ μ§μνκΈ°λ³΄λ€ ν ν¬λμ΄μ μΈλ κ΅μ²΄κ° λ¨Όμ λ€.
μ€μΈ‘ (μ΄ λκΈ μμ± νκ²½ μΈ‘μ κΈ°μ€, tiktoken 0.13.0)
μ§§μ ꡬ(ε₯)μμ λ°°μκ° λ ν¬κ² λ²μ΄μ§λ€. κ°μ μλ―Έμ "νκ΅μ΄μ ν ν° ν¨μ¨μ±"μ μμ΄ 5ν ν° λλΉ cl100k 16ν ν°(3.20x), o200k 8ν ν°(1.60x)μ΄λ€. λͺ μ¬κ΅¬Β·UI λ μ΄λΈΒ·νκ·Έμ²λΌ μ§§μ λ¬Έμμ΄μ λλμΌλ‘ λ€λ£¨λ νμ΄νλΌμΈμΌμλ‘ μ ν ν ν¬λμ΄μ μ μ΄λμ΄ ν¬λ€.
λ¨μ΄ λ¨μ λΆν΄ μ°¨μ΄
μλͺ¨ λ¨μλ‘ ν©μ΄μ§λ μμ μ΄ o200kμμλ 1~2ν ν°μΌλ‘ λ¬ΆμΈλ€. λ€λ§ "λ°μ΄ν°"μ²λΌ μ΄λ―Έ λ±λ‘λ λ¨μ΄λ μ°¨μ΄κ° μλ€. μ¦ μ ν ν ν¬λμ΄μ μ ν¨κ³Όλ κ³ λΉλ μΌμμ΄λ³΄λ€ μ‘°μ¬Β·μ΄λ―Έκ° λΆμ νμ©νμμ ν¬λ€.
"μ§μλ μμ΄, μΆλ ₯μ νκ΅μ΄" μ λ΅μ νκ³
μ΄ μ λ΅μ μ λ ₯ ν ν°λ§ μ€μΈλ€. λλΆλΆμ APIλ μΆλ ₯ λ¨κ°κ° μ λ ₯ λ¨κ°λ³΄λ€ λκ² μ± μ λλ―λ‘, μ€μ μ²κ΅¬μ‘μμ μ°¨μ§νλ λΉμ€μ΄ ν° νκ΅μ΄ "μΆλ ₯"μ κ·Έλλ‘ λ¨λλ€. μΆλ ₯ μͺ½μ μ€μ΄λ €λ©΄ λ€μμ΄ λ μ§μ μ μ΄λ€.
μ¬ν μ½λ
μ 리
ꡬν(cl100k κ³μ΄) λͺ¨λΈμ μ μ§νλ©΄ κ°μ μκΈμ μ ν λλΉ μ½ 1.6λ°° λ§μ ν ν°μ λΈλ€. λͺ¨λΈ κ΅μ²΄ λΉμ©μ΄ ν ν° μκΈ μ κ°μΌλ‘ μμλλμ§ κ³μ°ν κ°μΉκ° μλ€. λ€λ§ o200kμμλ νκ΅μ΄λ 1.36λ°°κ° λ¨μΌλ―λ‘ μμ ν΄μλ μλλ©°, μΆλ ₯ ν ν° μ μ΄μ ν¨κ» μ¨μΌ μ€μ§ μ κ°μ΄ λλ€.