Semua artikel

Why your Chinese and Japanese tags turn into ÖܽÜÂ×, and how to fix them for good

Artikel ini belum tersedia dalam Bahasa Indonesia — menampilkan versi asli bahasa Inggris.

You move a library you've been curating for fifteen years onto a new device, open it up, and a third of your artists have been replaced by ÖܽÜÂ× and ³¯«³¨³. Nothing is broken. Nothing is lost. You are simply looking at the right bytes through the wrong decoder.

What you're actually looking at

Text on disk is bytes. Turning bytes back into characters requires knowing which encoding produced them. When that information is missing, a player has to guess — and the historical default guess is Latin-1, which maps all 256 possible byte values to a character and therefore never fails. It just quietly produces nonsense.

Here is the same name, written three ways:

Original Encoding Bytes Read as Latin-1
周杰伦 GBK D6 DC BD DC C2 D7 ÖܽÜÂ×
陳奕迅 Big5 B3 AF AB B3 A8 B3 ³¯«³¨³
椎名林檎 Shift-JIS 92 C5 96 BC 97 D1 8C E7 <0x92>Å<0x96>¼<0x97>Ñ<0x8c>ç

The bytes are intact. The label is missing.

Why tags lose their encoding

It depends on the tag format, and the differences matter:

  • ID3v1 has no encoding field at all. Each field is a fixed 30-byte slot holding raw bytes. There is nowhere to record that those bytes are GBK.
  • ID3v2.3 has an encoding byte, but only offers two values: 0x00 for ISO-8859-1 and 0x01 for UTF-16. Chinese and Japanese taggers of that era needed neither — so they wrote GBK or Shift-JIS bytes into a frame declared as ISO-8859-1. The file says Latin-1 and lies.
  • ID3v2.4 finally added UTF-8 (0x03), which is why files tagged after roughly 2010 are usually fine.
  • APEv2, used by .ape and .wv, mandates UTF-8 in its specification. Garbling here means a tagger ignored the spec.
  • Vorbis comments, used by FLAC and Ogg, also mandate UTF-8. If your FLAC tags are garbled, they were almost certainly converted from an already-broken source rather than written wrong in place.

So the common case is narrow and specific: an MP3 tagged in the 2000s on a Chinese, Taiwanese or Japanese Windows machine, where the system codepage silently supplied the encoding that the file never recorded.

Telling the encodings apart

You can usually identify the source encoding by eye:

  • Lots of Ö Ü ½ Â × and other Latin-1 accented capitals — GBK. Chinese characters in GBK mostly start with bytes in the B0–F7 range, which land on accented uppercase letters.
  • Mixed accented letters and ordinary ASCII capitals, like ¾HÄR§g — Big5. Big5's second byte often falls in the ASCII range, so real letters show up in the middle of the garbage.
  • Unprintable boxes alternating with accented characters — Shift-JIS. Its lead bytes include 0x81–0x9F, a region Latin-1 fills with control characters that have no visible glyph.

If the string contains `` (U+FFFD, the replacement character), stop guessing. That one is different, and I'll come back to it.

Fixing the display is not fixing the file

Plenty of players let you override the assumed encoding for the library view. That solves the symptom on that one player, on that one device. The bytes in the file are unchanged, so the next player you try — or the same player after a reinstall — starts guessing again from scratch.

The repair that actually holds is a round trip performed on the file:

  1. Take the mis-decoded string and encode it back to bytes as Latin-1, recovering the original byte sequence.
  2. Decode those bytes with the encoding that really produced them — GBK, Big5, Shift-JIS.
  3. Write the corrected text back into the tag, in UTF-8, in a frame that honestly declares UTF-8.

Step 3 is the one people skip, and it's the one that makes the fix permanent and portable. After it, the file carries its own encoding forever and no player has to guess again.

Sovault does all three in a batch: it detects the likely source encoding, previews the corrected text so you can confirm it before committing, and writes the result back into the file as UTF-8. The preview matters — encoding detection is a heuristic, and Big5 and GBK produce plausible-looking but different results for some byte sequences.

When it genuinely can't be recovered

Two cases are not repairable, and it's worth knowing them so you don't waste an afternoon:

Truncated ID3v1 fields. Those slots are exactly 30 bytes. A Chinese title of 15 characters is 30 bytes in GBK and fits exactly; 16 characters gets cut mid-character, and the trailing half-character is simply gone. You can recover everything up to the cut, never past it.

Already double-converted text. If something previously "fixed" these tags by decoding them as Latin-1 and saving that result as UTF-8, the damage is now baked in — but this is still recoverable, because Latin-1 is lossless across all 256 byte values. What is not recoverable is a pass through Windows-1252, which leaves five byte values (0x81, 0x8D, 0x8F, 0x90, 0x9D) undefined. Anything that hit those bytes was replaced with U+FFFD and the original byte is gone for good. That's why the replacement character is the signal to stop: it means information was destroyed, not merely mislabelled.

For everything else, the data has been sitting there intact the whole time, waiting for someone to read it correctly.