Why Korean Text Turns Into Boxes (Font & Encoding Fixes)
What Causes Korean Text to Display as Boxes
Korean characters turning into empty boxes or squares happens for three distinct reasons: the font containing Korean glyphs is not embedded in the document, the file's character encoding is set incorrectly, or the conversion process between formats stripped the necessary font data. The cause determines which fix will actually work, so identifying it correctly matters more than trying random solutions.
Most often, font embedding is the culprit when you open a PDF or document on a different computer. The original creator's system had the Korean font installed, but the file itself does not contain that font data. Character encoding issues typically show up when opening text files or older document formats where the system guesses wrong about which character set to use. Conversion problems occur when moving between formats like DOCX to PDF or vice versa, especially through tools that do not preserve font information properly.
How to Identify Which Problem You Have
Open the file and look at the boxes. If every Korean character is a box but Latin characters appear normal, and the problem only happens on certain computers, you have a font embedding issue. If some Korean characters display correctly while others are boxes, or you see question marks mixed with boxes, that points to character encoding. If the boxes appeared immediately after converting the file from one format to another, the conversion process lost the font data.
Check the file properties if possible. In PDF readers like Adobe Acrobat, go to File > Properties > Fonts to see which fonts the document calls for and whether they are embedded. If a Korean font is listed as "not embedded" or missing entirely, you have confirmed a font problem. For text files, try opening in different programs or explicitly selecting UTF-8 encoding when opening to test for encoding issues.
Fixing Font Embedding Problems
If the font is not embedded, you have two options: install the missing font on your system or re-create the document with embedded fonts. Installing fonts only works as a temporary solution for viewing. Download and install common Korean fonts like Nanum Gothic, Malgun Gothic, or Batang from reputable sources. After installation, reopen the file and the characters should display.
For a permanent fix that works on any computer, the document creator must embed fonts. In Microsoft Word, go to File > Options > Save and check "Embed fonts in the file." In Adobe Acrobat or other PDF creation tools, ensure font embedding is enabled in export settings before creating the PDF. If you received the file from someone else and cannot get a properly embedded version, you will need to either keep the font installed or convert the document yourself with embedding enabled.
Fixing Character Encoding Mismatches
When encoding is wrong, you need to reopen the file with the correct character set specified. For plain text files, use a text editor that lets you choose encoding. In Notepad++, click Encoding > Character Sets > Korean > EUC-KR or UTF-8 depending on which displays correctly. In Sublime Text or VS Code, click the encoding indicator at the bottom and select the correct option.
For CSV files opened in Excel, do not double-click the file. Instead, open Excel first, use Data > Get External Data > From Text, and manually select UTF-8 or EUC-KR in the import wizard. Excel often defaults to the wrong encoding, causing immediate box display. Once correctly imported, save the file in a Unicode format to prevent future problems.
If you are coding or working with databases, ensure your entire pipeline uses UTF-8. Check that your text editor saves files as UTF-8, your database columns use UTF-8 collation, and your application explicitly requests UTF-8 when reading files. Mixing encodings at different stages creates box characters even when each individual component works correctly.
Fixing Conversion-Related Issues
When boxes appeared after converting between formats, the conversion tool stripped font information or used incompatible settings. The solution depends on your specific conversion path. Converting PDF to DOCX often loses embedded fonts because Word tries to substitute system fonts. Converting DOCX to PDF can lose fonts if the PDF creation engine does not embed them by default.
Use conversion tools that explicitly support font embedding and Unicode. When converting to PDF, choose settings like "PDF/A" which requires full font embedding, or manually enable font embedding in advanced options. When converting from PDF to editable formats, accept that you may need to reapply fonts manually after conversion, as this direction rarely preserves font data perfectly.
For web-based converters, the problem is worse because they run on servers without your fonts installed. If you frequently convert documents with Korean text, use desktop software where you control font installation, or verify that the online tool specifically states it handles CJK (Chinese, Japanese, Korean) fonts. Some professional conversion services maintain font libraries for this purpose, but free tools usually do not.
FAQ
Q: Will changing the system language to Korean fix box characters? No, system language settings do not embed fonts into existing files or change their encoding. You must address the specific cause directly.
Q: Can I extract text from a file showing boxes? Usually yes if the underlying character data is intact. Copy the text and paste into a program with proper font and encoding support to see if it displays correctly there.
Q: Why do boxes appear only when I email files? The recipient's computer lacks the necessary fonts and your file did not embed them. Email does not alter file contents, but it exposes font embedding problems when files move between systems.
Getting Reliable Conversions
When you need to convert documents containing Korean text and cannot risk losing font data, the format you choose and the tool's handling of Unicode matter significantly. Some formats inherently support font embedding better than others, and understanding these differences prevents repeated box-character problems. For reliable document handling across different formats while preserving font information, using specialized conversion services like those at OmniDesk formats ensures that character encoding and font data transfer correctly between file types, which matters especially when dealing with non-Latin scripts that standard converters often mishandle.