🌐 English
Open app

Blog

Why Formatting Breaks When You Convert Files (And How to Fix It)

Published 2026-09-15 · file conversion · formatting issues · document conversion · layout problems

Why Tables, Fonts, and Margins Break During Conversion

Formatting collapses during file conversion because different formats store layout information in fundamentally incompatible ways. A Word document saves tables as objects with cell dimensions and borders, while a plain-text format cannot represent those structures at all—and even between rich formats like DOCX and PDF, the underlying data models for fonts, spacing, and page layout rarely map one-to-one.

The Real Cause: Format Data Models Don't Align

Every file format uses its own method to encode visual elements. DOCX stores fonts as references to system-installed typefaces and encodes margins in XML tags. PDF embeds font subsets as vector outlines and positions every glyph with absolute coordinates. ODT uses different XML schemas for the same layout properties. When a converter translates one format to another, it must interpret the source structure and rebuild it in the target language—and that translation is rarely perfect.

Tables break because their internal representation varies widely. A DOCX table is a grid of cells with borders, padding, and alignment properties. A Markdown table is plain text with pipes and hyphens. A PDF table is often just text positioned to look like rows and columns, with no underlying table object at all. Converting from DOCX to PDF usually preserves the visual result because PDF can render the exact layout, but converting from PDF back to DOCX forces the software to guess where table boundaries are, often producing misaligned cells or missing borders.

Fonts fail when the target format does not support embedding or when the converter substitutes a missing typeface. If your source document uses a custom font that the conversion tool cannot access, it will replace it with a default—usually Arial or Times New Roman—which changes line breaks, paragraph flow, and sometimes page count. Even when fonts are embedded, some formats store only the characters actually used in the document, so editing the converted file with new text may reveal missing glyphs.

Margins shift because formats define page geometry differently. Word documents separate page size, margins, headers, and footers into distinct properties. PDF bakes everything into fixed page dimensions. HTML has no concept of pages at all, only a continuous flow that adapts to the viewport. Converting a Word file with one-inch margins to HTML will lose that margin data unless the converter writes it into CSS, and converting HTML to PDF requires the tool to invent page breaks and margins where none existed.

How to Preserve Formatting: Fixes That Match Your Situation

If You Control the Source Format

Use the closest possible target format. Converting DOCX to PDF preserves nearly all formatting because PDF is designed to freeze layout. Converting DOCX to RTF keeps most fonts and paragraph styles because both formats share similar text-encoding models. Avoid converting rich formats to plain text (TXT, CSV, Markdown) if you need tables or fonts—those formats cannot represent complex layout.

Simplify the source document before converting. Remove custom fonts and replace them with widely supported ones like Arial, Calibri, or Times New Roman. Flatten complex tables into simpler grids without merged cells or nested structures. Set margins and spacing using whole numbers rather than fractional measurements, which some converters round inconsistently.

If You Must Convert Between Specific Formats

DOCX to PDF: Use the original application (Microsoft Word, Google Docs, LibreOffice) to export rather than a third-party converter. Native export functions understand the full data model and produce the most accurate output. If the result still shows shifted margins, check the PDF page size setting—some tools default to A4 when the source is Letter, or vice versa.

PDF to DOCX: Expect imperfect results. PDF-to-Word converters use optical layout analysis to reconstruct tables and paragraphs, which works well for simple documents but fails on multi-column layouts, text boxes, or documents with background images. After conversion, manually review every table border, heading style, and page break. Some converters offer "flowing text" versus "preserve layout" modes—choose based on whether you plan to edit heavily (flowing) or print as-is (preserve).

DOCX to HTML: Fonts and margins require manual CSS adjustments. The converter will usually inline basic styles, but you will need to define page-width containers and margin rules separately if you want the HTML to visually match the printed Word document. Tables generally convert cleanly because both formats support similar table models.

Between Office Formats (DOCX, ODT, RTF): Font substitution is the main risk. If both systems have the same fonts installed, conversion is usually clean. If not, open the converted file and reapply styles manually. Margins and page size typically survive, but complex features like text boxes or watermarks may disappear.

When Workarounds Are Your Only Option

Some format pairs simply cannot preserve certain elements. Converting a PDF with embedded fonts to plain text will always lose formatting—that is not a bug, it is the nature of plain text. In these cases, accept the limitation and plan accordingly. If you need the content for editing, convert and reformat manually. If you need the content for display, convert to an image format or leave it as PDF.

If the converter offers settings, use them. Many tools let you choose image resolution, font embedding, or table-detection sensitivity. Adjusting these before conversion often reduces breakage.

FAQ

Why do my tables turn into gibberish text after conversion? The converter interpreted the table as plain text because the target format does not support table structures, or because the source table was actually text positioned to look like a table (common in PDFs). Use a target format that supports tables, or manually reconstruct the table after conversion.

Can I prevent font substitution during conversion? Embed fonts in the source document if the format allows it, or convert to a format that embeds fonts by default (PDF). If the target format does not support embedding, the converter will always substitute missing fonts.

Why do my page margins change even when I specify the same size? The converter may be rounding margin values, mixing up page-size standards (A4 versus Letter), or adding default margins because the target format requires them. Check the converted file's page setup and adjust manually.

Choosing a Converter That Handles Formatting Better

When you need reliable conversion that minimizes formatting loss, the tool matters as much as the technique. Some converters preserve more layout data because they understand a wider range of format specifications. Testing a few files through different tools will show you which handles your specific document types best. If you are comparing options or need to convert between many formats, checking which formats a service supports can save you from discovering limitations mid-project. For one-off conversions where formatting is critical, using the native application to export remains the safest approach—but when that is not available, a converter that explicitly maps tables, fonts, and margins between format data models will break fewer elements than one that treats everything as generic text.

Start free

Share