Every PDF file follows a specific internal layout. It begins with a header like %PDF-1.7 and ends with %%EOF. In between, the file holds numbered objects, a cross-reference table that records where each object sits, and a final directory that tells a viewer where to start reading. When a PDF breaks, the cross-reference table is the first thing to go. A corrupted xref table causes blank pages or errors even when every page, font, and image is intact. A rebuild tool scans the file for objects and writes a fresh map. That is the fix.

The PDF Header the First Bytes

The first line of a PDF is its header. It tells a viewer which version of the specification the file follows. A typical header reads %PDF-1.7. The version number matters. PDF 1.7 was standardised as ISO 32000-1 in 2008, and PDF 2.0 as ISO 32000-2. Newer versions introduced features that older viewers cannot handle. The header is short and rarely corrupted, but if the first bytes are missing or altered, a viewer may not recognise the file as a PDF at all.

PDF Objects the Content Building Blocks

Everything inside a PDF is stored as numbered objects. An object might hold text, fonts, images, page dimensions, or rendering instructions. Each object is written in a block that starts with an object number and a generation number, followed by the keyword obj, then the content, and finally endobj.

  • 12 0 obj ... endobj

Every object carries a unique number. The generation number tracks updates. When an object is edited, the generation increments. In a fresh file, it starts at 0. The file is a collection of numbered pieces held together by a directory. Grasp that, and the rest follows.

The Cross-Reference (Xref) Table the File's Map

How the Classic Xref Table Works

The cross-reference table is the directory. It lists the byte offset of every object in the file. In a classic xref table, each entry is exactly 20 bytes long. The first 10 digits give the byte offset, padded with leading zeros. The next 5 digits give the generation number. The final character is either n for an active object or f for a free or deleted one. The table opens with a header line that states the object number of the first entry and the total count. After the table, the keyword trailer appears, followed by the trailer dictionary.

This structure is critical because a viewer reads the file from the end. It finds the %%EOF marker, reads backward to locate the byte offset of the xref table, then uses that table to locate every object. If the xref table is damaged or missing, the viewer cannot find any object, even when every page is perfectly intact. A truncated file, where the end is cut off, becomes unreadable because it loses the xref table and the final directory first.

The Trailer and Startxref Pointing to the Map

The trailer dictionary sits after the xref table. It contains the /Root entry, which points to the document catalog object. The catalog is the starting point for locating pages, resources, and metadata. Immediately after the trailer dictionary, the keyword startxref appears, followed by the byte offset of the xref table. The file then ends with %%EOF.

Because viewers start reading from the end, the %%EOF marker and the startxref offset are the first things checked. If either is missing or corrupted, the viewer reports the file as damaged or refuses to open it. The startxref offset is the final navigation key the viewer uses to begin parsing the whole structure.

Incremental Updates and Advanced Structures

How Incremental Updates Work

PDFs can be edited without rewriting the entire file. When you add a comment or change a page, the viewer appends new objects, a new xref section, and a new final directory to the end of the file. The old xref table and directory remain in place, but the new directory points to the new xref section, which covers both old and new objects. This keeps edits fast. Over time it can build a long chain of updates.

Cross-Reference Streams and Object Streams

PDF 1.5 introduced cross-reference streams and object streams as a compressed alternative to the classic xref table and individual objects. Instead of a plain-text table, the cross-reference data is stored as a stream object that can be compressed, saving space. Object streams bundle multiple small objects into one compressed stream. These features are common in modern PDFs, but they serve the same purpose: a map from object numbers to their locations.

What Happens When the Xref Table is Damaged

If the xref table is damaged, missing, or points to wrong byte offsets, a viewer cannot find the objects that describe the pages. The file may still contain all the page content, fonts, and images, but the viewer has no way to locate them. The result: blank pages, error messages, or a flat refusal to open the file.

A rebuild tool scans the entire PDF from the beginning, looking for obj and endobj markers to reconstruct a new xref table. It reads every object it can find, determines its byte offset, and builds a fresh directory. The final directory is also regenerated with the correct /Root reference, provided the catalog object can be identified. The %%EOF marker is rewritten at the end. This process does not fix corrupted object content. It restores the map that lets the viewer reach the objects that are still intact.

Practical Repair Steps for a Damaged PDF

  1. Back up the original file. Before any repair, make a copy of the corrupted PDF so you can try different approaches.
  2. Check the file size. If the file is much smaller than expected, the end may be truncated. The xref table and final directory are likely missing.
  3. Open the file in a text editor. Scroll to the very end. Look for %%EOF. If it is missing or followed by garbage, the final directory is damaged.
  4. Use a repair tool. A tool that rebuilds the xref table scans the file for all objects and generates a new table. Online services offer this capability.
  5. Verify the result. Open the repaired file in a viewer and check that all pages display correctly.