Page migrations get planned. Document migrations get discovered. Somewhere in week three of a redesign, someone runs a proper crawl and finds that the site is not four hundred pages, it is four hundred pages and nine thousand PDFs, and a meaningful share of the organization’s search visibility is sitting inside files nobody has looked at since they were uploaded.
This is normal for municipalities, universities, health authorities, associations and any organization that publishes reports. It is also the part of a migration most likely to cause a visible drop in traffic, because documents rank, documents are linked from outside, and documents are almost always moved last and least carefully.
Why documents behave differently from pages
Four differences drive everything that follows.
- They are indexed on their own. A PDF is a document in its own right, with its own ranking, its own backlinks and its own entry traffic. Moving the page that links to it does not move it.
- Their links live outside your site. Reports get cited in other organizations’ pages, in news articles, in academic references and in printed material. Those links cannot be updated.
- Nobody owns them. Pages have an author who still works there. A 2016 tender document does not.
- They carry an accessibility liability. Untagged PDFs are the most common reason an otherwise compliant site fails an audit, and a migration is the moment that liability becomes visible.
Inventory first, and record more than the filename
Crawl the site and pull the file list from the server, then reconcile the two. They will disagree, and the difference is informative: files on the server that nothing links to are orphans, and files linked but missing are already broken.
For each document, record the current URL, file size, last modified date, which pages link to it internally, whether it has external backlinks, its impressions and clicks over the past twelve months, and whether the file is tagged. That last field is usually the one nobody has, and it is the one that determines the cost of the next phase.
The inventory is a deliverable, not a working note. It is what you will reconcile against at the end, and it is what protects you when someone asks in month six why a particular file is no longer there.
Triage against evidence, not instinct
With the inventory in hand, sort every document into four buckets:
- Move as is. Current, used, and either already tagged or scheduled for remediation.
- Convert to a page. Documents that should never have been documents. A two-page fee schedule, a set of office hours, a form that could be a form. These are the highest value conversions, because an HTML page is accessible, searchable, mobile-readable and editable, and the PDF was none of those things.
- Archive. Must remain retrievable for records reasons, but does not need to be in navigation or in the index.
- Retire. Superseded, duplicated, or zero traffic and zero links and no retention obligation.
Bucket two is where the real improvement is. It is common for a fifth of a document library to be content that was published as a PDF purely because that was easier for the author. Converting those is the difference between a migration and an upgrade.
Bucket four needs a named approver. Retiring content is a governance decision, and the person who signs it off should not be the person doing the migration.
Give documents real URLs
Most document libraries carry URLs shaped by whatever uploaded them: a date-based directory, an export path, a filename with a build number in it. A migration is your one chance to fix that, and the fix is worth making because a readable document URL is a ranking and trust signal in the same way a readable page URL is.
Whatever scheme you choose, hold to two rules. One redirect hop from old to new, never a chain. And a stable path that does not encode the year the file was uploaded, so next year’s version of the same report does not need a new address.
Script the move
Nine thousand files cannot be moved by hand, and the attempt is where errors enter. Script the transfer, script the redirect generation from the inventory, and script the verification. A scripted migration can be run against a staging environment, checked, corrected and run again. A manual one cannot be repeated, which means every mistake is permanent and every correction is another manual pass.
The redirect map should be generated from the inventory rather than written separately. If the two are maintained by different people in different files, they will diverge, and the divergence surfaces as 404s after launch.
Rewrite the internal links
Redirects keep external links working. They are not an acceptable permanent answer for your own links. A site that relies on its redirect table for internal navigation is slower, harder to audit, and one server migration away from losing the table entirely.
Rewrite every internal reference to point at the new address directly, including links inside other documents where the source files still exist. Then crawl again and confirm that no internal link passes through a redirect.
Reconcile, then prove it
The reconciliation report is what turns a migration from a claim into a fact. Every row in the original inventory should resolve to exactly one of: moved and verified at a new URL, converted to a named page, archived at a stated location, or retired with a named approver. No row should be unaccounted for, and the count in the report should match the count in the inventory.
Then run the checks that catch what reconciliation misses: request every old URL and confirm a single-hop redirect to a live file, crawl the new site for broken links and redirect chains, and confirm the new document URLs appear in your sitemap.
Watch the right numbers afterwards
Document reindexing is slower than page reindexing, so resist the urge to judge the outcome in week one. Track the old URLs’ impressions declining while the new URLs’ impressions rise, and confirm the total is roughly flat. Watch crawl errors and the count of indexed documents. Expect a dip. Expect it to recover over several weeks.
What should worry you is not a temporary dip but a total that never returns to its starting point. That almost always traces back to a redirect chain, a missing rule for one directory, or a batch of documents that quietly failed to transfer. All three are findable in the reconciliation report, which is the reason to produce one.
We run document migrations this way and hand over the inventory, the redirect map, the reconciliation report and the crawl report with the site. Our content and data migration page explains the process in more detail.