Migration Playbook

Confluence GitBook

Confluence to GitBook: The Complete Migration Playbook

A 39-step runbook across six phases — track your progress, and open the right tool at every step.

0 / 39 steps complete 0%
TL;DR

GitBook's native importer caps at 20 pages. Real Confluence migrations need XML-to-Markdown conversion, Git Sync, link rewriting, and staged QA — or a managed service.

There is no one-click migration path from Confluence to GitBook that preserves macros, attachments, and internal links. GitBook's built-in importer caps at 20 pages and scrapes rendered output, losing macro structure entirely. Confluence stores content in a proprietary XHTML-based storage format with custom XML elements for macros. GitBook is a Markdown-first, Git-backed documentation platform with no macro concept. Real migrations require XML-to-Markdown conversion, a SUMMARY.md for hierarchy mapping, attachment re-hosting through Git, and a deliberate redirect plan for old Confluence URLs.

Read this first

Pair-specific gotchas that catch teams out. Each one has cost somebody a weekend.

GitBook's sweet spot

External-facing developer docs, API references, and product guides. If your primary use case is broad internal knowledge management with deep Jira integration, GitBook may not be the right target. See our Notion vs. Confluence comparison for alternative paths.

The macro problem is the migration

If your Confluence instance uses only basic text, headings, lists, tables, and images, the conversion is manageable. If your pages are macro-heavy — and most enterprise Confluence instances are — every macro type requires a custom conversion rule or a conscious decision to drop the content.

Do not ignore 429s

Sustained abuse of any API will eventually get your token revoked. Build the backoff into your scripts from the start. Pushing data blindly into GitBook will result in dropped pages and a corrupted documentation hierarchy.

Deep nesting limits

Confluence allows virtually infinite page nesting. GitBook supports nesting, but overly deep hierarchies create a poor user experience in GitBook's sidebar. Plan to flatten your architecture during the transformation phase — keep nesting to three levels or fewer where possible. (gitbook.com)

Git Sync gotcha

If Git Sync is enabled, manage README.md in the repository, not in the GitBook UI, or you can create duplicate readmes and rendering conflicts. (gitbook.com)

The runbook

Work top to bottom. Tick steps as you go — your progress is saved in this browser.

01 Discovery Decide what content deserves to move before you plan how to move it. 0/5

Objective A content inventory with a keep/rewrite/retire decision on every article and an agreed URL strategy.

  1. Inventory all content in Confluence

    Content lead 2-3 days

    Count articles, categories, attachments, images and embedded media, and pull page views and last-updated dates for each article. Usage data is what makes the next decision defensible rather than political.

    Data Profiler Get real record counts instead of estimating from memory
  2. Make a keep, rewrite or retire call on every article

    Content lead 1-2 weeks

    Most knowledge bases are half stale. Migrating everything imports the staleness and doubles the work; use views and last-updated to triage, and get the owning team to confirm. This usually removes 30-50% of scope.

    Migrating stale content is the most common knowledge-base migration mistake — it costs effort and actively degrades the new site.

  3. Agree the URL and redirect strategy

    SEO / web 2-3 days

    Decide the GitBook URL structure and whether you can serve 301 redirects from the old paths. Public help centres carry real search traffic and inbound links; losing it is a measurable commercial impact.

    Without 301 redirects from old article URLs you lose accumulated search ranking and every external link and bookmark breaks.

  4. Map the information architecture

    Content lead 3-5 days

    Document the current category tree and design the target one, checking whether GitBook supports your nesting depth. Deeply nested hierarchies frequently have to be flattened, which changes navigation for everyone.

  5. Confirm permissions, audiences and localisation scope

    Content lead 2-3 days

    Establish which content is public, internal or restricted, and how GitBook models that. Then confirm how many locales you have and whether translation relationships between articles survive the move.

Confluence → GitBook specifics

Docs-as-code adoption
Engineering teams want documentation to live in Git, flow through pull requests, and ship alongside code. GitBook's native Git Sync makes this the default workflow. Confluence's storage format makes round-trips between Confluence and Git essentially impossible without serious engineering pain.
Simpler public-facing docs
Confluence was built as an internal wiki. Publishing external developer docs from Confluence requires fighting the platform's defaults. GitBook is purpose-built for published technical documentation with custom domains, branded sites, and clean URL structures.
Cost at scale
Confluence Cloud pricing compounds with user count. GitBook's per-site + per-user model can be cheaper for small, focused documentation teams — though it scales in two dimensions, so model the math for your team size.
Atlassian Data Center EOL pressure
Teams on Confluence Data Center are facing the 2029 end-of-life timeline, which forces a platform decision regardless.
GitBook has no equivalent of Confluence's macro ecosystem
Any migration from a macro-heavy Confluence instance involves permanent information loss unless you redesign those pages for GitBook's block model.

Don't move on until

  • Full content inventory with page views and last-updated dates
  • Keep / rewrite / retire decision recorded per article
  • URL and redirect strategy agreed with whoever owns SEO
02 Data Audit Audit the markup, the links and the assets — that is where KB migrations break. 0/9

Objective A content export with markup, internal links and every embedded asset accounted for.

  1. Export content and assess markup fidelity

    Content engineer 2-3 days

    Export articles in the richest format available and inspect what survived: tables, code blocks, callouts, nested lists, anchors and embedded video. Rich formatting is where fidelity is lost, and it is lost quietly.

    HTML-to-Markdown conversion routinely mangles nested lists, tables and code blocks. Inspect the output rather than trusting the converter.

    Data Profiler Profile the Confluence export for nulls, outliers and type drift
  2. Inventory every internal link and cross-reference

    Content engineer 2-3 days

    Extract all internal links, anchor links and article cross-references. These break by default: the target URL structure differs, so every internal link needs rewriting as part of the load, not afterwards.

    Internal links left pointing at old URLs turn the new knowledge base into a maze of 404s on day one.

  3. Inventory images, attachments and embedded media

    Content engineer 2 days

    List every asset with its URL, size and type, and confirm each still resolves. Assets hosted on the old platform's CDN will 404 the moment you decommission it, so they must be rehosted, not referenced.

    Images referenced from the source platform's CDN break when the old account closes. Download and rehost every asset.

  4. Find and fix broken links and orphans

    Content lead 2-3 days

    Crawl for existing broken internal and external links, and find articles no category links to. Fix them before migrating — a migration is a bad time to discover pre-existing rot.

  5. Check for PII and internal information in public content

    Compliance 1-2 days

    Scan for customer names, internal hostnames, credentials in code samples and screenshots containing real data. Republishing these on a public help centre is a disclosure, and screenshots are the usual culprit.

    PII & Compliance Scanner Find regulated fields before they land in a new system
  6. Normalise metadata

    Content engineer 1-2 days

    Standardise authors, tags, timestamps to UTC, and locale codes. Author mapping needs a decision for people who have left — attribution to a deleted user usually fails the import.

  7. Inventory macros

    Query using CQL or the API to identify which macros are used and how frequently. This determines your conversion scope.

  8. Identify orphan pages

    Pages not linked from the page tree that may be missed in a hierarchical crawl.

  9. Build the link map

    Create a lookup table of pageId → new Markdown path before starting conversion.

Confluence → GitBook specifics

REST API extraction (Recommended for developers)
Fetch page content via the Confluence REST API, which returns the raw storage format XML. This gives you the most programmatic control over macro translation and requires handling pagination.
Space HTML export
Confluence generates a ZIP of rendered HTML files. This is easier to parse into Markdown initially, but comments are excluded, blog posts are excluded in Cloud, and you lose the underlying macro data needed for high-fidelity conversion. (support.atlassian.com)
Backup Manager (XML Site/Space Export)
Generates a massive ZIP containing an entities.xml file with the raw storage format for the entire space. It provides absolute data fidelity but requires writing a robust XML parser to extract the pages without exhausting memory.
Count pages per space
Use the API: GET /wiki/api/v2/spaces/{id}/pages?limit=1 and check the total.
Catalog attachments
Total size determines whether inline download or bulk export is more practical. Flag files over 100 MB for external hosting.

Don't move on until

  • Content exported with markup fidelity assessed
  • Every internal link and asset reference inventoried
  • Broken links and missing assets fixed or logged
03 Field Mapping Map structure, metadata, permissions and — above all — URLs. 0/6

Objective A mapping covering article fields, taxonomy, permissions and a complete old-to-new URL map.

  1. Map the article schema

    Content engineer 2 days

    Map title, body, excerpt, author, dates, status, tags, SEO metadata and any custom properties. Confirm which fields GitBook lets you set on import versus which it computes — computed dates are a common surprise.

    Schema Mapper Opens pre-loaded with the Confluence → GitBook field pair
  2. Map the taxonomy and hierarchy

    Content lead 2-3 days

    Map categories, sections and tags to the target structure, resolving any nesting-depth limit explicitly. If you must flatten, decide how the lost level is preserved — usually as a tag or a title prefix.

    JSON to CSV Converter Flatten nested API responses into a reviewable sheet
  3. Map permissions and audience segmentation

    Content lead 2 days

    Map public, logged-in, and role-restricted visibility to GitBook's model. Verify the mapping deliberately: internal content accidentally published publicly is the highest-severity failure in this whole category.

    Permission mapping errors publish internal documentation to the open web. Verify visibility on every restricted article after load.

  4. Build the complete old-to-new URL map

    SEO / web 2-3 days

    Produce a row per article mapping the old URL to the new one, then confirm exactly where the 301s will be served — GitBook, a CDN, or your own web layer. Without this artifact the redirect step cannot be executed at all.

  5. Define the markup conversion and link-rewrite rules

    Content engineer 3-5 days

    Specify how each markup construct converts and how internal links are rewritten using the URL map. Write it as a repeatable transform, not manual edits — you will run it more than once.

    Data Format Converter Reshape the export into the format GitBook's importer expects
  6. Plan localisation and freeze the spec

    Content lead 1-2 days

    Confirm how translated articles link to their source language in GitBook, then version and sign off the mapping spec.

Don't move on until

  • Article schema and taxonomy mapped
  • Permission and audience model mapped to target equivalents
  • Complete URL map produced and redirect method confirmed
04 Test Migration Pilot the hardest articles, then read them. 0/6

Objective A pilot load whose formatting, links, assets and search all hold up under human review.

  1. Configure GitBook with the agreed structure

    Content engineer 3-5 days

    Create the category tree, permission groups, locales and branding before loading. Articles loaded before their categories exist land uncategorised and have to be moved by hand.

  2. Pick the most difficult articles as the pilot

    Content lead 0.5 day

    Choose 20-50 articles for difficulty: the longest, the most heavily formatted, ones with tables and code blocks, deep internal linking, many images, embedded video, restricted visibility, and non-Latin scripts. Easy articles prove nothing.

  3. Run the conversion and load with link rewriting

    Content engineer 2-3 days

    Apply the markup conversion, rewrite internal links from the URL map, upload and re-reference assets, then load. Log every conversion warning rather than suppressing it.

  4. Read every pilot article side by side

    Content lead 2-3 days

    Open source and target together and compare rendering. This step is manual on purpose: no automated check catches a table that collapsed into a paragraph or a code block that lost its indentation.

  5. Click every link and load every asset

    Content engineer 1-2 days

    Verify each internal link resolves, each image loads from the new host, each attachment downloads and each embed plays. Assets still served from the old CDN are the defect that surfaces only after decommissioning.

    Migration Validation Tool Diff the pilot batch against source before scaling up
  6. Test search and permissions

    Content lead 1-2 days

    Search for known terms and confirm the right articles rank, then verify every restricted pilot article is invisible to an anonymous browser. Test permissions from a logged-out session, not an admin one.

Don't move on until

  • Complex articles render correctly with formatting intact
  • Every internal link and asset in the pilot resolves
  • Search returns sensible results for the pilot content
05 Cutover Publish, redirect, and keep the search traffic. 0/6

Objective All in-scope content live in GitBook with redirects serving and search engines informed.

  1. Load the full content set ahead of the switch

    Content engineer 1-2 weeks

    Run the full conversion and load into GitBook, unpublished or on a staging domain. Content migration differs from data migration here: you can stage the whole thing before anyone sees it.

  2. Publish the runbook with the redirect step first-class

    Project manager 1 day

    Sequence the freeze, final delta, publish, redirect activation, sitemap submission and link updates, with owners for each. Redirect activation is the step with lasting commercial consequences, so it gets explicit ownership.

  3. Freeze editing and migrate the delta

    Content lead 2-4 hours

    Stop editing in Confluence, then convert and load anything changed since the full load. Announce the freeze to every team that publishes — content teams are used to editing whenever they like.

  4. Publish and verify permissions live

    Content lead 2-4 hours

    Publish the content set, then immediately verify restricted articles are not publicly reachable using an anonymous session. Do this before announcing the new site, not after.

    Verify restricted content from a logged-out browser. An admin session will show you everything and tell you nothing.

    Migration Validation Tool Confirm the final delta landed before you reopen
  5. Activate 301 redirects and submit the sitemap

    SEO / web 2-4 hours

    Turn on the redirects from the URL map, then spot-check a sample of high-traffic old URLs and confirm each returns 301 to the right article. Submit the new sitemap and keep the old one reachable until search engines have recrawled.

    Redirect chains and redirect loops both leak ranking. Verify each redirect resolves in a single hop.

  6. Repoint in-product and support links

    Content lead 1-2 days

    Update help links embedded in your product, in support macros, in email templates and in onboarding material. These are the links your existing customers actually use, and they are easy to forget.

Don't move on until

  • All content loaded, categorised and correctly permissioned
  • 301 redirects live and verified from a sample of old URLs
  • Sitemap submitted and support links repointed
06 Validation Watch traffic, links and search rankings for weeks, not hours. 0/7

Objective Verified content completeness, healthy redirects, and search traffic recovered to baseline.

  1. Reconcile content counts and assets

    Content engineer 1-2 days

    Compare article counts by category and status, plus asset counts, against source. Confirm every article in the keep list is present and every retired one genuinely is not.

    Migration Validation Tool Reconcile Confluence and GitBook record-for-record
  2. Crawl the new site for broken links and assets

    Content engineer 1-2 days

    Run a full crawl for 404s, broken images and missing attachments, and fix everything it finds. Repeat the crawl after the fixes rather than assuming they worked.

  3. Verify the redirects at scale

    SEO / web 1-2 days

    Test every mapped old URL for a single-hop 301 to the right destination. Chains and loops both leak ranking and are invisible unless you check the whole map, not a sample.

  4. Monitor organic traffic and rankings for four to eight weeks

    SEO / web 4-8 weeks

    Track organic sessions, impressions and rankings for your top articles against baseline. A dip in the first two weeks is normal; one that has not recovered by week six is a redirect or indexing problem to investigate.

    Do not decommission the old platform until search traffic has recovered — you may still need the old URLs to diagnose a ranking loss.

  5. Verify search, permissions and feedback loops

    Content lead 1-2 days

    Confirm on-site search returns good results for real queries, re-verify restricted content from a logged-out session, and check article feedback and analytics are collecting.

  6. Sign off and decommission on a delay

    Project sponsor 1 day

    Get acceptance against the Discovery criteria, keep Confluence available read-only until traffic has recovered, take a final export, and only then close the account.

  7. Run UAT with content owners

    The people who wrote the docs should verify accuracy.

Confluence → GitBook specifics

Page count
Compare the number of pages in Confluence vs. pages in GitBook.
Spot-check macros
Review 10–15 pages that were macro-heavy in Confluence. Verify the conversion rendered correctly.
Internal links
Click through cross-references. Broken links show up in GitBook's sidebar as warnings.
Images and attachments
Verify all images render. Missing images appear as broken Markdown image references.
Table integrity
Complex merged-cell tables from Confluence may not convert cleanly to Markdown tables.

Don't move on until

  • Article counts reconciled and no broken links remain
  • Redirects returning 301 with no chains or loops
  • Organic traffic recovered to within tolerance of baseline

Field mapping reference

The field-by-field mapping for each object. Use this as the starting point for your mapping spec.

Concept Equivalent 10 fields
Confluence fieldGitBook fieldNotes
Space Space 1:1 mapping. Each Confluence space becomes a GitBook space.
Space home page First page in the space TOC GitBook treats the first page as the landing page.
Page tree (nested pages) Pages with hierarchy Defined by folder structure + SUMMARY.md.
Blog posts Pages (flat) GitBook has no blog concept. Blog posts become regular pages.
Page labels — No equivalent. Use folder organization instead.
Comments — Not migrated. GitBook has comment threads, but no import path.
Attachments Files in Git repo Must be downloaded, placed in repo, and re-linked.
Space permissions Space member roles Coarser model. Map carefully.
Page restrictions — GitBook permissions are at space level, not page level.
Collections (space groups) Collections Conceptual equivalent exists.

Risk matrix

Per-object risk for this pair. Plan extra validation around anything marked high.

ObjectRiskNotes
Layout Macros (columns) high No Markdown equivalent — content must be linearized
Jira Issue Macros high No GitBook equivalent — become plain text links
Internal Links high All Confluence URLs break without a pageId-to-path mapping
Inline Comments high Stored separately from page body — not included in any export
Page Version History high Does not migrate — Git commit history starts fresh
Nested Tables high Markdown does not support nested tables — must restructure
Draw.io / Gliffy Diagrams medium Must be exported as PNG/SVG images and re-embedded
Page Body Content low Standard text, headings, lists, and simple tables convert cleanly
Code Blocks low Direct conversion to Markdown fenced code blocks
Info/Warning Panels low Map to GitBook hint blocks with type matching

The hard parts

What makes this specific migration difficult, beyond the mechanics.

Storage Format Conversion

Confluence's proprietary XML with ac:structured-macro elements must be parsed and converted to clean Markdown with custom rules per macro type.

Macro Translation

Each Confluence macro requires a dedicated conversion rule. Layout macros, Jira embeds, and page includes have no direct GitBook equivalent.

Native Import Cap

GitBook's built-in import panel is limited to 20 pages and 20 files per import, making it unusable for real migrations.

Attachment Re-hosting

Attachments must be downloaded from Confluence, placed in a Git repository or external storage, and re-linked in Markdown files.

Deep Nesting Limits

Confluence allows unlimited page nesting. GitBook recommends three levels or fewer for usable sidebar navigation.

Tools used in this playbook

All free, all run entirely in your browser — nothing is uploaded.

FAQ

Can I import Confluence pages directly into GitBook?

Yes, but with hard limits. GitBook's built-in import panel supports Confluence by URL or file upload, but it caps at 20 pages and 20 files per import. It scrapes the rendered page, not the storage format, so macro structure is lost. For larger migrations, GitBook recommends using Git Sync with a repository of pre-converted Markdown files.

What happens to Confluence macros when migrating to GitBook?

Confluence macros use a proprietary XML format (ac:structured-macro) that has no direct Markdown equivalent. Simple macros like code blocks and info panels map to GitBook blocks. Complex macros like Jira issue embeds, Draw.io diagrams, and multi-column layouts either require manual conversion, approximation as images or plain text, or are lost entirely. Each macro type needs a custom conversion rule.

What are the GitBook API rate limits for migration scripts?

GitBook's API returns HTTP 429 when rate limits are exceeded, with X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset headers. Migration scripts must implement exponential backoff. Sustained rate limit violations risk token revocation.

Can I keep Confluence and GitBook in sync during migration?

Bidirectional sync between Confluence and GitBook is not practical. Confluence's XML storage format and GitBook's Markdown are fundamentally incompatible for clean round-trips. The reliable approach is a one-way cutover: extract from Confluence, convert, import into GitBook, run a final delta sync, then decommission the Confluence space.

How long does a Confluence to GitBook migration take?

A small space (under 50 pages, minimal macros) can be migrated in a day using the Git Sync pipeline. Enterprise instances with 500+ pages, heavy macro usage, and complex hierarchies typically take 1–2 weeks including conversion script development, testing, and validation — or 2–5 days with a managed migration service.

Or skip all of this and let us handle it

Book a 30-minute call and we'll scope your migration in a single session.