MarkLogic Data Migration Tools and Scenarios
- Last Updated: October 6, 2026
- 3 minute read
- Progress Data Cloud
- Documentation
This topic explains the tools and migration scenarios for moving data into MarkLogic on Progress Data Cloud (PDC).
The tools used for migration planning and preparation are:
Tools
- MarkLogic Flux—The preferred tool for import, export, copy, and reprocessing. Data can be moved to and from local files, Amazon Simple Storage Service (S3) buckets, Azure, Java Database Connectivity (JDBC) sources, and MarkLogic instances. Flux also supports batching, streaming, and preview. See the MarkLogic Flux user guide for more details.
- Flux command-line interface (CLI)—Similar to Flux, but runs on an external host for long-running operations. Best used for ad hoc jobs or longer operations.
- MarkLogic Backup/Restore—Native cloud storage that supports encrypted backups, incremental backups, and journal archiving so you can restore to a specific backup or point in time.
- Journal archiving and replay—Saves the change logs so you can replay any updates that happened after the backup and keep the cutover window short.
- ml-gradle, or an equivalent Infrastructure as Code (IaC) tool—Recommended for declarative configuration and repeatable deployments of app servers, databases, forests, indexes, and security artifacts.
Migration scenarios
Choose a migration approach based on data size, acceptable downtime, and complexity. The following table summarizes common migration scenarios.
| # | Scenario | Downtime tolerance | Data size | Update frequency | Recommended option |
|---|---|---|---|---|---|
| 1 | Conduct a pilot, proof of concept, or migrate to production with a modest amount of data | Medium | Small to medium | Moderate | Flux |
| 2 | Migrate a large dataset | Minimal | Large | High | Backup/Restore with Journal Archiving |
| 3 | Migrate data within an available maintenance window | Several hours | Any, if the team can support the effort | Low to moderate | Simple Backup/Restore |
| 4 | Migrate data with zero or close to zero downtime | Near-zero | Any, if the team can support the effort | High | Blue/Green with dual-write |
Flux
Description: Use Flux to copy data from MarkLogic environments or other sources into MarkLogic on Progress Data Cloud (PDC). Flux supports iterative loads to import the data in stages and can run the same load again if necessary. You can also do a final catch-up load right before cutover to pick up late updates.
Size of migration: Small to moderate
Downtime: Low to moderate, usually seconds to minutes for the final cutover depending on the update rate.
Best suited for: Datasets with document counts in the low hundreds of millions, moderate change rates, or situations where the ease of the migration is more important than achieving the lowest possible downtime.
Backup/Restore with journal archiving
Description: Use this option to migrate large datasets when you need very low downtime and a predictable recovery state. When using backup/restore with journal archiving, create a full backup from the source, restore it into Progress Data Cloud (PDC), continuously archive and ship journals, then replay them immediately before cutover.
Size of migration: Large
Downtime: Very low. The only downtime will be the time required to pause writes on the source, ship the last journals, and replay them.
Best suited for: Large datasets with higher write rates or when you need a predictable restore process.
Simple Backup/Restore
Description: With this option, changes from the application to the source MarkLogic system are paused, a final backup of the current data is created, and then the backup is restored into Progress Data Cloud (PDC). The restored data is then verified and the application is brought back online.
Size of migration: Any practical scale
Downtime: High. Downtime includes backup time, data transfer, restore, and verification.
Best suited for: Simpler environments or where a maintenance window is available and downtime is acceptable.
Blue/Green (A/B) Parallel Systems
Description: With this option, the old and new systems are kept in sync during the transition. At the end of the migration, traffic is switched over.
Size of migration: Any practical scale
Downtime: Near-zero.
Best suited for: Mission-critical systems that need continuous operation. This option requires teams that can support dual-write and data reconciliation.