📣
TiDB Cloud Premium is now in public preview. Unlimited growth, instant elasticity, advanced security for enterprise workloads. Try it out →

Data Pipeline



TiDB Cloud Data Pipeline replicates full data and incremental changes from your TiDB Cloud instance to TiDB Cloud Lake, without requiring a third-party ETL tool. For continuous replication, it first exports a full snapshot of the selected source data, and then continuously replicates row changes so that the data in TiDB Cloud Lake stays up to date.

You can use Data Pipeline for the following scenarios:

  • One-time data loading: export a full snapshot to TiDB Cloud Lake for initial data loading or migration.
  • Continuous data synchronization: keep TiDB Cloud Lake in sync with your TiDB Cloud instance for analytics and reporting workloads.

How it works

A data pipeline with continuous replication operates in two phases:

  1. Full snapshot export: a one-time export of the selected source tables to the external stage, from which TiDB Cloud Lake loads the snapshot.
  2. Incremental replication: continuous capture and replication of row changes (inserts, updates, and deletes) so that TiDB Cloud Lake stays current with the source.

You can also configure a pipeline to export only a full snapshot without ongoing replication.

An external stage (Amazon S3 or Alibaba Cloud OSS) is used as the intermediate storage between your TiDB Cloud instance and TiDB Cloud Lake. TiDB Cloud writes exported snapshots and captured row changes to the stage, and TiDB Cloud Lake loads data from the stage into the target warehouse. This decouples the write rate from the consumption rate, improving reliability and giving you control over cost and latency.

For more information, see Data Pipeline FAQ.

Availability

PlanStatus
TiDB Cloud PremiumData Pipeline is available in the TiDB Cloud console in private preview, and it is available upon request.
TiDB Cloud DedicatedData Pipeline is not available in the TiDB Cloud console yet. To use a data pipeline, you need to set it up manually.
TiDB Cloud EssentialData Pipeline is not available in the TiDB Cloud console yet. To use a data pipeline, you need to set it up manually.

Create a data pipeline

Refer to the guide for your plan:

View the Data Pipeline page

To view and manage your data pipelines, take the following steps:

  1. In the TiDB Cloud console, navigate to the My TiDB page.

  2. Click the name of your target TiDB Cloud Premium instance to go to its overview page, and then click Data > Data Pipeline in the left navigation pane. The Data Pipeline page is displayed.

On the Data Pipeline page, you can create a data pipeline, view a list of existing data pipelines, and manage existing data pipelines, such as pausing, resuming, editing, and deleting a pipeline.

Manage a data pipeline

Pause and resume a data pipeline

  • Pause: stops data replication and marks the pipeline as Paused. No data is lost, and the replication progress is preserved. A pipeline cannot be paused while it is being created or while the full snapshot is being exported.
  • Resume: continues replication from where it was paused, including ingestion into TiDB Cloud Lake.

To pause and resume a data pipeline, navigate to the Data Pipeline page of your target TiDB Cloud Premium instance, click ... in the row of the pipeline, and then click Pause or Resume.

Edit a data pipeline

To edit a data pipeline, navigate to the Data Pipeline page of your target TiDB Cloud Premium instance, click ... in the row of the pipeline, and then click Edit.

Editing is disabled while a data pipeline is Running. Pause the pipeline first, then edit it, and resume it afterwards to apply the changes.

The destination type and the sync mode cannot be changed after the pipeline is created.

Delete a data pipeline

To delete a data pipeline, take the following steps:

  1. Navigate to the Data Pipeline page of your target TiDB Cloud Premium instance, click ... in the row of the pipeline, and then click Delete.

  2. Read the warning and confirm the operation. Deleting a data pipeline:

    • Immediately stops all data replication.
    • Attempts to remove the TiDB Cloud Lake data source and integration task associated with the pipeline. If the removal fails, these resources might remain and require manual cleanup.
    • Does not delete the data already written to TiDB Cloud Lake.
    • Does not delete the target databases or tables in the warehouse.

This action cannot be undone.

See also

Was this page helpful?