📣
TiDB Cloud Premium is now in public preview. Unlimited growth, instant elasticity, advanced security for enterprise workloads. Try it out →

Sink to TiDB Cloud Lake



In TiDB Cloud, you can use Data Pipeline to replicate full data and incremental changes from your TiDB Cloud Premium instance to TiDB Cloud Lake, without requiring a third-party ETL tool. It first exports a full snapshot of the selected source data, and then can continuously replicate row changes so that the data in TiDB Cloud Lake stays up to date.

Restrictions

  • The TiDB Cloud Lake warehouse must be in the same region as your TiDB Cloud instance.
  • Only tables with a primary key can be replicated incrementally. Tables without a primary key are listed in the Filter results panel during pipeline creation. If included in the sync scope, their incremental replication is skipped.
  • You can create up to 100 changefeeds per TiDB Cloud Premium instance. Each data pipeline with incremental replication consumes one changefeed slot.
  • Deleting a data pipeline does not delete the data already written to TiDB Cloud Lake, nor the target databases and tables in your warehouse.

Prerequisites

Before you begin, make sure that you have:

  • A TiDB Cloud Premium instance. Note the region in which it is deployed.
  • A warehouse in TiDB Cloud Lake that is in the same region as your instance. If you do not have one yet, create it in the TiDB Cloud Lake console first. Only warehouses in the same region as your instance can be selected when you create the data pipeline.
  • An external stage bucket: an Amazon S3 bucket or an Alibaba Cloud OSS bucket. Create it in the same region as your instance.
  • The user name and password of a TiDB database user that can read the source tables.

Create a data pipeline

To create a data pipeline, you need to configure the destination, the external stage, and the replication settings.

Step 1. Configure the destination

  1. In the TiDB Cloud console, navigate to the overview page of the target TiDB Cloud Premium instance, click Data > Data Pipeline in the left navigation pane, and then click Create Data Pipeline in the upper-right corner.

  2. In the Destination area, configure the following fields:

    • Destination: select TiDB Cloud Lake.
    • Warehouse: select the target warehouse. Only warehouses in the same region as your instance are listed. If no warehouse is available, create one in TiDB Cloud Lake first, and then refresh the list.
  3. (Optional) Configure the target naming convention in the Database Prefix, Database Suffix, Table Prefix, and Table Suffix fields. All four fields are empty by default, which means that the databases and tables created in TiDB Cloud Lake keep the same names as their sources:

    • Database name: <database prefix><source database name><database suffix>
    • Table name: <table prefix><source table name><table suffix>

Step 2. Configure the external stage

An external stage is the object storage that bridges the two sides of a data pipeline: TiDB Cloud writes the exported snapshot and the captured row changes to the stage, and TiDB Cloud Lake loads the data from the stage into the target warehouse. For more information, see Why does a data pipeline require an external stage?.

TiDB Cloud Data Pipeline supports Amazon S3 and Alibaba Cloud OSS as the external stage. Create the bucket in the same region as your TiDB Cloud Premium instance, and complete the provider-side setup first. The configuration steps vary depending on your cloud provider:

    1. In the External Stage area, enter the Bucket URI of your S3 bucket in the s3://<bucket-name>/<path-to-data>/ format.

      Leave the other fields empty for now. You will fill them in after you configure bucket access using one of the methods described in the following sections.

    2. To let TiDB Cloud write data to the external stage and TiDB Cloud Lake read data from it, configure the bucket access in the Bucket Access area. Select one of the following methods and complete the authorization accordingly.

      • Method 1: Use an AWS Role ARN (recommended)

        One IAM role is shared by TiDB Cloud (which writes to the stage) and TiDB Cloud Lake (which reads from the stage), so that you configure the authorization only once and no long-lived access key is stored. You can create the role with the CloudFormation template provided by TiDB Cloud, or set it up manually in AWS.

        For the complete AWS-side setup, see Set Up an Amazon S3 External Stage for TiDB Cloud Data Pipeline. After the role is created, in the TiDB Cloud console, paste the RoleARN output value in the Role ARN field, and, if you also created an SQS queue, copy the queue URL into the SQS Queue URL field.

      • Method 2: Use an AWS access key

        For the complete AWS-side setup, including the IAM user, its permissions, and the optional SQS queue, see Bucket Access with Access Key. Then, in the TiDB Cloud console, select AWS Access Key, and fill in Access Key ID and Secret Access Key.

      After you have filled in the required information for the method you selected, click Test Connection to verify that TiDB Cloud can access the bucket. If the check fails, verify the bucket region and the permissions granted to the role or the access key, and then test the connection again.

    For the complete OSS-side setup, including the RAM user, its permissions, and the access keys, see Set Up an Alibaba Cloud OSS External Stage for TiDB Cloud Data Pipeline.

    1. In the External Stage area, enter the Bucket URI of your OSS bucket in the oss://<bucket-name>/<path-to-data>/ format.

    2. Fill in the following fields:

      • Access Key ID: the AccessKey ID of the RAM user.
      • Access Key Secret: the AccessKey Secret of the RAM user.
    3. Click Test Connection to verify that TiDB Cloud can access the bucket. If the check fails, verify the bucket region and the permissions granted to the RAM user, and then test the connection again.

    Step 3. Configure replication

    In the Replication Data area, configure how the data is replicated:

    1. Sync Mode: select the synchronization mode.

      • Full Data + Incremental Data (default): exports a full snapshot of the selected source data, and then continuously replicates row changes. This is the recommended mode for ongoing synchronization.
      • Full Data: exports a one-time full snapshot of the selected source data only. No incremental data is replicated, and changes made to the source after the snapshot are ignored.
    2. Sync Interval: the end-to-end latency target for the data pipeline. The changefeed flush cycle and the TiDB Cloud Lake polling cycle both contribute to the end-to-end latency. Shorter intervals reduce data latency but increase the number of API calls to cloud storage. The default value is shown in the console.

    3. Changefeed Capacity Units: the processing power allocated for incremental replication, shown together with the maximum replication throughput it supports. For example, 2 CCUs (the maximum replication throughput is 5,000 rows/s).

    4. TiDB Username and TiDB Password: fill in the user name and password of a TiDB database user. The data pipeline uses this account to export the full snapshot, so the account must have read access to the source tables. Incremental row changes are captured separately by the changefeed.

    5. Sync Objects: select which objects are replicated.

      • Customize (default): specify explicit rules in Table Filter Rules. The rule syntax is the same as the TiCDC table filter rules. By default, one *.* rule replicates all non-system tables. The Filter results panel shows the databases and tables that match the rules.
      • All: replicate all tables from all databases. The table filter rules configuration is hidden.

      Optionally, select Case-sensitive to make the matching of database and table names in the filter rules case-sensitive. By default, matching is case-insensitive.

    6. Pipeline Name: enter a name for the data pipeline.

    7. Click Create.

      The pipeline enters Creating while the full snapshot is being exported. For Full Data + Incremental Data, the status changes to Running when incremental replication starts.

    Manage the data pipeline

    Edit a data pipeline

    To edit a data pipeline, go to the Data Pipeline of your target TiDB Cloud Premium instance, click ... in the row of the pipeline, and then click Edit.

    Editing is disabled while a data pipeline is Running. Pause the pipeline first, then edit it, and resume it afterwards to apply the changes.

    The destination type and the sync mode cannot be changed after the pipeline is created.

    Pause and resume a data pipeline

    • Pause: stops data replication and marks the pipeline as Paused. No data is lost, and the replication progress is preserved. A pipeline cannot be paused while it is being created or while the full snapshot is being exported.
    • Resume: continues replication from where it was paused, including ingestion into TiDB Cloud Lake.

    To pause and resume a data pipeline, go to the Data Pipeline of your target TiDB Cloud Premium instance, click ... in the row of the pipeline, and then click Pause or Resume.

    Delete a data pipeline

    To delete a data pipeline, take the following steps:

    1. Go to the Data Pipeline of your target TiDB Cloud Premium instance, click ... in the row of the pipeline, and then click Delete.

    2. Read the warning and confirm the operation. Deleting a data pipeline:

      • Immediately stops all data replication.
      • Attempts to remove the TiDB Cloud Lake data source and integration task associated with the pipeline. If the removal fails, these resources might remain and require manual cleanup.
      • Does not delete the data already written to TiDB Cloud Lake.
      • Does not delete the target databases or tables in the warehouse.

    This action cannot be undone.

    See also

    Was this page helpful?