📣
TiDB Cloud Premium is now in public preview. Unlimited growth, instant elasticity, advanced security for enterprise workloads. Try it out →

Sink to TiDB Cloud Lake



This guide walks you through the end-to-end setup of a data pipeline from a TiDB Cloud Dedicated cluster to TiDB Cloud Lake: export a full snapshot to Amazon S3 with Dumpling, create a changefeed to continuously write incremental changes to the same S3 location, and configure TiDB Cloud Lake to load both the snapshot and incremental data.

Restrictions

  • The TiDB Cloud Lake warehouse must be in the same region as your TiDB Cloud Dedicated cluster.
  • Only tables with a primary key can be replicated incrementally.
  • To create the cloud storage changefeed, your TiDB Cloud Dedicated cluster must run v7.1.1 or later. For more information, see Sink to Cloud Storage.
  • The pipeline requires manual setup and maintenance of AWS IAM resources and credentials, a changefeed, and TiDB Cloud Lake integrations.
  • For more information on DDL, DML, and column type support, see Data Pipeline SQL Compatibility for TiDB Cloud Lake.

Prerequisites

Before you begin, make sure that you have:

  • A TiDB Cloud Dedicated cluster. Note the region in which it is deployed.
  • A network connection from the machine that runs Dumpling to your cluster. You can use a public connection, VPC peering, or a private endpoint. This guide uses a public connection as an example.
  • A SQL user that can read the source tables. This guide uses root as an example. For the required privileges, see Required privileges.
  • An Amazon S3 bucket (for example, s3://my-datapipeline-bucket) in the same region as your TiDB Cloud Dedicated cluster.
  • A TiDB Cloud Lake warehouse in the same region as your TiDB Cloud Dedicated cluster.

Step 1. Prepare S3 bucket access

The data pipeline components (Dumpling, the changefeed, and TiDB Cloud Lake) all require access to the same S3 bucket. Create an IAM user with the required permissions for the target S3 bucket, create an access key for the user, and use the same access key for all three components.

  1. Open the IAM Console and create an IAM user (for example, tidb-cloud-datapipeline-user).

  2. Attach the following permissions policy to the user. Replace <your-bucket-name> and <your-prefix> with your actual values:

    { "Version": "2012-10-17", "Statement": [ { "Sid": "S3BucketAccess", "Effect": "Allow", "Action": [ "s3:ListBucket", "s3:GetBucketLocation" ], "Resource": "arn:aws:s3:::<your-bucket-name>" }, { "Sid": "S3ObjectAccess", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:GetObjectVersion", "s3:DeleteObjectVersion" ], "Resource": "arn:aws:s3:::<your-bucket-name>/<your-prefix>/*" } ] }
  3. Create an access key for the user, and record the Access Key ID and Secret Access Key. You need them when you export the snapshot, create the changefeed, and configure TiDB Cloud Lake.

Step 2. Export a full snapshot with Dumpling

TiDB Cloud Dedicated does not provide the export feature in the TiDB Cloud console, so export the full snapshot with Dumpling.

1. Prepare the network and SQL user

  1. Make sure your TiDB Cloud Dedicated cluster is reachable from the machine that runs Dumpling. This guide uses a public connection. If you use a public connection, add the machine's IP address to the IP access list of the cluster. For more information, see Connect to TiDB Cloud Dedicated via Public Connection and Configure an IP Access List.
  2. In the TiDB Cloud console, click Connect on the overview page of your cluster, and record the host and port of the connection. You need them in the Dumpling command.
  3. Prepare a SQL user that has the privileges required by Dumpling. This guide uses root as an example. To use a dedicated user instead, grant it the privileges required by Dumpling.

2. Export the snapshot with Dumpling

Run Dumpling on a machine that can connect to your TiDB Cloud Dedicated cluster. Pass the AWS access key in the -o URI with the access-key and secret-access-key parameters:

tiup dumpling \ -h "<host>" \ -P <port> \ -u "<username>" \ -p "<password>" \ --filetype csv \ --csv-output-dialect snowflake \ --escape-backslash=false \ -o "s3://<bucket>/<prefix>/snapshot/?access-key=<access-key-id>&secret-access-key=<secret-access-key>" \ --s3.region "<region>"

Parameter descriptions:

  • --filetype csv with --csv-output-dialect snowflake: export data in CSV format the Snowflake dialect.
  • --escape-backslash=false: disable backslash escaping.
  • Compression is disabled by default.
  • -o and --s3.region: write the exported files to your S3 bucket. The access-key and secret-access-key parameters in the -o URI provide the credentials for the bucket.

After the export completes successfully, the command output includes a JSON summary. Find the SessionParams.tidb_snapshot field in the output, and record its value. This value is the snapshot TSO, which you need when you create the changefeed so that incremental replication continues from the exported snapshot.

Step 3. Create a changefeed for incremental data

In the TiDB Cloud console, you can create a cloud storage changefeed for your TiDB Cloud Dedicated cluster. For the full procedure, see Sink to Cloud Storage. When you configure the changefeed, pay attention to the following settings:

  • S3 URI: use the incremental/ sub-path under the same prefix as the snapshot, for example, s3://<bucket>/<prefix>/incremental/.
  • Bucket Access: select AWS Access Key, and enter the access key from Step 1. Prepare S3 bucket access. Make sure the permissions cover the incremental/ path.
  • Start Replication Position: select Start replication from a specific TSO, and enter the snapshot TSO recorded in Step 2. Export a full snapshot with Dumpling.
  • Data Format: select Canal-JSON, and enable both Enable TiDB Extension and Enable Canal Content Compatibility. These settings produce data in a format that is compatible with the TiDB Cloud Lake integration.

Step 4. Configure TiDB Cloud Lake

In TiDB Cloud Lake, you need to create a data source and an integration to load the data from the S3 bucket.

1. Create a data source

  1. In the TiDB Cloud Lake console, navigate to Data > Data Sources > Create.
  2. Select Service: TiDB.
  3. Select the access key authentication and fill in:
    • Access Key ID and Secret Access Key: the credentials from Prepare S3 bucket access.
    • S3 Bucket Name: the bucket name only (for example, my-datapipeline-bucket, not the full URI).
    • S3 Region: the same region as your TiDB Cloud Dedicated cluster.
  4. SQS Queue URL is optional. If you want to enable event-driven ingestion, set up an SQS queue, configure the S3 bucket notification, and grant the IAM user the required SQS permissions. For more information, see Amazon SQS and S3 IAM Role for TiDB Cloud Lake.

2. Create an integration

  1. In the TiDB Cloud Lake console, navigate to Data > Integration > Create.
  2. Fill in the following fields:
    • Data Source: select the data source created above.
    • Name: a name for this integration task.
    • Sync Mode: select Snapshot + CDC to load the full snapshot first and then continuously apply incremental changes.
    • Table Rules: *.* to sync all exported tables.
    • Changefeed S3 Prefix: <prefix>/incremental/.
    • Dumpling S3 Prefix: <prefix>/snapshot/.
    • Poll Interval: the interval at which TiDB Cloud Lake scans the external stage for new data. The default value is 60 seconds. A shorter interval reduces data latency but increases TiDB Cloud Lake hosting cost.
    • Merge Interval: the interval at which TiDB Cloud Lake merges incremental data into the warehouse. The default value is 30 seconds. A shorter interval reduces data latency but increases TiDB Cloud Lake hosting cost.
    • Warehouse: select the target warehouse.
  3. Click Create.
  4. After creation, the integration is Stopped by default. Click the integration action button > Start to begin data loading.

See also

Was this page helpful?