Sink to TiDB Cloud Lake
This guide walks you through the end-to-end setup of a data pipeline from a TiDB Cloud Essential instance to TiDB Cloud Lake: export a full snapshot to Amazon S3, create a changefeed to continuously write incremental changes to the same S3 location, and configure TiDB Cloud Lake to load both the snapshot and incremental data.
Restrictions
- The TiDB Cloud Lake warehouse must be in the same region as your Essential instance.
- Only tables with a primary key can be replicated incrementally.
- The pipeline requires manual setup and maintenance of AWS IAM resources and credentials, a changefeed, and TiDB Cloud Lake integrations.
- For details on DDL, DML, and column type support, see Data Pipeline SQL Compatibility for TiDB Cloud Lake.
Prerequisites
Before you begin, make sure that you have:
- A TiDB Cloud Essential instance deployed in a specific region.
- Access to the TiDB Cloud API for this organization. You can create an API key on the TiDB Cloud API Keys page in the TiDB Cloud console. Make sure to save the Public Key and Private Key because they are required for all API calls in this guide.
- An Amazon S3 bucket in the same region (for example,
s3://my-datapipeline-bucket) as your TiDB Cloud Essential instance. - A TiDB Cloud Lake warehouse in the same region as your TiDB Cloud Essential instance.
Step 1. Prepare S3 bucket access
The data pipeline components (export, changefeed, and TiDB Cloud Lake) all require access to the same S3 bucket. Choose one of the following methods to access the S3 bucket:
- Role ARN (for TiDB Cloud Essential instances hosted on AWS): a single IAM role shared by all three components. This method avoids long-lived credentials and provides stronger security.
- Access Key: simpler to set up and required when Role ARN is unavailable; requires manual credential management and rotation.
Method 1: Use a Role ARN
In this method, you first use the CloudFormation stack provided by the Export feature in the TiDB Cloud console to create an IAM role configured for Export. Then, you extend the same role's trust policy and permissions so that the changefeed and TiDB Cloud Lake can also use it to access the S3 bucket.
1. Create the role with Export CloudFormation
- In the TiDB Cloud console, navigate to the overview page for your TiDB Cloud Essential instance.
- Click Data > Import in the left navigation pane, and then click Export Data to in the upper-right corner.
- Choose Amazon S3. When configuring the S3 destination with Role ARN authentication, TiDB Cloud provides a CloudFormation link. Use it to create the IAM role.
After the stack is created, record the Role ARN from the stack Outputs (for example, arn:aws:iam::<account-id>:role/<role-name>).
2. Consolidate trust relationships
The IAM role created in the previous step is initially configured for exporting data from TiDB Cloud Essential to S3. Because the same role is also used by the changefeed and TiDB Cloud Lake to access the S3 bucket, update its trust policy to allow these components to assume the role as well.
Collect the following values required for the additional trust relationships, and then update the role's trust policy. In the AWS Console, navigate to the role created in 1. Create the role with Export CloudFormation, go to the Trust relationships tab, and click Edit trust policy.
Export: before replacing the trust policy, record the existing Export AWS account ID and external ID so that you can preserve this trust relationship in the consolidated policy.
Changefeed: call the TiDB Cloud API to get the required values:
curl -L -X GET 'https://serverless.tidbapi.com/v1beta1/clusters/{clusterId}/changefeeds:getCloudStorageAuthConfig' \ -u '<Public Key>:<Private Key>' --digestRecord
tidbCloudAccountIdandtidbCloudAccountExternalIdfrom the response.TiDB Cloud Lake: in the TiDB Cloud Lake console, navigate to Data > Data Sources > Create. In the Basic Info section, select Service: TiDB, and then under Trust Cloud Platform roles, record the following values:
- Lake Setup & Validation Role ARN
- Lake Data Loading Role ARN
- Lake External ID
Replace the role's trust policy with the consolidated policy below:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowExportAssumeRole",
"Effect": "Allow",
"Principal": {
"AWS": "arn:aws:iam::<Export_TiDB_Cloud_Account_ID>:root"
},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {
"sts:ExternalId": "<Export_External_ID>"
}
}
},
{
"Sid": "AllowChangefeedAssumeRole",
"Effect": "Allow",
"Principal": {
"AWS": "arn:aws:iam::<Changefeed_TiDB_Cloud_Account_ID>:root"
},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {
"sts:ExternalId": "<Changefeed_External_ID>"
}
}
},
{
"Sid": "AllowLakeSetupAssumeRole",
"Effect": "Allow",
"Principal": {
"AWS": "<Lake_Setup_and_Validation_Role_ARN>"
},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {
"sts:ExternalId": "<Lake_External_ID>"
}
}
},
{
"Sid": "AllowLakeLoadAssumeRole",
"Effect": "Allow",
"Principal": {
"AWS": "<Lake_Data_Loading_Role_ARN>"
},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {
"sts:ExternalId": "<Lake_External_ID>"
}
}
}
]
}
3. Expand permissions to cover the full pipeline prefix
The CloudFormation-created permissions policy is scoped to the snapshot export path only. The changefeed writes to {prefix}/incremental/, and TiDB Cloud Lake reads from both {prefix}/snapshot/ and {prefix}/incremental/, so the policy must cover the parent prefix.
In the AWS Console, navigate to the role created in Step 1, under the Permissions tab, click the policy name, and edit the policy to replace the resource scope.
Replace the role's inline permissions policy with:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "S3BucketAccess",
"Effect": "Allow",
"Action": [
"s3:ListBucket",
"s3:GetBucketLocation"
],
"Resource": "arn:aws:s3:::<your-bucket-name>"
},
{
"Sid": "S3ObjectAccess",
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject",
"s3:GetObjectVersion",
"s3:DeleteObjectVersion"
],
"Resource": "arn:aws:s3:::<your-bucket-name>/<your-prefix>/*"
}
]
}
Method 2: Use an Access Key
If you prefer Access Key authentication, create an IAM user with the following permissions and provide its credentials when configuring Export, changefeed, and TiDB Cloud Lake.
Permissions policy:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "S3BucketAccess",
"Effect": "Allow",
"Action": [
"s3:ListBucket",
"s3:GetBucketLocation"
],
"Resource": "arn:aws:s3:::<your-bucket-name>"
},
{
"Sid": "S3ObjectAccess",
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject",
"s3:GetObjectVersion",
"s3:DeleteObjectVersion"
],
"Resource": "arn:aws:s3:::<your-bucket-name>/<your-prefix>/*"
}
]
}
Record the Access Key ID and Secret Access Key for use in later steps.
Step 2. Export full snapshot to Amazon S3
In the TiDB Cloud console, navigate to Data > Import, click Export Data to in the upper-right corner, and choose Amazon S3 to create a new export task.
Configuration:
- Selected Data: choose the databases and tables to export.
- Data Format:
CSV- Click Edit CSV Configuration and set:
- Dialect:
Snowflake - Escape backslash:
false
- Dialect:
- Click Edit CSV Configuration and set:
- Compression:
None - Amazon S3 Settings:
- Bucket URI:
s3://<bucket>/<prefix>/snapshot/(use the recommended snapshot sub-path) - Role ARN or Access Key: use the credentials from Prepare Bucket Access.
- Bucket URI:
After the export task completes, open the task details and record the Snapshot TSO value. This TSO is required when creating the changefeed.
Step 3. Create a changefeed for incremental data
Currently, TiDB Cloud Essential does not support creating a cloud-storage sink via TiDB Cloud console, so you need to use the TiDB Cloud API.
Call the changefeed creation API with the following required fields:
Request example:
curl -L -X POST 'https://serverless.tidbapi.com/v1beta1/clusters/{clusterId}/changefeeds' \
-H 'Content-Type: application/json' \
-u '<Public Key>:<Private Key>' --digest \
-d '{
"displayName": "<changefeed-name>",
"sink": {
"type": "CLOUD_STORAGE",
"cloudStorage": {
"storage": {
"type": "S3",
"s3": {
"uri": "s3://<bucket>/<prefix>/incremental/",
"authType": "ROLE_ARN",
"roleArn": "<role-arn>"
}
},
"dataFormat": {
"protocol": "CANAL_JSON",
"contentCompatible": true,
"canalJsonConfig": {
"enableTidbExtension": true
}
}
}
},
"filter": {
"mode": "IGNORE_NOT_SUPPORT_TABLE",
"filterRule": ["*.*"]
},
"startPosition": {
"mode": "FROM_TSO",
"tso": "<snapshot_tso>"
},
"rcu": 2
}'
Step 4. Configure TiDB Cloud Lake
In TiDB Cloud Lake, you need to create a data source and an integration to load the data from the S3 bucket.
1. Create a data source
- In the TiDB Cloud Lake console, navigate to Data > Data Sources > Create.
- Select Service: TiDB.
- Choose Role ARN or Access Key authentication and fill in:
- Role ARN: the ARN from 1. Create the role with Export CloudFormation, or Access Key ID / Secret Access Key from Method 2: Use an Access Key.
- S3 Bucket Name: the bucket name only (for example,
my-datapipeline-bucket, not the full URI). - S3 Region: the same region as your Essential instance.
- SQS Queue URL is optional. If you want to enable event-driven ingestion, set up an SQS queue and configure the S3 bucket notification first. For details, see Amazon SQS and S3 IAM Role for TiDB Cloud Lake.
- Under Trust Cloud Platform roles, verify the TiDB Cloud Lake platform roles and external ID match the values you added to the consolidated trust policy in 2. Consolidate trust relationships.
2. Create an integration
- In the TiDB Cloud Lake console, navigate to Data > Integration > Create.
- Fill in the following fields:
- Data Source: select the data source created above.
- Name: a name for this integration task.
- Sync Mode: select
Snapshot + CDCto load the full snapshot first and then continuously apply incremental changes. - Table Rules:
*.*to sync all exported tables. - Changefeed S3 Prefix:
<prefix>/incremental/. - Dumpling S3 Prefix:
<prefix>/snapshot/. - Poll Interval: the interval at which TiDB Cloud Lake scans the external stage for new data. The default value is 60 seconds. A shorter interval reduces data latency but increases TiDB Cloud Lake hosting cost.
- Merge Interval: the interval at which TiDB Cloud Lake merges incremental data into the warehouse. The default value is 30 seconds. A shorter interval reduces data latency but increases TiDB Cloud Lake hosting cost.
- Warehouse: select the target warehouse.
- Click Create.
- After creation, the integration is Stopped by default. Click the integration action button > Start to begin data loading.
See also
- For details on DDL, DML, and column type support, see Data Pipeline SQL Compatibility for TiDB Cloud Lake.