From 4d94a534d398d83ae3c7df0f4e6face0986bc151 Mon Sep 17 00:00:00 2001 From: lilin90 Date: Tue, 22 Sep 2026 16:52:33 +0800 Subject: [PATCH 1/7] cloud: expand TiDB Cloud Lake overview --- tidb-cloud-lake/lake-overview.md | 61 ++++++++++++++++++++++++++------ 1 file changed, 51 insertions(+), 10 deletions(-) diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index 69a3498ec5820..22d0e3fbb90b3 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -15,16 +15,57 @@ TiDB Cloud Lake supports ANSI SQL, semi-structured data processing, vector searc ## Why {{{ .lake }}}? -{{{ .lake }}} brings analytics, vector, search, and geo workloads together in one cloud-native platform. With storage-compute separation, ANSI SQL support, and managed infrastructure, teams can work on multi-modal data with better flexibility, performance, and cost efficiency. - -| Feature | Description | Learn more | -|---|---|---| -| **Unified Engine** | Analytics, vector, search, and geo share one optimizer and runtime. | [TiDB Cloud Lake Architecture](/tidb-cloud-lake/guides/tidb-cloud-lake-architecture.md) | -| **Unified Data** | Structured, semi-structured, unstructured, and vector data share object storage. | [TiDB Cloud Lake Architecture](/tidb-cloud-lake/guides/tidb-cloud-lake-architecture.md) | -| **Analytics Native** | ANSI SQL, windowing, incremental aggregates, and streaming power BI run on the same platform. | [Worksheets](/tidb-cloud-lake/guides/worksheet.md) | -| **Vector Native** | Embeddings, vector indexes, and semantic retrieval all run in SQL. | [Vector Search](/tidb-cloud-lake/guides/vector-search-guide.md) | -| **Search Native** | Full-text search and inverted indexes power hybrid retrieval. | [Full-Text Index](/tidb-cloud-lake/guides/full-text-index.md) | -| **Geo Native** | Geospatial indexes and functions power map and location services. | [Geo Analytics](/tidb-cloud-lake/guides/geo-analytics.md) | +{{{ .lake }}} brings analytics, data engineering, search, and AI workloads together in one cloud-native platform. It combines independently scalable compute, object-storage-based data management, SQL access, and built-in multimodal capabilities in a managed service. + +### Scale compute independently from storage + +{{{ .lake }}} separates compute from storage. Data is stored in durable, cost-effective object storage, while warehouses provide independently managed compute resources. You can choose a warehouse size for each workload, resize it as demand changes, and use separate warehouses for data loading and query execution. This architecture lets you scale compute without moving or duplicating the underlying data. + +For more information, see [TiDB Cloud Lake Architecture](/tidb-cloud-lake/guides/tidb-cloud-lake-architecture.md) and [Warehouses](/tidb-cloud-lake/guides/warehouse.md). + +### Keep costs transparent and predictable + +The main cost components are straightforward: + +- **Warehouse compute** is billed per second while a warehouse is running. By default, a warehouse automatically suspends after five minutes of inactivity. A suspended warehouse does not consume compute resources. +- **Storage** is billed according to the amount of data stored in object storage. +- **Cloud services** are billed according to API request usage. +- **Data integration services** are billed per second while they are running. + +You can also set a monthly spending limit and monitor usage and billing history. For details, see [TiDB Cloud Lake Pricing & Billing](/tidb-cloud-lake/guides/pricing-billing.md) and [Managing Costs](/tidb-cloud-lake/guides/manage-costs.md). + +### Run high-performance analytics at scale + +{{{ .lake }}} is designed for analytical workloads such as interactive dashboards, ad hoc exploration, large-scale scans, and concurrent queries. Its optimizer uses techniques such as predicate pushdown, join reordering, and scan pruning to reduce unnecessary work. + +For workload-specific optimization, you can use cluster keys to organize related data into adjacent storage blocks, materialized views to persist recurring query results, and specialized indexes and result caching to accelerate common access patterns. The appropriate strategy depends on your query filters, data distribution, and update patterns. + +For more information, see [Performance Optimization](/tidb-cloud-lake/guides/performance-optimization.md), [Cluster Key](/tidb-cloud-lake/guides/cluster-key-performance.md), and [Materialized View](/tidb-cloud-lake/sql/materialized-view.md). + +### Build data pipelines in the platform + +{{{ .lake }}} provides built-in primitives for data ingestion and incremental processing: + +- [Stages](/tidb-cloud-lake/guides/stage-overview.md) provide managed locations and interfaces for loading, querying, and unloading files. +- [Streams](/tidb-cloud-lake/guides/track-and-transform-data-via-streams.md) capture table changes for incremental processing. +- [Tasks](/tidb-cloud-lake/guides/automate-data-loading-with-tasks.md) run SQL on a schedule or when a stream contains new rows. +- [Data integration](/tidb-cloud-lake/guides/data-integration-overview.md) provides a visual interface for importing or continuously synchronizing data from supported external systems. + +Together, these capabilities support batch ingestion, change data capture (CDC), incremental ETL, scheduled transformations, and downstream analytical tables without requiring every pipeline to rely on separate orchestration tooling. + +### Work with structured and semi-structured data + +You can use SQL to store, query, clean, and transform relational and semi-structured data in the same platform. The `VARIANT` data type preserves nested **JSON** structures, while JSON path expressions, virtual columns, and inverted indexes help you retrieve and search frequently accessed fields efficiently. + +This is useful for application events, logs, user activity data, API payloads, and AI agent traces, which often have evolving schemas and can be costly to flatten up front. + +To protect sensitive data, you can apply row access policies to filter rows at query time and masking policies to redact column values, including selected keys in `VARIANT` data. For more information, see [JSON & Search](/tidb-cloud-lake/guides/json-search.md), [Row Access Policy](/tidb-cloud-lake/guides/row-access-policy.md), and [Masking Policy](/tidb-cloud-lake/guides/masking-policy.md). + +### Combine analytics, search, and AI workloads + +Full-text search, vector search, geospatial analysis, and SQL analytics run on the same data platform. You can use inverted indexes for keyword-oriented retrieval, vector indexes for semantic similarity search, and SQL predicates and joins to combine retrieval results with structured business data. + +This unified approach supports use cases such as product and business analytics, application-event analysis, search and recommendations, retrieval-augmented generation (RAG), and AI-agent trace analysis without requiring a separate data copy for each workload. For an end-to-end example, see [Multimodal Data Analytics](/tidb-cloud-lake/guides/multimodal-data-analytics.md). ## Get Started From 56d6952c898185c1f0782ba254b10cb7842ffa0c Mon Sep 17 00:00:00 2001 From: Lilian Lee Date: Tue, 22 Sep 2026 17:14:46 +0800 Subject: [PATCH 2/7] Improve clarity on handling structured and semi-structured data Clarified the description of working with structured and semi-structured data, emphasizing the role of JSON path expressions, virtual columns, and inverted indexes. Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --- tidb-cloud-lake/lake-overview.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index 22d0e3fbb90b3..0db1d761e7410 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -55,7 +55,7 @@ Together, these capabilities support batch ingestion, change data capture (CDC), ### Work with structured and semi-structured data -You can use SQL to store, query, clean, and transform relational and semi-structured data in the same platform. The `VARIANT` data type preserves nested **JSON** structures, while JSON path expressions, virtual columns, and inverted indexes help you retrieve and search frequently accessed fields efficiently. +You can use SQL to store, query, clean, and transform relational and semi-structured data in the same platform. The `VARIANT` data type preserves nested **JSON** structures, while JSON path expressions let you access nested fields, virtual columns accelerate frequently queried paths, and inverted indexes support text search. This is useful for application events, logs, user activity data, API payloads, and AI agent traces, which often have evolving schemas and can be costly to flatten up front. From a9456abf31b3f3f0830b94bd49d274f16ea1719c Mon Sep 17 00:00:00 2001 From: Lilian Lee Date: Tue, 22 Sep 2026 17:15:28 +0800 Subject: [PATCH 3/7] Apply suggestions from code review Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> --- tidb-cloud-lake/lake-overview.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index 0db1d761e7410..c3fb49c14593e 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -32,7 +32,7 @@ The main cost components are straightforward: - **Cloud services** are billed according to API request usage. - **Data integration services** are billed per second while they are running. -You can also set a monthly spending limit and monitor usage and billing history. For details, see [TiDB Cloud Lake Pricing & Billing](/tidb-cloud-lake/guides/pricing-billing.md) and [Managing Costs](/tidb-cloud-lake/guides/manage-costs.md). +Administrators can set a monthly spending limit. You can monitor usage and billing history. For details, see [TiDB Cloud Lake Pricing & Billing](/tidb-cloud-lake/guides/pricing-billing.md) and [Managing Costs](/tidb-cloud-lake/guides/manage-costs.md). ### Run high-performance analytics at scale From 7081c6f1f3ebe3138eb9ce1814156246252eb3ef Mon Sep 17 00:00:00 2001 From: Lilian Lee Date: Tue, 22 Sep 2026 17:18:28 +0800 Subject: [PATCH 4/7] Apply suggestions from code review Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> --- tidb-cloud-lake/lake-overview.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index c3fb49c14593e..bd5e5c6cecd64 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -30,7 +30,7 @@ The main cost components are straightforward: - **Warehouse compute** is billed per second while a warehouse is running. By default, a warehouse automatically suspends after five minutes of inactivity. A suspended warehouse does not consume compute resources. - **Storage** is billed according to the amount of data stored in object storage. - **Cloud services** are billed according to API request usage. -- **Data integration services** are billed per second while they are running. +- **Service hosting** for data integration services is billed per second while they are running. Administrators can set a monthly spending limit. You can monitor usage and billing history. For details, see [TiDB Cloud Lake Pricing & Billing](/tidb-cloud-lake/guides/pricing-billing.md) and [Managing Costs](/tidb-cloud-lake/guides/manage-costs.md). From f7079e45098164b8f29f2a6cedb63a299e354d4d Mon Sep 17 00:00:00 2001 From: lilin90 Date: Tue, 22 Sep 2026 17:37:14 +0800 Subject: [PATCH 5/7] Update wording to address comments --- tidb-cloud-lake/lake-overview.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index bd5e5c6cecd64..3b30b26fddd0b 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -59,11 +59,11 @@ You can use SQL to store, query, clean, and transform relational and semi-struct This is useful for application events, logs, user activity data, API payloads, and AI agent traces, which often have evolving schemas and can be costly to flatten up front. -To protect sensitive data, you can apply row access policies to filter rows at query time and masking policies to redact column values, including selected keys in `VARIANT` data. For more information, see [JSON & Search](/tidb-cloud-lake/guides/json-search.md), [Row Access Policy](/tidb-cloud-lake/guides/row-access-policy.md), and [Masking Policy](/tidb-cloud-lake/guides/masking-policy.md). +To protect sensitive data, you can apply row access policies (an experimental feature) to filter rows at query time and masking policies to redact column values, including selected keys in `VARIANT` data. For more information, see [JSON & Search](/tidb-cloud-lake/guides/json-search.md), [Row Access Policy](/tidb-cloud-lake/guides/row-access-policy.md), and [Masking Policy](/tidb-cloud-lake/guides/masking-policy.md). ### Combine analytics, search, and AI workloads -Full-text search, vector search, geospatial analysis, and SQL analytics run on the same data platform. You can use inverted indexes for keyword-oriented retrieval, vector indexes for semantic similarity search, and SQL predicates and joins to combine retrieval results with structured business data. +Full-text search, [vector search](/tidb-cloud-lake/guides/vector-search-guide.md), [geospatial analysis](/tidb-cloud-lake/guides/geo-analytics.md), and SQL analytics run on the same data platform. You can use [full-text indexes](/tidb-cloud-lake/guides/full-text-index.md) (inverted indexes) for keyword-oriented retrieval, vector indexes for semantic similarity search, and SQL predicates and joins to combine retrieval results with structured business data. This unified approach supports use cases such as product and business analytics, application-event analysis, search and recommendations, retrieval-augmented generation (RAG), and AI-agent trace analysis without requiring a separate data copy for each workload. For an end-to-end example, see [Multimodal Data Analytics](/tidb-cloud-lake/guides/multimodal-data-analytics.md). From 50f0a698c5e10131d3c2fb04dae946fe81de0f8d Mon Sep 17 00:00:00 2001 From: lilin90 Date: Tue, 22 Sep 2026 17:41:41 +0800 Subject: [PATCH 6/7] Add Why Lake to TOC --- TOC-tidb-cloud-lake.md | 1 + tidb-cloud-lake/lake-overview.md | 2 +- 2 files changed, 2 insertions(+), 1 deletion(-) diff --git a/TOC-tidb-cloud-lake.md b/TOC-tidb-cloud-lake.md index 90d850110c7f0..604a192be7ba5 100644 --- a/TOC-tidb-cloud-lake.md +++ b/TOC-tidb-cloud-lake.md @@ -6,6 +6,7 @@ ## Get Started - [Overview](/tidb-cloud-lake/lake-overview.md) +- [Why TiDB Cloud Lake](/tidb-cloud-lake/lake-overview.md#why-lake) - [Quick Start](/tidb-cloud-lake/lake-quick-start.md) ## Guides diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index 3b30b26fddd0b..13e539e67590e 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -13,7 +13,7 @@ TiDB Cloud Lake supports ANSI SQL, semi-structured data processing, vector searc > > TiDB Cloud Lake is currently in **public preview**. Feature availability and service limits might change as we continue to improve the product. -## Why {{{ .lake }}}? +## Why {{{ .lake }}}? {#why-lake} {{{ .lake }}} brings analytics, data engineering, search, and AI workloads together in one cloud-native platform. It combines independently scalable compute, object-storage-based data management, SQL access, and built-in multimodal capabilities in a managed service. From 58463b9eddb9c51f17bec1a170d846a49b46eb9c Mon Sep 17 00:00:00 2001 From: lilin90 Date: Thu, 24 Sep 2026 10:35:25 +0800 Subject: [PATCH 7/7] Add TiDB Cloud Lake continuous replication docs Added a new overview section describing TiDB Cloud Data Pipeline replication from TiDB Cloud to TiDB Cloud Lake. It explains the full snapshot plus incremental CDC flow, common use cases, and the external-stage dependency and plan availability restrictions. --- tidb-cloud-lake/lake-overview.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/tidb-cloud-lake/lake-overview.md b/tidb-cloud-lake/lake-overview.md index 13e539e67590e..908681730b259 100644 --- a/tidb-cloud-lake/lake-overview.md +++ b/tidb-cloud-lake/lake-overview.md @@ -53,6 +53,18 @@ For more information, see [Performance Optimization](/tidb-cloud-lake/guides/per Together, these capabilities support batch ingestion, change data capture (CDC), incremental ETL, scheduled transformations, and downstream analytical tables without requiring every pipeline to rely on separate orchestration tooling. +### Replicate TiDB Cloud data continuously + +TiDB Cloud Data Pipeline replicates full and incremental data from a TiDB Cloud instance to {{{ .lake }}} without requiring a third-party ETL tool. It first exports a full snapshot of the selected tables, and then continuously replicates row changes, including inserts, updates, and deletes, so that analytical data stays current. + +You can use this integration to: + +- Replicate operational TiDB data to an analytical warehouse. +- Keep dashboards and reports up to date. +- Make continuously updated data available to analytics, search, vector, and AI workloads in {{{ .lake }}}. + +Data Pipeline uses TiCDC for incremental replication and an external stage to transfer data between TiDB Cloud and {{{ .lake }}}. Availability and restrictions depend on your TiDB Cloud plan. For details, see [Data Pipeline to TiDB Cloud Lake](/tidb-cloud/data-pipeline/data-pipeline-overview.md). + ### Work with structured and semi-structured data You can use SQL to store, query, clean, and transform relational and semi-structured data in the same platform. The `VARIANT` data type preserves nested **JSON** structures, while JSON path expressions let you access nested fields, virtual columns accelerate frequently queried paths, and inverted indexes support text search.