> For the complete documentation index, see [llms.txt](https://docs.starrocks.io/llms.txt). This page is also available as Markdown at its `.md` URL.

# Loading options

Data loading is the process of cleansing and transforming raw data from various data sources based on your business requirements and loading the resulting data into StarRocks to facilitate analysis.

StarRocks provides a variety of options for data loading:

* Loading methods: Insert, Stream Load, Broker Load, Pipe, Routine Load, and Spark Load
* Ecosystem tools: StarRocks Connector for Apache Kafka® (Kafka connector for short), StarRocks Connector for Apache Spark™ (Spark connector for short), StarRocks Connector for Apache Flink® (Flink connector for short), and other tools such as SMT, DataX, CloudCanal, and Kettle Connector
* API: Stream Load transaction interface

These options each have its own advantages and support its own set of data source systems to pull from.

This topic provides an overview of these options, along with comparisons between them to help you determine the loading option of your choice based on your data source, business scenario, data volume, data file format, and loading frequency.

## Introduction to loading options[​](#introduction-to-loading-options "Direct link to Introduction to loading options")

This section mainly describes the characteristics and business scenarios of the loading options available in StarRocks.

![Loading options overview](/assets/images/loading_intro_overview-01dc28ae3c1adabb8349da9fbfa0140f.png)

note

In the following sections, "batch" or "batch loading" refers to the loading of a large amount of data from a specified source all at a time into StarRocks, whereas "stream" or "streaming" refers to the continuous loading of data in real time.

## Loading methods[​](#loading-methods "Direct link to Loading methods")

### [Insert](https://docs.starrocks.io/docs/loading/InsertInto.md)[​](#insert "Direct link to insert")

**Business scenario:**

* INSERT INTO VALUES: Append to an internal table with small amounts of data.

* INSERT INTO SELECT:

  * INSERT INTO SELECT FROM `<table_name>`: Append to a table with the result of a query on an internal or external table.

  * INSERT INTO SELECT FROM FILES(): Append to a table with the result of a query on data files in remote storage.

    note

    For AWS S3, this feature is supported from v3.1 onwards. For HDFS, Microsoft Azure Storage, Google GCS, and S3-compatible storage (such as MinIO), this feature is supported from v3.2 onwards.

**File format:**

* INSERT INTO VALUES: SQL

* INSERT INTO SELECT:

  * INSERT INTO SELECT FROM `<table_name>`: StarRocks tables
  * INSERT INTO SELECT FROM FILES(): Parquet and ORC

**Data volume:** Not fixed (The data volume varies based on the memory size.)

### [Stream Load](https://docs.starrocks.io/docs/loading/StreamLoad.md)[​](#stream-load "Direct link to stream-load")

**Business scenario:** Batch load data from a local file system.

**File format:** CSV and JSON

**Data volume:** 10 GB or less

### [Broker Load](https://docs.starrocks.io/docs/sql-reference/sql-statements/loading_unloading/BROKER_LOAD.md)[​](#broker-load "Direct link to broker-load")

**Business scenario:**

* Batch load data from HDFS or cloud storage like AWS S3, Microsoft Azure Storage, Google GCS, and S3-compatible storage (such as MinIO).
* Batch load data from a local file system or NAS.

**File format:** CSV, Parquet, ORC, and JSON (supported since v3.2.3)

**Data volume:** Dozens of GB to hundreds of GB

### [Pipe](https://docs.starrocks.io/docs/sql-reference/sql-statements/loading_unloading/pipe/CREATE_PIPE.md)[​](#pipe "Direct link to pipe")

**Business scenario:** Batch load or stream data from HDFS or AWS S3.

note

This loading method is supported from v3.2 onwards.

**File format:** Parquet and ORC

**Data volume:** 100 GB to 1 TB or more

### [Routine Load](https://docs.starrocks.io/docs/sql-reference/sql-statements/loading_unloading/routine_load/CREATE_ROUTINE_LOAD.md)[​](#routine-load "Direct link to routine-load")

**Business scenario:** Stream data from Kafka.

**File format:** CSV, JSON, and Avro (supported since v3.0.1)

**Data volume:** MBs to GBs of data as mini-batches

### [Spark Load](https://docs.starrocks.io/docs/sql-reference/sql-statements/loading_unloading/SPARK_LOAD.md)[​](#spark-load "Direct link to spark-load")

**Business scenario:** Batch load data of Apache Hive™ tables stored in HDFS by using Spark clusters.

**File format:** CSV, Parquet (supported since v2.0), and ORC (supported since v2.0)

**Data volume:** Dozens of GB to TBs

## Ecosystem tools[​](#ecosystem-tools "Direct link to Ecosystem tools")

### [Kafka connector](https://docs.starrocks.io/docs/loading/kafka/Kafka-connector-starrocks.md)[​](#kafka-connector "Direct link to kafka-connector")

**Business scenario:** Stream data from Kafka.

### [Spark connector](https://docs.starrocks.io/docs/loading/spark/Spark-connector-starrocks.md)[​](#spark-connector "Direct link to spark-connector")

**Business scenario:** Batch load data from Spark.

### [Flink connector](https://docs.starrocks.io/docs/loading/Flink-connector-starrocks.md)[​](#flink-connector "Direct link to flink-connector")

**Business scenario:** Stream data from Flink.

### [SMT](https://docs.starrocks.io/docs/integrations/loading_tools/SMT.md)[​](#smt "Direct link to smt")

**Business scenario:** Load data from data sources such as MySQL, PostgreSQL, SQL Server, Oracle, Hive, ClickHouse, and TiDB through Flink.

### [DataX](https://docs.starrocks.io/docs/integrations/loading_tools/DataX-starrocks-writer.md)[​](#datax "Direct link to datax")

**Business scenario:** Synchronize data between various heterogeneous data sources, including relational databases (for example, MySQL and Oracle), HDFS, and Hive.

### [CloudCanal](https://docs.starrocks.io/docs/integrations/loading_tools/CloudCanal.md)[​](#cloudcanal "Direct link to cloudcanal")

**Business scenario:** Migrate or synchronize data from source databases (for example, MySQL, Oracle, and PostgreSQL) to StarRocks.

### [Kettle Connector](https://github.com/StarRocks/starrocks-connector-for-kettle)[​](#kettle-connector "Direct link to kettle-connector")

**Business scenario:** Integrate with Kettle. By combining Kettle's robust data processing and transformation capabilities with StarRocks's high-performance data storage and analytical abilities, more flexible and efficient data processing workflows can be achieved.

## API[​](#api "Direct link to API")

### [Stream Load transaction interface](https://docs.starrocks.io/docs/loading/Stream_Load_transaction_interface.md)[​](#stream-load-transaction-interface "Direct link to stream-load-transaction-interface")

**Business scenario:** Implement two-phase commit (2PC) for transactions that are run to load data from external systems such as Flink and Kafka, while improving the performance of highly concurrent stream loads. This feature is supported from v2.4 onwards.

**File format:** CSV and JSON

**Data volume:** 10 GB or less

## Choice of loading options[​](#choice-of-loading-options "Direct link to Choice of loading options")

This section lists the loading options available for common data sources, helping you choose the option that best suits your situation.

### Object storage[​](#object-storage "Direct link to Object storage")

| **Data source**                       | **Available loading options**                                                                                                                                                                                                               |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| AWS S3                                | - (Batch) INSERT INTO SELECT FROM FILES() (supported since v3.1)<br />- (Batch) Broker Load<br />- (Batch or streaming) Pipe (supported since v3.2)See [Load data from AWS S3](https://docs.starrocks.io/docs/loading/objectstorage/s3.md). |
| Microsoft Azure Storage               | - (Batch) INSERT INTO SELECT FROM FILES() (supported since v3.2)<br />- (Batch) Broker LoadSee [Load data from Microsoft Azure Storage](https://docs.starrocks.io/docs/loading/objectstorage/azure.md).                                     |
| Google GCS                            | - (Batch) INSERT INTO SELECT FROM FILES() (supported since v3.2)<br />- (Batch) Broker LoadSee [Load data from GCS](https://docs.starrocks.io/docs/loading/objectstorage/gcs.md).                                                           |
| S3-compatible storage (such as MinIO) | - (Batch) INSERT INTO SELECT FROM FILES() (supported since v3.2)<br />- (Batch) Broker LoadSee [Load data from MinIO](https://docs.starrocks.io/docs/loading/objectstorage/minio.md).                                                       |

### Local file system (including NAS)[​](#local-file-system-including-nas "Direct link to Local file system (including NAS)")

| **Data source**                   | **Available loading options**                                                                                                                   |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| Local file system (including NAS) | - (Batch) Stream Load<br />- (Batch) Broker LoadSee [Load data from a local file system](https://docs.starrocks.io/docs/loading/StreamLoad.md). |

### HDFS[​](#hdfs "Direct link to HDFS")

| **Data source** | **Available loading options**                                                                                                                                                                                                      |
| --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| HDFS            | - (Batch) INSERT INTO SELECT FROM FILES() (supported since v3.2)<br />- (Batch) Broker Load<br />- (Batch or streaming) Pipe (supported since v3.2)See [Load data from HDFS](https://docs.starrocks.io/docs/loading/hdfs_load.md). |

### Flink, Kafka, and Spark[​](#flink-kafka-and-spark "Direct link to Flink, Kafka, and Spark")

| **Data source** | **Available loading options**                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| --------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Apache Flink®   | - [Flink connector](https://docs.starrocks.io/docs/loading/Flink-connector-starrocks.md)<br />- [Stream Load transaction interface](https://docs.starrocks.io/docs/loading/Stream_Load_transaction_interface.md)                                                                                                                                                                                                                                                                                                                                                                                                                        |
| Apache Kafka®   | - (Streaming) [Kafka connector](https://docs.starrocks.io/docs/loading/kafka/Kafka-connector-starrocks.md)<br />- (Streaming) [Routine Load](https://docs.starrocks.io/docs/loading/kafka/RoutineLoad.md)<br />- [Stream Load transaction interface](https://docs.starrocks.io/docs/loading/Stream_Load_transaction_interface.md) **NOTE**<br />If the source data requires multi-table joins and extract, transform and load (ETL) operations, you can use Flink to read and pre-process the data and then use [Flink connector](https://docs.starrocks.io/docs/loading/Flink-connector-starrocks.md) to load the data into StarRocks. |
| Apache Spark™   | - [Spark connector](https://docs.starrocks.io/docs/loading/spark/Spark-connector-starrocks.md)<br />- [Spark Load](https://docs.starrocks.io/docs/loading/spark/SparkLoad.md)                                                                                                                                                                                                                                                                                                                                                                                                                                                           |

### Data lakes[​](#data-lakes "Direct link to Data lakes")

| **Data source** | **Available loading options**                                                                                                                                                                                                                                                                                                                                            |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Apache Hive™    | - (Batch) Create a [Hive catalog](https://docs.starrocks.io/docs/data_source/catalog/hive_catalog.md) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).<br />- (Batch) [Spark Load](https://docs.starrocks.io/docs/loading/SparkLoad/). |
| Apache Iceberg  | (Batch) Create an [Iceberg catalog](https://docs.starrocks.io/docs/data_source/catalog/iceberg.md) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).                                                                                    |
| Apache Hudi     | (Batch) Create a [Hudi catalog](https://docs.starrocks.io/docs/data_source/catalog/hudi_catalog.md) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).                                                                                   |
| Delta Lake      | (Batch) Create a [Delta Lake catalog](https://docs.starrocks.io/docs/data_source/catalog/deltalake_catalog.md) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).                                                                        |
| Elasticsearch   | (Batch) Create an [Elasticsearch catalog](https://docs.starrocks.io/docs/data_source/catalog/elasticsearch_catalog.md) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).                                                                |
| Apache Paimon   | (Batch) Create a [Paimon catalog](https://docs.starrocks.io/docs/data_source/catalog/paimon_catalog.md) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).                                                                               |

Note that StarRocks provides [unified catalogs](https://docs.starrocks.io/docs/data_source/catalog/unified_catalog/) from v3.2 onwards to help you handle tables from Hive, Iceberg, Hudi, and Delta Lake data sources as a unified data source without ingestion.

### Internal and external databases[​](#internal-and-external-databases "Direct link to Internal and external databases")

| **Data source**                                                              | **Available loading options**                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| StarRocks                                                                    | (Batch) Create a [StarRocks external table](https://docs.starrocks.io/docs/data_source/External_table.md#starrocks-external-table) and then use [INSERT INTO VALUES](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-via-insert-into-values) to insert a few data records or [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table) to insert the data of a table.<br />**NOTE**<br />StarRocks external tables only support data writes. They do not support data reads. |
| MySQL                                                                        | - (Batch) Create a [JDBC catalog](https://docs.starrocks.io/docs/data_source/catalog/jdbc_catalog.md) (recommended) or a [MySQL external table](https://docs.starrocks.io/docs/data_source/External_table.md#deprecated-mysql-external-table) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).<br />- (Streaming) Use [SMT, Flink CDC connector, Flink, and Flink connector](https://docs.starrocks.io/docs/loading/Flink_cdc_load.md).                                |
| Other databases such as Oracle, PostgreSQL, SQL Server, ClickHouse, and TiDB | - (Batch) Create a [JDBC catalog](https://docs.starrocks.io/docs/data_source/catalog/jdbc_catalog.md) (recommended) or a [JDBC external table](https://docs.starrocks.io/docs/data_source/External_table.md#external-table-for-a-jdbc-compatible-database) and then use [INSERT INTO SELECT FROM `<table_name>`](https://docs.starrocks.io/docs/loading/InsertInto.md#insert-data-from-an-internal-or-external-table-into-an-internal-table).<br />- (Streaming) Use [SMT, Flink CDC connector, Flink, and Flink connector](https://docs.starrocks.io/docs/loading/loading_tools.md).                    |
