---
title: Parquet Export
slug: reference/parquet-export
docTags: 
createdAt: 2025-09-09T09:39:41.042Z
---

[Apache Parquet](https://parquet.apache.org/) is an open source, column-oriented data file format designed for efficient data storage and retrieval. It provides efficient data compression and encoding schemes with enhanced performance to handle complex data in bulk.

To create a parquet export, navigate to `Configuration > Parquet export`. The screen shows an overview of the created export schedules, and allows for the creation of a one-off export or the creation of an export schedule. The export is not filtered on quality: points of every status are exported.

:::hint{type="warning"}
Exporting parquet files can take a significant amount of resources.&#x20;
Plan the scheduled exports accordingly.

Exporting (many) measurements over great timeranges can take up a lot of storage.
Choose the export destination accordingly.
:::

## One-off export

To create a one-off export task that only runs once, click the `Parquet export` button in the upper right corner. A modal will open where the export task can be configured.

1. Select a time range for which to export data
2. Choose an interval to split up the range into multiple files
3. Choose a directory where the exported parquet file(s) will be saved
4. Choose whether to allow overwriting existing files or not
5. Choose to include the full asset tree or not.&#x20;
   When enabling this option, the relational structure of the asset tree is exported as well. This relational structure consists of a file containing the assets' data, a file containing the asset properties, and a file containing measurements data.&#x20;
6. Select the measurement(s) that need to be exported either by choosing a set of labels, or by selecting from the asset tree.
   When selecting from the asset tree, the selected asset structure is automatically also exported.

## Export on a schedule

To create a recurring export, press the `Create Task scheduler` button in the top right corner.

1. Fill in the task details
   1. Choose a name and description for the scheduler
   2. Write an RRULE to define when and how often the task should be scheduled.
2. Fill in the export details
   1. Fill in the start offset.
      The start offset determines the start of the interval for the exported data, relative to the RRule’s trigger event.
   2. Fill in the period.
      The period defines how long the time period should be that's included in the export, starting from the start time.
   3. Fill in the end time offset.
      This value defines until what time the export should include data, relative to the RRule's trigger event.
   4. Choose an interval to split up the range into multiple files.
3. Choose a directory where the exported parquet file(s) will be saved
4. Choose whether to allow overwriting existing files or not
5. Choose to include the full asset tree or not.&#x20;
   When enabling this option, the relational structure of the asset tree is exported as well. This relational structure consists of a file containing the assets' data, a file containing the asset properties, and a file containing measurements data.&#x20;
6. Select the measurement(s) that need to be exported either by choosing a set of labels, or by selecting from the asset tree.
   When selecting from the asset tree, the selected asset structure is automatically also exported.
7. Click save

:::hint{type="info"}
An RRULE is way to define a recurrence set. The rule defines a pattern to generate a series of timestamps for events. Their syntax is defined by [section 3.8.5.3 of RFC 5545](https://datatracker.ietf.org/doc/html/rfc5545#section-3.8.5.3).
Use the [RRULE tool](https://icalendar.org/rrule-tool.html) for help in creating RRules
:::

### Example

The following is an example configuration of an exporter.

The used RRULE is `RRULE:FREQ=HOURLY;INTERVAL=1;BYMINUTE=0;BYSECOND=0`, which triggers every hour, on the hour.
By setting the start offset to `-1h10m` and the period to `1h`, the scheduler
will export an hour’s worth of data, starting from 10 minutes before the previous hour until 10 minutes before the current hour.
Setting the stop offset to `-10m` would be equivalent to setting the period to `1h` in this example.

| trigger at | export start time | export stop time |
| ---------- | ----------------- | ---------------- |
| 12:00      | 10:50             | 11:50            |
| 13:00      | 11:50             | 12:50            |
| 14:00      | 12:50             | 13:50            |

## Advanced Settings

Advanced settings can be configured for both the one-off export and the scheduled export. By default, these settings contain sane defaults, but they can be overwritten if the use case asks for it. &#x20;

### Batch Size

**Default:** 100.000
**Minimum:** 10.000****
**Description:** The number of records to process in each batch during export. A larger batch size can improve performance and compression ratio but may use more memory. A parquet file can contain multiple batches.

### Max rows per file

**Default:** 0
**Description:** The maximum number of rows allowed in each Parquet file. If set to 0, there is no limit and all data of a measurement will be written to a single file.

### Parquet V2 data pages supported

**Default:** true
**Description:** Enable support for Parquet V2 data pages, which can improve performance and reduce file size for certain workloads. Disable it when not supported by downstream pipelines.

### Compression codec

**Default:** zstd
**Options:** none, snappy, gzip, brotli, **zstd**, lz4raw
**Description:** The compression codec to use for the Parquet files. Different codecs offer various trade-offs between compression ratio and speed.&#x20;

### Compression level

**Default:** default
**Options:&#x20;**&#x20;best speed, **default,** better compression, best compression
**Description:** The level of compression to apply when using the selected codec. Higher levels typically result in better compression but may take longer to process and require more system resources.

## Export progress

Creating a parquet export can take some time, depending on the amount of data that needs to be exported. Task progress can be viewed in the [Tasks](docId\:OUywPxc3NeEVmwRsZfhlE) screen.
