Parquet file output transform Icon Parquet file output

Description

The Parquet file output transform writes data into the Apache Parquet file format.

For more information on this see: Apache Parquet.

Options

The dialog is split across four tabs so it fits a 1080p screen at 100% zoom: File, Options, Partitioning, and Fields.

Notes:

  • The date optionally referenced in the output file name(s) will be the start of the pipeline execution.

  • Rows are buffered in memory until a row group is full (see the Row group size option), then encoded, compressed and written. Memory use is therefore bounded by the row group size (times the number of copies and, when partitioning, the number of open partitions), not by the split size.

File

Parquet file output File tab

Option Description

Transform name

Name of the transform this name has to be unique in a single pipeline.

Base file name

Specify the base filename. This is composed of where you want to write the Parquet file to as well as the start of the filename. Examples:

Write to Amazon AWS S3 : s3://my-bucket-name/transactions

Write to a local folder : /my/folder/customer-data

Extension

This is the extension of the file. Usually this is simply parquet

Include date?

Check this box if you want to include the date in the filename with mask yyyMMdd

Include time?

Check this box if you want to include the time in the filename with mask HHmmss

Include date-time-format?

Check this box if you want to include a specific custom date-time format in the filename

Include transform copy number?

Enable this option if you run this transform in multiple copies to not have multiple threads write to the same file. The copy number is formatted with mask 00

Split into parts and include number?

Enable this option if you want to split the output into multiple parts. Specify a split size larger than 0 and this is then the number of rows per file. The file part (split) number will be included in the filename to make sure that the same file is not being overwritten. The split number is formatted with mask 0000

Create parent folders?

Create missing parent folders of the output path.

Include compression codec before extension?

When enabled (the default for new transforms), the compression codec is placed before the file extension (for example file.snappy.parquet), which matches the naming used by Spark and other Parquet tools. When disabled, the compression codec is appended after the extension (for example file.parquet.snappy). Existing pipelines that do not have this option set keep the previous behavior.

Options

Parquet file output Options tab

Option Description

Compression codec

Here you can indicate which compression codec you want to use. The default is UNCOMPRESSED.

Version

Choose the protocol version of Parquet (1.0 or 2.0)

Row group size (bytes)

The size of a row group in bytes, not rows. A row group is the unit of work for a reader: the transform buffers rows in memory until the row group is full, then encodes, compresses and writes it. The default is 268435456 (256 MB); Parquet itself defaults to 128 MB, and the Parquet project recommends 512 MB to 1 GB for large files. See Row group size and the file footer below before lowering this value.

Data page size (bytes)

The size of a data page in bytes, the unit of encoding and compression inside a column chunk. The default is 8192 (8 kB), which the Parquet project recommends for good read performance; Parquet itself defaults to 1048576 (1 MB).

Dictionary page size (bytes)

The maximum size of a dictionary page in bytes. When the dictionary of a column grows past this size the writer falls back to plain encoding for that column chunk. The default is 1048576 (1 MB).

The three size options accept human-friendly numbers such as 256m (256,000,000), 1.5m or 1e6, and variables.

Row group size and the file footer

Every row group adds an entry per column to the file footer: offsets, sizes, encodings and min/max statistics. The footer therefore grows with the number of row groups times the number of columns, and with the size of the string values in the statistics. Readers such as Dremio and Apache Drill refuse files whose footer is larger than 16 MB (Max supported footer size is 16777216).

Small row groups are the usual cause of an oversized footer. Hop 2.3 and 2.4 saved a row group size of 20000 into every new Parquet file output transform, so pipelines created with those versions write a row group every 20 kB and can produce footers of tens of megabytes. Open such transforms and set the row group size back to the default, or at least to a few megabytes. The transform logs a warning at startup when the row group size is below 1 MB.

With detailed logging enabled the transform logs every file it closes with its size, its number of rows and row groups and the size of its footer, so you can see the effect of the settings on your data without inspecting the files.

Partitioning

Parquet file output Partitioning tab

Leave Partition by empty to write a single file set, as before.

When one or more fields are listed, the transform writes Hive-style name=value folders under the base file name, and those fields are not stored inside the Parquet files.

Option Description

Existing data

What to do when a partition folder already exists: append, overwrite that partition, fail if it exists, or overwrite the whole base folder. This option is enabled when at least one partition field is set.

Maximum open partitions

How many partition files may be written at the same time. Each open file buffers up to one row group in memory, so a lower number uses less memory but writes more files. This option is enabled when at least one partition field is set.

Partition by

Incoming fields used as directory levels. Leave empty to write a single file set.

Fields

Parquet file output Fields tab

Option Description

Fields

You can specify which fields to write and in which order. You can use the "Get Fields" button to populate the dialog. Leave empty to output all input fields.

Hop to Parquet types

Every field is written as an optional column of the Parquet type below.

Hop type Parquet column

Integer

INT64

Number

DOUBLE

BigNumber with a length of 1 to 38

DECIMAL(length, precision), rounded with the rounding type of the field. A value with more digits before the decimal point than the column holds stops the transform with an error.

BigNumber without a length, or longer than 38

STRING, the number as text

String

STRING

Boolean

BOOLEAN

Date

TIMESTAMP(MILLIS), adjusted to UTC

Timestamp

TIMESTAMP(MICROS), adjusted to UTC. The nanoseconds beyond the microsecond are dropped.

Binary

BINARY

JSON

JSON, the JSON text

UUID

UUID, the 16 bytes of the UUID

Other Hop types can’t be written to Parquet.