Parquet file output
Description
The Parquet file output transform writes data into the Apache Parquet file format.
For more information on this see: Apache Parquet.
Options
The dialog is split across four tabs so it fits a 1080p screen at 100% zoom: File, Options, Partitioning, and Fields.
Notes:
-
The date optionally referenced in the output file name(s) will be the start of the pipeline execution.
-
Rows are buffered in memory until a row group is full (see the Row group size option), then encoded, compressed and written. Memory use is therefore bounded by the row group size (times the number of copies and, when partitioning, the number of open partitions), not by the split size.
File

| Option | Description |
|---|---|
Transform name |
Name of the transform this name has to be unique in a single pipeline. |
Base file name |
Specify the base filename. This is composed of where you want to write the Parquet file to as well as the start of the filename. Examples: Write to Amazon AWS S3 : Write to a local folder : |
Extension |
This is the extension of the file.
Usually this is simply |
Include date? |
Check this box if you want to include the date in the filename with mask |
Include time? |
Check this box if you want to include the time in the filename with mask |
Include date-time-format? |
Check this box if you want to include a specific custom date-time format in the filename |
Include transform copy number? |
Enable this option if you run this transform in multiple copies to not have multiple threads write to the same file.
The copy number is formatted with mask |
Split into parts and include number? |
Enable this option if you want to split the output into multiple parts.
Specify a split size larger than 0 and this is then the number of rows per file.
The file part (split) number will be included in the filename to make sure that the same file is not being overwritten.
The split number is formatted with mask |
Create parent folders? |
Create missing parent folders of the output path. |
Include compression codec before extension? |
When enabled (the default for new transforms), the compression codec is placed before the file extension (for example |
Options

| Option | Description |
|---|---|
Compression codec |
Here you can indicate which compression codec you want to use. The default is UNCOMPRESSED. |
Version |
Choose the protocol version of Parquet (1.0 or 2.0) |
Row group size (bytes) |
The size of a row group in bytes, not rows.
A row group is the unit of work for a reader: the transform buffers rows in memory until the row group is full, then encodes, compresses and writes it.
The default is |
Data page size (bytes) |
The size of a data page in bytes, the unit of encoding and compression inside a column chunk.
The default is |
Dictionary page size (bytes) |
The maximum size of a dictionary page in bytes.
When the dictionary of a column grows past this size the writer falls back to plain encoding for that column chunk.
The default is |
The three size options accept human-friendly numbers such as 256m (256,000,000), 1.5m or 1e6, and variables.
Row group size and the file footer
Every row group adds an entry per column to the file footer: offsets, sizes, encodings and min/max statistics.
The footer therefore grows with the number of row groups times the number of columns, and with the size of the string values in the statistics.
Readers such as Dremio and Apache Drill refuse files whose footer is larger than 16 MB (Max supported footer size is 16777216).
Small row groups are the usual cause of an oversized footer.
Hop 2.3 and 2.4 saved a row group size of 20000 into every new Parquet file output transform, so pipelines created with those versions write a row group every 20 kB and can produce footers of tens of megabytes.
Open such transforms and set the row group size back to the default, or at least to a few megabytes.
The transform logs a warning at startup when the row group size is below 1 MB.
With detailed logging enabled the transform logs every file it closes with its size, its number of rows and row groups and the size of its footer, so you can see the effect of the settings on your data without inspecting the files.
Partitioning

Leave Partition by empty to write a single file set, as before.
When one or more fields are listed, the transform writes Hive-style name=value folders under the base file name, and those fields are not stored inside the Parquet files.
| Option | Description |
|---|---|
Existing data |
What to do when a partition folder already exists: append, overwrite that partition, fail if it exists, or overwrite the whole base folder. This option is enabled when at least one partition field is set. |
Maximum open partitions |
How many partition files may be written at the same time. Each open file buffers up to one row group in memory, so a lower number uses less memory but writes more files. This option is enabled when at least one partition field is set. |
Partition by |
Incoming fields used as directory levels. Leave empty to write a single file set. |
Fields

| Option | Description |
|---|---|
Fields |
You can specify which fields to write and in which order. You can use the "Get Fields" button to populate the dialog. Leave empty to output all input fields. |
Hop to Parquet types
Every field is written as an optional column of the Parquet type below.
| Hop type | Parquet column |
|---|---|
Integer |
INT64 |
Number |
DOUBLE |
BigNumber with a length of 1 to 38 |
DECIMAL(length, precision), rounded with the rounding type of the field. A value with more digits before the decimal point than the column holds stops the transform with an error. |
BigNumber without a length, or longer than 38 |
STRING, the number as text |
String |
STRING |
Boolean |
BOOLEAN |
Date |
TIMESTAMP(MILLIS), adjusted to UTC |
Timestamp |
TIMESTAMP(MICROS), adjusted to UTC. The nanoseconds beyond the microsecond are dropped. |
Binary |
BINARY |
JSON |
JSON, the JSON text |
UUID |
UUID, the 16 bytes of the UUID |
Other Hop types can’t be written to Parquet.