Parquet file input
Description
The Parquet file input transform reads (primitive) values from an Apache Parquet file.
For more information on this see: Apache Parquet.
Options
Notes:
-
Files are read in place through Apache VFS: local files directly, other locations (S3, Azure, …) by seeking in the file.
-
All input values are passed to the output.
-
A column must have the same type in every file read by the transform, or be read into a Hop type both of its types convert to (see the table below). A value which can’t be converted stops the transform with an error naming the Parquet column, its type and the field.
-
Timestamps annotated as adjusted to UTC are instants. The ones which aren’t hold the date and time fields of a local time, taken in the time zone of the Hop JVM.
| Option | Description |
|---|---|
Transform name |
Name of the transform this name has to be unique in a single pipeline. |
Filename field |
Specify the input field. Use a transform like Get File Names to obtain file names. Any supported file location is fine. |
Metadata filename |
If you specify a filename here, you can leave the fields section empty and Hop will automatically determine the output fields. It prevents you from having to define all the fields when this metadata is already in a parquet file schema. |
Output null row when empty |
In case you want to extract file metadata and there were no rows found in the parquet file(s) you will receive one empty row. This row can then be used to extract metadata with the Metadata structure of stream transform. |
Fields |
In this table you can specify all the fields you want to obtain from the parquet files as well as their desired Hop output type. |
Get fields button |
With this button you can select a parquet file from which we’ll read the schema to populate the Fields grid. |
Parquet to Hop types
The Get fields button proposes a Hop type for every column, based on its logical type or, without one, its physical type. You can pick another Hop type from the last column instead.
| Parquet column | Get fields proposes | Can also be read as |
|---|---|---|
BOOLEAN |
Boolean |
String, Integer (1 or 0) |
INT32, INT64, INT(8/16/32/64, signed), INT(8/16/32, unsigned) |
Integer |
String, Number, BigNumber |
INT(64, unsigned) |
BigNumber |
String, Number, Integer (up to 9223372036854775807) |
DECIMAL (on INT32, INT64, FIXED_LEN_BYTE_ARRAY or BINARY) |
BigNumber, with the precision as length and the scale as precision |
String, Number, Integer (the fraction is dropped), Binary (the stored bytes) |
FLOAT, DOUBLE |
Number |
String, BigNumber, Integer (rounded) |
FLOAT16 |
Number |
String, BigNumber, Integer (rounded), Binary (the stored bytes) |
DATE |
Date, midnight in the local time zone |
Timestamp. String, Integer, Number and BigNumber get the stored number of days. |
TIME (MILLIS, MICROS, NANOS) |
Timestamp, the time on 1970-01-01 |
Date. String, Integer, Number and BigNumber get the stored number. |
TIMESTAMP (MILLIS, MICROS, NANOS) |
Timestamp, with its full precision |
Date. String, Integer, Number and BigNumber get the stored number. |
INT96 |
Timestamp |
Date |
STRING, ENUM |
String |
BigNumber (text holding a number), JSON (text holding JSON), Binary |
JSON |
JSON |
String, BigNumber (text holding a number), Binary |
UUID |
UUID, or String when the UUID value type isn’t installed |
String, Binary |
BINARY, FIXED_LEN_BYTE_ARRAY without a logical type |
Binary |
String (UTF-8 text), JSON, BigNumber (text holding a number, or else the unscaled bytes of a decimal with the length and precision of the field) |
BSON, INTERVAL, GEOMETRY, GEOGRAPHY |
Binary |
|
LIST, MAP, VARIANT and other nested columns |
Not supported |
|