HDFS Advanced Source Options
Besides the specified options in the HDFS Source Options section, you can also specify these options that are available for file system sources:- COMPRESSION_METHOD
- START_FILENAME
- END_FILENAME
- START_CREATED_TIMESTAMP
- END_CREATED_TIMESTAMP
- START_MODIFIED_TIMESTAMP
- END_MODIFIED_TIMESTAMP
- SORT_BY
- SORT_DIRECTION
- SORT_REWRITE
- PRINCIPAL
- KEYTAB_FILE
- TOKEN
- TOKEN_FILE
- RENEW_INTERVAL
CONFIG source option, the most useful properties are the ones with the dfs.client prefix. For the full property reference, see HDFS Default and Core Default. Only properties that affect client connections and read operations apply to data pipeline loading. This table describes some common properties.
Kerberos Authentication for HDFS Sources
HDFS sources support Kerberos authentication with these two methods:- Keytab file — Specify the
PRINCIPALandKEYTAB_FILEoptions. - Delegation token — Specify the
TOKENorTOKEN_FILEoption and, optionally, theRENEW_INTERVALoption.
CREATE PIPELINE SQL statement fails.
In most cases, you must also specify the dfs.namenode.kerberos.principal and dfs.data.transfer.protection configuration properties in the CONFIG option so that the loading process can identify the Kerberos principal of the namenode server.
HDFS only supports one Key Distribution Center (KDC) path for each JVM®. Therefore, the KDC path affects all HDFS pipelines within the same system instance. If multiple HDFS pipelines are running, they must share the krb5.conf file by listing the principals and realms in it, or you must run them sequentially.
A valid Kerberos configuration file (krb5.conf) must exist on all Loader Nodes regardless of the authentication method. For help locating the file, see Locating the krb5.conf Configuration File. If you do not configure the file location, the system searches the default operating system locations, such as the /etc/krb5.conf file on Linux®. To specify a different file location, or to refresh the file after you update its contents, set the streamloader.extractorEngineParameters.configurationOption.kerberosConfig configuration setting. This setting affects all data pipelines that run on the Loader Node. For details, see Configuration Settings for Data Pipelines.
This example sets the Kerberos configuration file location to the /etc/krb5.conf file.
SQL
RENEW_INTERVAL option specifies. The KDC cannot renew a token after the token reaches the renewal lifetime (renew_lifetime) that you configure in the Kerberos configuration. If the credentials expire before the load finishes, such as during a long-running or continuous data pipeline, the pipeline fails, and the load cannot progress with the expired credentials. For these cases, use keytab file authentication instead. To resume a failed load with a new delegation token, update the token using the ALTER PIPELINE SQL statement and restart the pipeline.
For examples, see “Load Data from HDFS with Kerberos Keytab Authentication” and “Load Data from HDFS with Kerberos Delegation Token Authentication” on the Data Pipelines reference page.
HDFS Loading Example
Follow these steps to load JSON data from an HDFS source into the Ocient System.Step 1: Create a New Database
Connect to a SQL Node using the Commands Supported by the Ocient JDBC CLI Program. Then, execute theCREATE DATABASE SQL statement to create the geo database.
SQL
Step 2: Create a New Table in the Database
Create thelocations table in the public schema to store location data with a name, zip code, and a point with latitude and longitude:
name— Location name as a not nullable stringzipcode— Zip code as an integerlocation— Location latitude and longitude as a point
SQL
Step 3: Preview and Create a Data Pipeline
Preview thelocations data pipeline to load JSON data from HDFS. This data pipeline uses the HDFS endpoint hdfs-namenode:9000, which consists of the namenode and port number, and the /locations/2026/**/*.json filter to load data from JSON files. The pipeline selects the name, zip code, and point data. For the point data construction, see ST_POINT.
SQL
locations data pipeline to load JSON data from HDFS.
SQL
locations pipeline, execute the START PIPELINE SQL statement to start the data pipeline.
Step 4: Observe the Load Progress
With your pipeline running, data begins to load immediately from the JSON files. If there are many files in each file group, the load process first sorts the files into batches, partitions them for parallel processing, and assigns them to Loader Nodes. You can check the pipeline status and progress by querying theinformation_schema.pipeline_status system catalog table or executing the SHOW PIPELINE_STATUS SQL statement.
COMPLETED, all data is available in the target table.
After a few seconds, the data is available for query in the public.locationstable.
DROP PIPELINE locations; SQL statement. Execution of this statement leaves the data in your target table, but removes metadata about the pipeline execution from the system.

