Using the HDFS (Hadoop Distributed File System) Adapter

The HDFS adapter resides in dfx-esp-hdfs-adapter.jar, which bundles the Java publisher and subscriber SAS Event Stream Processing clients. The subscriber client receives event blocks and writes events in CSV format to an HDFS file. The publisher client reads events in CSV format from an HDFS file and injects event blocks into a source window of an engine.

The target HDFS and the name of the file within the file system are both passed as required parameters to the adapter.

The subscriber client enables you to specify values for HDFS block size and number of replicas. You can configure the subscriber client to periodically write the HDFS file using the optional periodicity or maxfilesize parameters. If so configured, a timestamp is appended to the filename of each written file.

You can configure the publisher client to read from a growing file. In that case, the publisher runs indefinitely and publishes event blocks whenever the HDFS file size increases.

You must define the DFESP_HDFS_JARS environment variable for the client target platform. This variable specifies the location of the Hadoop JAR files.

List the JAR files in this order: hadoop-common-*.jar , hadoop-hdfs-*.jar , common hadoop JARs.

For example, in Linux you would run this command line:

$ export DFESP_HDFS_JARS=/usr/local/hadoop-2.5.0/share/hadoop/common/hadoop-common-2.5.0.jar:
/usr/local/hadoop-2.5.0/share/hadoop/hdfs/hadoop-hdfs-2.5.0.jar:
/usr/local/hadoop-2.5.0/share/hadoop/common/lib/*

Subscriber usage:

$DFESP_HOME/bin/dfesp_hdfs_subscriber -u url -f hdfs -t outputfile <-b hdfsblocksize >
<-n hdfsnumreplicas > <-m maxfilesize > <-p periodicity >
<-d dateformat > <- g gdconfigfile > <-l native | solace | tervela | rabbitmq | kafka> <-o severe | warning | info>
<-c [ configfilesection]> <-O tokenlocation > <-j jaasconf> <-k krb5conf>

Parameter

Definition

—u url

Specifies the dfESP subscribe standard URL in the form dfESP://host:port/project/continuousquery/window?snapshot=true | false.

Append the following if needed:

?collapse=true | false

?rmretdel=true | false

—f hdfs

Specifies the target file system, in the form "hdfs://host:port”. Specifies the file system normally configured in property fs.defaultFS in file core-site.xml.

—t outputfile

Specifies the output CSV file, in the form “/path/filename.csv”.

—b hdfsblocksize

Specifies the HDFS block size in MB. The default value is 64MB.

—n hdfsnumreplicas

Specifies the HDFS number of replicas. The default value is 1.

—m maxfilesize

Specifies the output file periodicity in bytes.

-p periodicity

Specifies the output file periodicity in seconds.

—d dateformat

Specifies the format of ESP_DATETIME and ESP_TIMESTAMP fields in CSV events. The default behavior is these fields are interpreted as an integer number of seconds (ESP_DATETIME) or microseconds (ESP_TIMESTAMP) since epoch.

—l native | solace | tervela | rabbitmq | kafka

Specifies the transport type. When you specify solace, tervela, rabbitmq, or kafka transports instead of the default native transport, use the required client configuration files specified in Using Alternative Transport Libraries for Java Clients in SAS Event Stream Processing: Publish/Subscribe API.

-g gdconfigfile

Specifies the guaranteed delivery configuration file for the client.

-o severe | warning | info

Specifies the application logging level.

-c [configfilesection]

Specifies the name of the section in the file /etc/javaadapters.config to parse for configuration parameters.

-O tokenlocation

Specifies the location of the file in the local file system that contains the OAuth token required for authentication by the publish/subscribe server.

-j jaasconf

Specifies the location of the JAAS configuration on the local file system. This file is used for Kerberos authentication to the Hadoop grid. By default, there is no authentication.

-k krb5conf

Specifies the location of the Kerberos 5 configuration file on the local file system. This file is required for Kerberos authentication to the Hadoop grid. The default value is /etc/krb5.conf.

Publisher usage:



$DFESP_HOME/bin/dfesp_hdfs_publisher -u url -f hdfs –i inputfile <-b blocksize >
<-t> <-d dateformat > <- g gdconfigfile > <-l native | solace | tervela | rabbitmq | kafka>
<-o severe | warning | info> <-c [ configfilesection]> < -e> < -m csvfielddelimiter >
<-n > <-O tokenlocation > <-Q> <-C> <-F eventtype> <-s> <-q> <-j jaasconf> <-k krb5conf>

Parameter

Definition

—u url

Specifies the publish standard URL in the form "dfESP://host:port/project/continuousquery/window".

—f hdfs

Specifies the target file system, in the form “hdfs://host:port”.

—i inputfile

Specifies the input CSV file, in the form “/path/filename.csv”.

—b blocksize

Specifies the number of events per event block.

-t

Specifies that event blocks are transactional. The default is normal.

—d dateformat

Specifies the format of ESP_DATETIME and ESP_TIMESTAMP fields in CSV events. The default behavior is these fields are interpreted as an integer number of seconds (ESP_DATETIME) or microseconds (ESP_TIMESTAMP) since epoch.

-g gdconfigfile

Specifies the guaranteed delivery configuration file for the client.

—l native | solace | tervela | rabbitmq | kafka

Specifies the transport type. When you specify solace, tervela, rabbitmq, or kafka transports instead of the default native transport, use the required client configuration files specified in Using Alternative Transport Libraries for Java Clients in SAS Event Stream Processing: Publish/Subscribe API.

-o severe | warning | info

Specifies the application logging level.

-c [configfilesection]

Specifies the name of the section in file /etc/javaadapters.config to parse for configuration parameters.

-e

Specifies that when a field in an input CSV event cannot be parsed, the event is dropped, an error is logged, and publishing continues.

-m csvfielddelimiter

Specifies the character delimiter for field data in input CSV events. The default delimiter is the , character.

-n

Specifies that input events are missing the key field that is autogenerated by the source window.

-O tokenlocation

Specifies the location of the file in the local file system that contains the OAuth token required for authentication by the publish/subscribe server.

-Q

Specify to quiesce the project after all events are injected into the source window.

-q

Specifies that a received CSV line should be treated as an opaque string.

-s

Specifies to build events with opcode = Upsert instead of Insert.

-C

Prepends an opcode and comma to read CSV events. The opcode is Insert unless -s is enabled.

-F eventtype

Specifies the event type to Insert into input CSV events (with comma). Valid values are "normal" and "partialupdate".

-j jaasconf

Specifies the location of the JAAS configuration on the local file system. This file is used for Kerberos authentication to the Hadoop grid. By default, there is no authentication.

-k krb5conf

Specifies the location of the Kerberos 5 configuration file on the local file system. This file is required for Kerberos authentication to the Hadoop grid. The default value is /etc/krb5.conf.