Data Science Pilot Action Set

Provides actions for automating data science workflows, including automatic machine learning pipeline exploration, execution and ranking.

exploreData Action

Exploration, automatic variable analysis and grouping using comprehensive statistical profiling of the variables..

dataSciencePilot.exploreData <result=results> <status=rc> /
required parameter casOut
={
caslib="string",
indexVars={"variable-name-1" <, "variable-name-2", ...>},
lifetime=64-bit-integer,
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
},
ecdfTolerance=double,
event="string",
explorationPolicy
={
cv
={
lowMoment=double
lowRobust=double
},
dateTimeVariables={"variable-name-1" <, "variable-name-2", ...>},
dateVariables={"variable-name-1" <, "variable-name-2", ...>},
nominal
={
includeNegative=TRUE | FALSE
includeNonIntegral=TRUE | FALSE
intervals={"variable-name-1" <, "variable-name-2", ...>}
nominals={"variable-name-1" <, "variable-name-2", ...>}
},
timeVariables={"variable-name-1" <, "variable-name-2", ...>}
},
freq="variable-name",
inputs
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
misraGries=TRUE | FALSE,
required parameter table
={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
singlePass=TRUE | FALSE,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
target="variable-name",
weight="variable-name"
;
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

required parametercasOut

—

specifies the CAS table to store the analysis results.

Parameter Descriptions

* casOut={casouttable}

specifies the CAS table to store the analysis results.

Long formcasOut={name="table-name"}
Shortcut formcasOut="table-name"

The casouttable value can be one or more of the following:

caslib="string"

specifies the name of the caslib for the output table.

indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data.

lifetime=64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
memoryFormat="DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

name="table-name"

specifies the name for the output table.

promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table that has the same name.

DefaultFALSE
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

distinctCountLimit=integer

specifies the distinct count limit. If the limit is exceeded, and the misraGries parameter is set to True, the Misra-Gries frequency sketch algorithm is used to estimate the frequency distribution. Otherwise, the distinct count operation is aborted.

Default10000
Minimum value256

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function. This value is used by the quantile sketch algorithm.

Default0.001
Range1E-06–0.1

event="string"

specifies the target variable level that you want to model. Multilevel classification problems are cast into a one-versus-all binary classification problem, where the value of the event parameter denotes the level that you are modeling.

explorationPolicy={avaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) policy.

Aliasesexploration
avapt

The avaptPolicy value can be one or more of the following:

cardinality={cardinalityAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) cardinality policy.

The cardinalityAvaptPolicy value can be one or more of the following:

lowMediumCutoff=double

specifies the cardinality threshold for the low-medium cutoff.

Default32
Range2–256
mediumHighCutoff=double

specifies the cardinality threshold for the medium-high cutoff.

Default64
Range2–1024
minNObsPerTargetLevel=double

specifies the minimum number of observations for each target level.

Default10
Range5–100
cv={cvAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) coefficient of variation policy.

AliascoefficientVariation

The cvAvaptPolicy value can be one or more of the following:

lowMoment=double

specifies the absolute value of the low-high percentage threshold for the moment coefficient of variation (CV).

Default1.0
Minimum value0
lowRobust=double

specifies the absolute value of the low-high percentage threshold for the robust coefficient of variation (CV).

Default1.0
Minimum value0
dateTimeVariables={"variable-name-1" <, "variable-name-2", ...>}

specifies the datetime variables.

AliasdateTime
dateVariables={"variable-name-1" <, "variable-name-2", ...>}

specifies the date variables.

Aliasdate
entropy={entropyAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) entropy policy.

The entropyAvaptPolicy value can be one or more of the following:

giniLowMediumCutoff=double

specifies the Gini entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
giniMediumHighCutoff=double

specifies the Gini entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
shannonLowMediumCutoff=double

specifies the Shannon entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
shannonMediumHighCutoff=double

specifies the Shannon entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
iqv={iqvAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) index of qualitative variation policy.

AliasqualitativeVariationIndex

The iqvAvaptPolicy value can be one or more of the following:

highTopBottom=double

specifies the low-high cutoff frequency ratio threshold between the most frequent and least frequent levels of a nominal variable.

AliashighTop1Bottom1
Default100
Minimum value1
highTopTwo=double

specifies the low-high cutoff frequency ratio threshold between the most frequent and second most frequent levels of a nominal variable.

AliashighTop1Top2
Default10
Minimum value1
highVariationRatio=double

specifies the variation ratio threshold for the low-high cutoff.

AliashighModVr
Default0.5
Range(0–1]
kurtosis={kurtosisAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) kurtosis policy.

The kurtosisAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the absolute value of the moment kurtosis threshold for the low-medium cutoff.

Default5
Minimum value0
momentMediumHighCutoff=double

specifies the absolute value of the moment kurtosis threshold for the medium-high cutoff.

Default10
Minimum value0
robustLowMediumCutoff=double

specifies the absolute value of the robust kurtosis threshold for the low-medium cutoff.

Default2
Minimum value0
robustMediumHighCutoff=double

specifies the absolute value of the robust kurtosis threshold for the medium-high cutoff.

Default3
Minimum value0
missing={missingAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) missing grouping policy.

The missingAvaptPolicy value can be one or more of the following:

lowMediumCutoff=double

specifies the missing percentage threshold for the low-medium cutoff.

Default5
Range0–100
mediumHighCutoff=double

specifies the missing percentage threshold for the medium-high cutoff.

Default25
Range0–100
nominal={nominalAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) nominal policy.

The nominalAvaptPolicy value can be one or more of the following:

cardinalityRatio=double

specifies the AVAPT nominal policy cardinality ratio threshold.

Default0.25
Range(0–1]
cardinalityThreshold=double

specifies the AVAPT nominal policy cardinality threshold.

Default1024
Minimum value32
includeNegative=TRUE | FALSE

when set to True, includes numeric variables with some negative values in the nominal analysis.

DefaultFALSE
includeNonIntegral=TRUE | FALSE

when set to True, includes numeric variables with some nonintegral values in the nominal analysis.

DefaultFALSE
intervals={"variable-name-1" <, "variable-name-2", ...>}

specifies variables to consider as intervals.

nominals={"variable-name-1" <, "variable-name-2", ...>}

specifies variables to consider as nominals.

outlier={outlierAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) outlier policy.

The outlierAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the z-score outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
momentMediumHighCutoff=double

specifies the z-score outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
robustLowMediumCutoff=double

specifies the modified interquartile range outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
robustMediumHighCutoff=double

specifies the modified interquartile range outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
skewness={skewnessAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) skewness policy.

The skewnessAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the moment skewness threshold for the low-medium cutoff.

Default2
Range0–100
momentMediumHighCutoff=double

specifies the moment skewness threshold for the medium-high cutoff.

Default10
Range0–100
robustLowMediumCutoff=double

specifies the robust skewness threshold for the low-medium cutoff.

Default0.75
Range0–3
robustMediumHighCutoff=double

specifies the robust skewness threshold for the medium-high cutoff.

Default2
Range0–3
timeVariables={"variable-name-1" <, "variable-name-2", ...>}

specifies the time variables.

Aliastime

freq="variable-name"

specifies the frequency variable.

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

Aliasvars

misraGries=TRUE | FALSE

when set to True, uses the Misra-Gries algorithm for the frequency distribution estimation, if the distinct count limit is exceeded.

DefaultTRUE

* table={castable}

specifies the table name, caslib, and other common parameters.

Long formtable={name="table-name"}
Shortcut formtable="table-name"

The castable value can be one or more of the following:

caslib="string"

specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.

AliascompVars

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasesoptions
dataSource
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the input table.

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the input data.

whereTable={groupbytable}

specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.

The groupbytable value can be one or more of the following:

casLib="string"

specifies the caslib for the filter table. By default, the active caslib is used.

dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}

specifies data source options.

Aliasesoptions
dataSource

For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter (Appendix A: Common Parameters).

importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the filter table.

vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variable names to use from the filter table.

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the data from the filter table.

target="variable-name"

specifies the target variable.

AliasevalVar

weight="variable-name"

specifies the weight variable.

exploreData Action

Exploration, automatic variable analysis and grouping using comprehensive statistical profiling of the variables..

results, info = s:dataSciencePilot_exploreData{
required parameter casOut
={
caslib="string",
indexVars={"variable-name-1" <, "variable-name-2", ...>},
lifetime=64-bit-integer,
name="table-name",
promote=true | false,
replace=true | false,
},
ecdfTolerance=double,
event="string",
explorationPolicy
={
cv
={
lowMoment=double
lowRobust=double
},
dateTimeVariables={"variable-name-1" <, "variable-name-2", ...>},
dateVariables={"variable-name-1" <, "variable-name-2", ...>},
nominal
={
includeNegative=true | false
includeNonIntegral=true | false
intervals={"variable-name-1" <, "variable-name-2", ...>}
nominals={"variable-name-1" <, "variable-name-2", ...>}
},
timeVariables={"variable-name-1" <, "variable-name-2", ...>}
},
freq="variable-name",
inputs
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
misraGries=true | false,
required parameter table
={
caslib="string",
computedOnDemand=true | false,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
singlePass=true | false,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
target="variable-name",
weight="variable-name"
}
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

required parametercasOut

—

specifies the CAS table to store the analysis results.

Parameter Descriptions

* casOut={casouttable}

specifies the CAS table to store the analysis results.

Long formcasOut={name="table-name"}
Shortcut formcasOut="table-name"

The casouttable value can be one or more of the following:

caslib="string"

specifies the name of the caslib for the output table.

indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data.

lifetime=64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
memoryFormat="DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

name="table-name"

specifies the name for the output table.

promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table that has the same name.

Defaultfalse
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

distinctCountLimit=integer

specifies the distinct count limit. If the limit is exceeded, and the misraGries parameter is set to True, the Misra-Gries frequency sketch algorithm is used to estimate the frequency distribution. Otherwise, the distinct count operation is aborted.

Default10000
Minimum value256

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function. This value is used by the quantile sketch algorithm.

Default0.001
Range1E-06–0.1

event="string"

specifies the target variable level that you want to model. Multilevel classification problems are cast into a one-versus-all binary classification problem, where the value of the event parameter denotes the level that you are modeling.

explorationPolicy={avaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) policy.

Aliasesexploration
avapt

The avaptPolicy value can be one or more of the following:

cardinality={cardinalityAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) cardinality policy.

The cardinalityAvaptPolicy value can be one or more of the following:

lowMediumCutoff=double

specifies the cardinality threshold for the low-medium cutoff.

Default32
Range2–256
mediumHighCutoff=double

specifies the cardinality threshold for the medium-high cutoff.

Default64
Range2–1024
minNObsPerTargetLevel=double

specifies the minimum number of observations for each target level.

Default10
Range5–100
cv={cvAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) coefficient of variation policy.

AliascoefficientVariation

The cvAvaptPolicy value can be one or more of the following:

lowMoment=double

specifies the absolute value of the low-high percentage threshold for the moment coefficient of variation (CV).

Default1.0
Minimum value0
lowRobust=double

specifies the absolute value of the low-high percentage threshold for the robust coefficient of variation (CV).

Default1.0
Minimum value0
dateTimeVariables={"variable-name-1" <, "variable-name-2", ...>}

specifies the datetime variables.

AliasdateTime
dateVariables={"variable-name-1" <, "variable-name-2", ...>}

specifies the date variables.

Aliasdate
entropy={entropyAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) entropy policy.

The entropyAvaptPolicy value can be one or more of the following:

giniLowMediumCutoff=double

specifies the Gini entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
giniMediumHighCutoff=double

specifies the Gini entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
shannonLowMediumCutoff=double

specifies the Shannon entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
shannonMediumHighCutoff=double

specifies the Shannon entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
iqv={iqvAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) index of qualitative variation policy.

AliasqualitativeVariationIndex

The iqvAvaptPolicy value can be one or more of the following:

highTopBottom=double

specifies the low-high cutoff frequency ratio threshold between the most frequent and least frequent levels of a nominal variable.

AliashighTop1Bottom1
Default100
Minimum value1
highTopTwo=double

specifies the low-high cutoff frequency ratio threshold between the most frequent and second most frequent levels of a nominal variable.

AliashighTop1Top2
Default10
Minimum value1
highVariationRatio=double

specifies the variation ratio threshold for the low-high cutoff.

AliashighModVr
Default0.5
Range(0–1]
kurtosis={kurtosisAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) kurtosis policy.

The kurtosisAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the absolute value of the moment kurtosis threshold for the low-medium cutoff.

Default5
Minimum value0
momentMediumHighCutoff=double

specifies the absolute value of the moment kurtosis threshold for the medium-high cutoff.

Default10
Minimum value0
robustLowMediumCutoff=double

specifies the absolute value of the robust kurtosis threshold for the low-medium cutoff.

Default2
Minimum value0
robustMediumHighCutoff=double

specifies the absolute value of the robust kurtosis threshold for the medium-high cutoff.

Default3
Minimum value0
missing={missingAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) missing grouping policy.

The missingAvaptPolicy value can be one or more of the following:

lowMediumCutoff=double

specifies the missing percentage threshold for the low-medium cutoff.

Default5
Range0–100
mediumHighCutoff=double

specifies the missing percentage threshold for the medium-high cutoff.

Default25
Range0–100
nominal={nominalAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) nominal policy.

The nominalAvaptPolicy value can be one or more of the following:

cardinalityRatio=double

specifies the AVAPT nominal policy cardinality ratio threshold.

Default0.25
Range(0–1]
cardinalityThreshold=double

specifies the AVAPT nominal policy cardinality threshold.

Default1024
Minimum value32
includeNegative=true | false

when set to True, includes numeric variables with some negative values in the nominal analysis.

Defaultfalse
includeNonIntegral=true | false

when set to True, includes numeric variables with some nonintegral values in the nominal analysis.

Defaultfalse
intervals={"variable-name-1" <, "variable-name-2", ...>}

specifies variables to consider as intervals.

nominals={"variable-name-1" <, "variable-name-2", ...>}

specifies variables to consider as nominals.

outlier={outlierAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) outlier policy.

The outlierAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the z-score outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
momentMediumHighCutoff=double

specifies the z-score outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
robustLowMediumCutoff=double

specifies the modified interquartile range outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
robustMediumHighCutoff=double

specifies the modified interquartile range outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
skewness={skewnessAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) skewness policy.

The skewnessAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the moment skewness threshold for the low-medium cutoff.

Default2
Range0–100
momentMediumHighCutoff=double

specifies the moment skewness threshold for the medium-high cutoff.

Default10
Range0–100
robustLowMediumCutoff=double

specifies the robust skewness threshold for the low-medium cutoff.

Default0.75
Range0–3
robustMediumHighCutoff=double

specifies the robust skewness threshold for the medium-high cutoff.

Default2
Range0–3
timeVariables={"variable-name-1" <, "variable-name-2", ...>}

specifies the time variables.

Aliastime

freq="variable-name"

specifies the frequency variable.

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

Aliasvars

misraGries=true | false

when set to True, uses the Misra-Gries algorithm for the frequency distribution estimation, if the distinct count limit is exceeded.

Defaulttrue

* table={castable}

specifies the table name, caslib, and other common parameters.

Long formtable={name="table-name"}
Shortcut formtable="table-name"

The castable value can be one or more of the following:

caslib="string"

specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=true | false

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
Defaultfalse
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.

AliascompVars

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasesoptions
dataSource
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the input table.

singlePass=true | false

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

Defaultfalse
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the input data.

whereTable={groupbytable}

specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.

The groupbytable value can be one or more of the following:

casLib="string"

specifies the caslib for the filter table. By default, the active caslib is used.

dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}

specifies data source options.

Aliasesoptions
dataSource

For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter (Appendix A: Common Parameters).

importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the filter table.

vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variable names to use from the filter table.

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the data from the filter table.

target="variable-name"

specifies the target variable.

AliasevalVar

weight="variable-name"

specifies the weight variable.

exploreData Action

Exploration, automatic variable analysis and grouping using comprehensive statistical profiling of the variables..

results=s.dataSciencePilot.exploreData(
required parameter casOut
={
"caslib":"string",
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"lifetime":64-bit-integer,
"name":"table-name",
"promote":True | False,
"replace":True | False,
},
ecdfTolerance=double,
event="string",
explorationPolicy
={
"cv"
:{
"lowMoment":double
"lowRobust":double
},
"dateTimeVariables":["variable-name-1" <, "variable-name-2", ...>],
"dateVariables":["variable-name-1" <, "variable-name-2", ...>],
"iqv"
:{
"highTopBottom":double
"highTopTwo":double
},
"missing"
:{
"lowMediumCutoff":double
},
"nominal"
:{
"includeNegative":True | False
"includeNonIntegral":True | False
"intervals":["variable-name-1" <, "variable-name-2", ...>]
"nominals":["variable-name-1" <, "variable-name-2", ...>]
},
"timeVariables":["variable-name-1" <, "variable-name-2", ...>]
},
freq="variable-name",
inputs
=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
misraGries=True | False,
required parameter table
={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"singlePass":True | False,
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression",
"whereTable"
:{
"casLib":"string"
"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter "name":"table-name"
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>]
"where":"where-expression"
}
},
target="variable-name",
weight="variable-name"
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

required parametercasOut

—

specifies the CAS table to store the analysis results.

Parameter Descriptions

* casOut={casouttable}

specifies the CAS table to store the analysis results.

Long formcasOut={"name":"table-name"}
Shortcut formcasOut="table-name"

The casouttable value can be one or more of the following:

"caslib":"string"

specifies the name of the caslib for the output table.

"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data.

"lifetime":64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
"memoryFormat":"DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

"name":"table-name"

specifies the name for the output table.

"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table that has the same name.

DefaultFalse
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

distinctCountLimit=integer

specifies the distinct count limit. If the limit is exceeded, and the misraGries parameter is set to True, the Misra-Gries frequency sketch algorithm is used to estimate the frequency distribution. Otherwise, the distinct count operation is aborted.

Default10000
Minimum value256

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function. This value is used by the quantile sketch algorithm.

Default0.001
Range1E-06–0.1

event="string"

specifies the target variable level that you want to model. Multilevel classification problems are cast into a one-versus-all binary classification problem, where the value of the event parameter denotes the level that you are modeling.

explorationPolicy={avaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) policy.

Aliasesexploration
avapt

The avaptPolicy value can be one or more of the following:

"cardinality":{cardinalityAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) cardinality policy.

The cardinalityAvaptPolicy value can be one or more of the following:

"lowMediumCutoff":double

specifies the cardinality threshold for the low-medium cutoff.

Default32
Range2–256
"mediumHighCutoff":double

specifies the cardinality threshold for the medium-high cutoff.

Default64
Range2–1024
"minNObsPerTargetLevel":double

specifies the minimum number of observations for each target level.

Default10
Range5–100
"cv":{cvAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) coefficient of variation policy.

AliascoefficientVariation

The cvAvaptPolicy value can be one or more of the following:

"lowMoment":double

specifies the absolute value of the low-high percentage threshold for the moment coefficient of variation (CV).

Default1.0
Minimum value0
"lowRobust":double

specifies the absolute value of the low-high percentage threshold for the robust coefficient of variation (CV).

Default1.0
Minimum value0
"dateTimeVariables":["variable-name-1" <, "variable-name-2", ...>]

specifies the datetime variables.

AliasdateTime
"dateVariables":["variable-name-1" <, "variable-name-2", ...>]

specifies the date variables.

Aliasdate
"entropy":{entropyAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) entropy policy.

The entropyAvaptPolicy value can be one or more of the following:

"giniLowMediumCutoff":double

specifies the Gini entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
"giniMediumHighCutoff":double

specifies the Gini entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
"shannonLowMediumCutoff":double

specifies the Shannon entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
"shannonMediumHighCutoff":double

specifies the Shannon entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
"iqv":{iqvAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) index of qualitative variation policy.

AliasqualitativeVariationIndex

The iqvAvaptPolicy value can be one or more of the following:

"highTopBottom":double

specifies the low-high cutoff frequency ratio threshold between the most frequent and least frequent levels of a nominal variable.

AliashighTop1Bottom1
Default100
Minimum value1
"highTopTwo":double

specifies the low-high cutoff frequency ratio threshold between the most frequent and second most frequent levels of a nominal variable.

AliashighTop1Top2
Default10
Minimum value1
"highVariationRatio":double

specifies the variation ratio threshold for the low-high cutoff.

AliashighModVr
Default0.5
Range(0–1]
"kurtosis":{kurtosisAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) kurtosis policy.

The kurtosisAvaptPolicy value can be one or more of the following:

"momentLowMediumCutoff":double

specifies the absolute value of the moment kurtosis threshold for the low-medium cutoff.

Default5
Minimum value0
"momentMediumHighCutoff":double

specifies the absolute value of the moment kurtosis threshold for the medium-high cutoff.

Default10
Minimum value0
"robustLowMediumCutoff":double

specifies the absolute value of the robust kurtosis threshold for the low-medium cutoff.

Default2
Minimum value0
"robustMediumHighCutoff":double

specifies the absolute value of the robust kurtosis threshold for the medium-high cutoff.

Default3
Minimum value0
"missing":{missingAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) missing grouping policy.

The missingAvaptPolicy value can be one or more of the following:

"lowMediumCutoff":double

specifies the missing percentage threshold for the low-medium cutoff.

Default5
Range0–100
"mediumHighCutoff":double

specifies the missing percentage threshold for the medium-high cutoff.

Default25
Range0–100
"nominal":{nominalAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) nominal policy.

The nominalAvaptPolicy value can be one or more of the following:

"cardinalityRatio":double

specifies the AVAPT nominal policy cardinality ratio threshold.

Default0.25
Range(0–1]
"cardinalityThreshold":double

specifies the AVAPT nominal policy cardinality threshold.

Default1024
Minimum value32
"includeNegative":True | False

when set to True, includes numeric variables with some negative values in the nominal analysis.

DefaultFalse
"includeNonIntegral":True | False

when set to True, includes numeric variables with some nonintegral values in the nominal analysis.

DefaultFalse
"intervals":["variable-name-1" <, "variable-name-2", ...>]

specifies variables to consider as intervals.

"nominals":["variable-name-1" <, "variable-name-2", ...>]

specifies variables to consider as nominals.

"outlier":{outlierAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) outlier policy.

The outlierAvaptPolicy value can be one or more of the following:

"momentLowMediumCutoff":double

specifies the z-score outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
"momentMediumHighCutoff":double

specifies the z-score outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
"robustLowMediumCutoff":double

specifies the modified interquartile range outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
"robustMediumHighCutoff":double

specifies the modified interquartile range outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
"skewness":{skewnessAvaptPolicy}

specifies the automatic variable analysis and grouping (AVAPT) skewness policy.

The skewnessAvaptPolicy value can be one or more of the following:

"momentLowMediumCutoff":double

specifies the moment skewness threshold for the low-medium cutoff.

Default2
Range0–100
"momentMediumHighCutoff":double

specifies the moment skewness threshold for the medium-high cutoff.

Default10
Range0–100
"robustLowMediumCutoff":double

specifies the robust skewness threshold for the low-medium cutoff.

Default0.75
Range0–3
"robustMediumHighCutoff":double

specifies the robust skewness threshold for the medium-high cutoff.

Default2
Range0–3
"timeVariables":["variable-name-1" <, "variable-name-2", ...>]

specifies the time variables.

Aliastime

freq="variable-name"

specifies the frequency variable.

inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

Aliasvars

misraGries=True | False

when set to True, uses the Misra-Gries algorithm for the frequency distribution estimation, if the distinct count limit is exceeded.

DefaultTrue

* table={castable}

specifies the table name, caslib, and other common parameters.

Long formtable={"name":"table-name"}
Shortcut formtable="table-name"

The castable value can be one or more of the following:

"caslib":"string"

specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

"computedOnDemand":True | False

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFalse
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.

AliascompVars

The casinvardesc value can be one or more of the following:

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of the format field plus the length of the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"computedVarsProgram":"string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}

specifies data source options.

Aliasesoptions
dataSource
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the input table.

"singlePass":True | False

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFalse
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use in the action.

The casinvardesc value can be one or more of the following:

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of the format field plus the length of the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"where":"where-expression"

specifies an expression for subsetting the input data.

"whereTable":{groupbytable}

specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.

The groupbytable value can be one or more of the following:

"casLib":"string"

specifies the caslib for the filter table. By default, the active caslib is used.

"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}

specifies data source options.

Aliasesoptions
dataSource

For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter (Appendix A: Common Parameters).

"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the filter table.

"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variable names to use from the filter table.

The casinvardesc value can be one or more of the following:

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of the format field plus the length of the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"where":"where-expression"

specifies an expression for subsetting the data from the filter table.

target="variable-name"

specifies the target variable.

AliasevalVar

weight="variable-name"

specifies the weight variable.

exploreData Action

Exploration, automatic variable analysis and grouping using comprehensive statistical profiling of the variables..

results <– cas.dataSciencePilot.exploreData(s,
required parameter casOut
=list(
caslib="string",
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
lifetime=64-bit-integer,
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
),
ecdfTolerance=double,
event="string",
explorationPolicy
=list(
cv
=list(
lowMoment=double
lowRobust=double
),
dateTimeVariables=list("variable-name-1" <, "variable-name-2", ...>),
dateVariables=list("variable-name-1" <, "variable-name-2", ...>),
iqv
=list(),
nominal
=list(
includeNegative=TRUE | FALSE
includeNonIntegral=TRUE | FALSE
intervals=list("variable-name-1" <, "variable-name-2", ...>)
nominals=list("variable-name-1" <, "variable-name-2", ...>)
),
timeVariables=list("variable-name-1" <, "variable-name-2", ...>)
),
freq="variable-name",
inputs
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
misraGries=TRUE | FALSE,
required parameter table
=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
singlePass=TRUE | FALSE,
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression",
whereTable
=list(
casLib="string"
dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
required parameter name="table-name"
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>)
where="where-expression"
)
),
target="variable-name",
weight="variable-name"
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

required parametercasOut

—

specifies the CAS table to store the analysis results.

Parameter Descriptions

* casOut=list(casouttable)

specifies the CAS table to store the analysis results.

Long formcasOut=list(name="table-name")
Shortcut formcasOut="table-name"

The casouttable value can be one or more of the following:

caslib="string"

specifies the name of the caslib for the output table.

indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data.

lifetime=64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
memoryFormat="DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

name="table-name"

specifies the name for the output table.

promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table that has the same name.

DefaultFALSE
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

distinctCountLimit=integer

specifies the distinct count limit. If the limit is exceeded, and the misraGries parameter is set to True, the Misra-Gries frequency sketch algorithm is used to estimate the frequency distribution. Otherwise, the distinct count operation is aborted.

Default10000
Minimum value256

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function. This value is used by the quantile sketch algorithm.

Default0.001
Range1E-06–0.1

event="string"

specifies the target variable level that you want to model. Multilevel classification problems are cast into a one-versus-all binary classification problem, where the value of the event parameter denotes the level that you are modeling.

explorationPolicy=list(avaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) policy.

Aliasesexploration
avapt

The avaptPolicy value can be one or more of the following:

cardinality=list(cardinalityAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) cardinality policy.

The cardinalityAvaptPolicy value can be one or more of the following:

lowMediumCutoff=double

specifies the cardinality threshold for the low-medium cutoff.

Default32
Range2–256
mediumHighCutoff=double

specifies the cardinality threshold for the medium-high cutoff.

Default64
Range2–1024
minNObsPerTargetLevel=double

specifies the minimum number of observations for each target level.

Default10
Range5–100
cv=list(cvAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) coefficient of variation policy.

AliascoefficientVariation

The cvAvaptPolicy value can be one or more of the following:

lowMoment=double

specifies the absolute value of the low-high percentage threshold for the moment coefficient of variation (CV).

Default1.0
Minimum value0
lowRobust=double

specifies the absolute value of the low-high percentage threshold for the robust coefficient of variation (CV).

Default1.0
Minimum value0
dateTimeVariables=list("variable-name-1" <, "variable-name-2", ...>)

specifies the datetime variables.

AliasdateTime
dateVariables=list("variable-name-1" <, "variable-name-2", ...>)

specifies the date variables.

Aliasdate
entropy=list(entropyAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) entropy policy.

The entropyAvaptPolicy value can be one or more of the following:

giniLowMediumCutoff=double

specifies the Gini entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
giniMediumHighCutoff=double

specifies the Gini entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
shannonLowMediumCutoff=double

specifies the Shannon entropy threshold for the low-medium cutoff.

Default0.25
Range0–1
shannonMediumHighCutoff=double

specifies the Shannon entropy threshold for the medium-high cutoff.

Default0.75
Range0–1
iqv=list(iqvAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) index of qualitative variation policy.

AliasqualitativeVariationIndex

The iqvAvaptPolicy value can be one or more of the following:

highTopBottom=double

specifies the low-high cutoff frequency ratio threshold between the most frequent and least frequent levels of a nominal variable.

AliashighTop1Bottom1
Default100
Minimum value1
highTopTwo=double

specifies the low-high cutoff frequency ratio threshold between the most frequent and second most frequent levels of a nominal variable.

AliashighTop1Top2
Default10
Minimum value1
highVariationRatio=double

specifies the variation ratio threshold for the low-high cutoff.

AliashighModVr
Default0.5
Range(0–1]
kurtosis=list(kurtosisAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) kurtosis policy.

The kurtosisAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the absolute value of the moment kurtosis threshold for the low-medium cutoff.

Default5
Minimum value0
momentMediumHighCutoff=double

specifies the absolute value of the moment kurtosis threshold for the medium-high cutoff.

Default10
Minimum value0
robustLowMediumCutoff=double

specifies the absolute value of the robust kurtosis threshold for the low-medium cutoff.

Default2
Minimum value0
robustMediumHighCutoff=double

specifies the absolute value of the robust kurtosis threshold for the medium-high cutoff.

Default3
Minimum value0
missing=list(missingAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) missing grouping policy.

The missingAvaptPolicy value can be one or more of the following:

lowMediumCutoff=double

specifies the missing percentage threshold for the low-medium cutoff.

Default5
Range0–100
mediumHighCutoff=double

specifies the missing percentage threshold for the medium-high cutoff.

Default25
Range0–100
nominal=list(nominalAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) nominal policy.

The nominalAvaptPolicy value can be one or more of the following:

cardinalityRatio=double

specifies the AVAPT nominal policy cardinality ratio threshold.

Default0.25
Range(0–1]
cardinalityThreshold=double

specifies the AVAPT nominal policy cardinality threshold.

Default1024
Minimum value32
includeNegative=TRUE | FALSE

when set to True, includes numeric variables with some negative values in the nominal analysis.

DefaultFALSE
includeNonIntegral=TRUE | FALSE

when set to True, includes numeric variables with some nonintegral values in the nominal analysis.

DefaultFALSE
intervals=list("variable-name-1" <, "variable-name-2", ...>)

specifies variables to consider as intervals.

nominals=list("variable-name-1" <, "variable-name-2", ...>)

specifies variables to consider as nominals.

outlier=list(outlierAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) outlier policy.

The outlierAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the z-score outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
momentMediumHighCutoff=double

specifies the z-score outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
robustLowMediumCutoff=double

specifies the modified interquartile range outlier percentage threshold for the low-medium cutoff.

Default1
Range0–100
robustMediumHighCutoff=double

specifies the modified interquartile range outlier percentage threshold for the medium-high cutoff.

Default2.5
Range0–100
skewness=list(skewnessAvaptPolicy)

specifies the automatic variable analysis and grouping (AVAPT) skewness policy.

The skewnessAvaptPolicy value can be one or more of the following:

momentLowMediumCutoff=double

specifies the moment skewness threshold for the low-medium cutoff.

Default2
Range0–100
momentMediumHighCutoff=double

specifies the moment skewness threshold for the medium-high cutoff.

Default10
Range0–100
robustLowMediumCutoff=double

specifies the robust skewness threshold for the low-medium cutoff.

Default0.75
Range0–3
robustMediumHighCutoff=double

specifies the robust skewness threshold for the medium-high cutoff.

Default2
Range0–3
timeVariables=list("variable-name-1" <, "variable-name-2", ...>)

specifies the time variables.

Aliastime

freq="variable-name"

specifies the frequency variable.

inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

Aliasvars

misraGries=TRUE | FALSE

when set to True, uses the Misra-Gries algorithm for the frequency distribution estimation, if the distinct count limit is exceeded.

DefaultTRUE

* table=list(castable)

specifies the table name, caslib, and other common parameters.

Long formtable=list(name="table-name")
Shortcut formtable="table-name"

The castable value can be one or more of the following:

caslib="string"

specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.

AliascompVars

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)

specifies data source options.

Aliasesoptions
dataSource
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the input table.

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use in the action.

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the input data.

whereTable=list(groupbytable)

specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.

The groupbytable value can be one or more of the following:

casLib="string"

specifies the caslib for the filter table. By default, the active caslib is used.

dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)

specifies data source options.

Aliasesoptions
dataSource

For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter (Appendix A: Common Parameters).

importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

For more information about specifying the importOptions parameter, see the common importOptions parameter (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the filter table.

vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variable names to use from the filter table.

The casinvardesc value can be one or more of the following:

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of the format field plus the length of the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the data from the filter table.

target="variable-name"

specifies the target variable.

AliasevalVar

weight="variable-name"

specifies the weight variable.

Last updated: August 04, 2026