Data Preprocess Action Set: Syntax

Provides actions for data preprocessing and transformation

transform Action

Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.

dataPreprocess.transform <result=results> <status=rc> /
casOut
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
casOutBinDetails
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
casOutLevelBinMap
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
casOutVarTransInfo
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
code
={
casOut
={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
lifetime=64-bit-integer
maxMemSize=64-bit-integer
memoryFormat="DVR" | "INHERIT" | "STANDARD"
name="table-name"
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"
threadBlockSize=64-bit-integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
comment=TRUE | FALSE,
fmtWdth=integer,
indentSize=integer,
labelId=integer,
lineSize=integer,
noTrim=TRUE | FALSE,
tabForm=TRUE | FALSE
},
copyAllVars=TRUE | FALSE,
copyVars={"variable-name-1" <, "variable-name-2", ...>},
evaluationStats=TRUE | FALSE,
freq="variable-name",
fuzzyCompare=double,
includeInputVars=TRUE | FALSE,
includeMissingGroup=TRUE | FALSE,
maxIterations=integer,
misraGries=TRUE | FALSE,
outputTableOptions
={
forceTableReturn=TRUE | FALSE,
tableNames={"string-1" <, "string-2", ...>}
},
overrides
={
alpha=double,
binMapping="LEFT" | "RIGHT",
binMissing=TRUE | FALSE,
binOutliers=TRUE | FALSE,
emptyBins=TRUE | FALSE,
enforceBinaryLevels=TRUE | FALSE,
ivFactor=double,
minNObsInBin=64-bit-integer,
missingBinStats=TRUE | FALSE,
missingEvalNonEvent=TRUE | FALSE,
noDataLowerUpperBound=TRUE | FALSE,
outlierBinsStats=TRUE | FALSE,
woeAdjust=double,
woeDefinition="EVENT" | "NONEVENT"
},
requestPackages
={{
catTrans
={
arguments
={
maxNBins=integer
minNBins=integer
nBinsArray={integer-1 <, integer-2, ...>} | integer
overrides
={
binMissing=TRUE | FALSE
emptyBins=TRUE | FALSE
enforceBinaryLevels=TRUE | FALSE
ivFactor=double
minNObsInBin=64-bit-integer
missingBinStats=TRUE | FALSE
missingEvalNonEvent=TRUE | FALSE
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
}
preprocessRare=TRUE | FALSE
}
},
dateTime
={
method={"ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"}
},
discretize
={
arguments
={
binEnds={double-1 <, double-2, ...>}
binStarts={double-1 <, double-2, ...>}
binWidths={double-1 <, double-2, ...>}
cutPoints={double-1 <, double-2, ...>}
maxNBins=integer
minNBins=integer
nBinsArray={integer-1 <, integer-2, ...>} | integer
overrides
={
alpha=double
binMapping="LEFT" | "RIGHT"
binMissing=TRUE | FALSE
binOutliers=TRUE | FALSE
emptyBins=TRUE | FALSE
enforceBinaryLevels=TRUE | FALSE
ivFactor=double
minNObsInBin=64-bit-integer
missingBinStats=TRUE | FALSE
missingEvalNonEvent=TRUE | FALSE
outlierBinsStats=TRUE | FALSE
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
}
}
},
evaluationStats=TRUE | FALSE | {evaluationStatsOptions},
events={"string-1" <, "string-2", ...>},
hash
={
arguments
={
nBuckets=integer
}
method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
},
impute
={
maxRandom=double
minRandom=double
valuesInterval={double-1 <, double-2, ...>}
valuesNominal={"string-1" <, "string-2", ...>}
},
inputs
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
inputsInheritFormats=TRUE | FALSE,
name="string",
outlier
={
arguments
={
aadLocationUseMean=TRUE | FALSE
max=double
min=double
replacements={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {double-1 <, double-2, ...>}
userDefinedLimits={double-1 <, double-2, ...>}
}
},
output
={
noScoreCode=TRUE | FALSE
noScoreTable=TRUE | FALSE
scoreWOE=TRUE | FALSE
},
targets
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
}, {...}},
sasVarNameLength=TRUE | FALSE,
saveState
={
caslib="string",
indexVars={"variable-name-1" <, "variable-name-2", ...>},
lifetime=64-bit-integer,
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
},
seed=integer,
required parameter table
={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
computedVarsProgram="string",
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
orderBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
tolerance=double,
weight="variable-name"
;
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOut

—

specifies the settings for an output table.

 casOutBinDetails

—

specifies the settings for an output table that includes information about the binning results.

 casOutLevelBinMap

—

specifies the settings for an output table that contains the nominal level bin mapping information.

 casOutVarTransInfo

—

specifies the settings for an output table that includes information for the variable transformations.

 code

casOut

specifies the settings for generating SAS DATA step scoring code.

 saveState

—

specifies the settings for an output table that contains the transformation model table.

Parameter Descriptions

casOut={casouttable}

specifies the settings for an output table.

For more information about specifying the casOut parameter, see the common casouttable parameter.

casOutBinDetails={casouttable}

specifies the settings for an output table that includes information about the binning results.

For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.

casOutLevelBinMap={casouttable}

specifies the settings for an output table that contains the nominal level bin mapping information.

For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.

casOutVarTransInfo={casouttable}

specifies the settings for an output table that includes information for the variable transformations.

For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.

code={codegen}

specifies the settings for generating SAS DATA step scoring code.

For more information about specifying the code parameter, see the common codegen parameter.

copyAllVars=TRUE | FALSE

when set to True, all the variables from the input table are copied to the scored output table.

AliasallIdVars
DefaultFALSE

copyVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.

distinctCountLimit=integer

specifies the distinct count limit.

evaluationStats=TRUE | FALSE

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats
DefaultFALSE

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05

includeInputVars=TRUE | FALSE

when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.

DefaultFALSE

includeMissingGroup=TRUE | FALSE

when set to True, missing values are allowed as group-by keys.

DefaultFALSE

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

misraGries=TRUE | FALSE

specifies that the Misra-Gries algorithm be used for most frequent estimation.

DefaultFALSE

outputTableOptions={outputTableOptions}

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliastblOpts

The outputTableOptions value can be one or more of the following:

forceTableReturn=TRUE | FALSE

when set to True, result tables are returned to the client even if the output is also saved as an output table.

DefaultFALSE
tableNames={"string-1" <, "string-2", ...>}

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

alpha=double

specifies the significance level.

Default0.05
binMapping="LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
binMissing=TRUE | FALSE

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFALSE
binOutliers=TRUE | FALSE

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

DefaultFALSE
emptyBins=TRUE | FALSE

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFALSE
enforceBinaryLevels=TRUE | FALSE

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTRUE
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=TRUE | FALSE

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTRUE
missingEvalNonEvent=TRUE | FALSE

when set to True, missing values of the target variables are considered as non-event values.

DefaultFALSE
noDataLowerUpperBound=TRUE | FALSE

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

DefaultFALSE
outlierBinsStats=TRUE | FALSE

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

DefaultTRUE
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch={quantileSketchOptions}

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

compressionFactor=double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
epsilon=double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

rank={rankOptions}

specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.

Long formrank={intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"}
Shortcut formrank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

The rankOptions value can be one or more of the following:

intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

AD

Anderson-Darling statistic.

AVGQUANKURT

average quantile skewness.

AVGQUANSKEW

average quantile kurtosis.

CLASSICALKURT

moment kurtosis.

CLASSICALSKEW

moment skewness.

CVM

Cramer-von Mises statistic.

KS

Kolmogorov-Smirnov statistic.

PEARSON

Pearson correlation.

VARIANCE

variance.

nominalStat="CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"

specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

CHISQ

chi-square statistic.

CRAMERSV

Cramer's V

FTEST

F-test statistic.

G2

g2 statistic.

GINI

gini Index.

IV

information value.

WELCHTTEST

Welch's t-test statistic.

WOE

weight of evidence.

topKInteractions=integer
Default10
Minimum value1
topKSave=integer
Default1
Minimum value1

requestPackages={{transformRequestPackage-1} <, {transformRequestPackage-2}, ...>}

specifies an array of transform request packages to be processed by the action.

Aliasespipelines
reqPacks

The transformRequestPackage value can be one or more of the following:

catTrans={catTransPhase}

specifies the parameters to use for the categorical transformation phase.

The catTransPhase value can be one or more of the following:

arguments={catTransArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The catTransArguments value can be one or more of the following:

contingencyTblOpts={contingencyTableOptions}

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

inputsNLevels=integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
inputsRawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
maxNBins=integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

minNBins=integer

specifies the minimum number of bins.

nBinsArray={integer-1 <, integer-2, ...>} | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

AliasesmiscellaneousOpts
opts

The globalOverrides value can be one or more of the following:

binMissing=TRUE | FALSE

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFALSE
emptyBins=TRUE | FALSE

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFALSE
enforceBinaryLevels=TRUE | FALSE

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTRUE
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=TRUE | FALSE

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTRUE
missingEvalNonEvent=TRUE | FALSE

when set to True, missing values of the target variables are considered as non-event values.

DefaultFALSE
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
preprocessRare=TRUE | FALSE

when set to True, rare levels are grouped into a single group at the start of the grouping process.

Aliaspreprocess
DefaultFALSE
rareThreshold=integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
rareThresholdPercent=double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Default5
Range(0, 100)
treeCrit="ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

DefaultGAINRATIO
ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
method="DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"

specifies the binning technique to use.

DTREE

groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.

GROUPRARE

groups rare levels of the analysis variable. This is an unsupervised technique. If you do not specify one of maxNLevels, rareFreqCutOff, or rareThresholdPer, then rareThresholdPer is set to 5.

ONEHOT

one hot encoding of categorical variables.

AliasLABEL
RTREE

groups based on a one-level regression tree. The criterion is the sum of squared error (SSE).

WOE

groups based on maximizing the information value (IV).

dateTime={dateTimePhase}

specifies the parameters to use for the date-time transformation phase.

The dateTimePhase value can be one or more of the following:

inputType="DATE" | "DATETIME" | "TIME"

input variable type as one of date, time or datetime.

DATE

date variable type

DATETIME

datetime variable type

TIME

time variable type

method={"ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"}

specifies the binning technique to use.

Aliastech
ALL

all (year, day, month, ...) of a datetime variable.

DAYMONTH

day of the month of a datetime variable.

DAYWEEK

day of the week of a datetime variable.

HOUR

hour of a datetime variable.

LEAPYEAR

datetime is a leap year or not.

MINUTE

minute of a datetime variable.

MONTH

month of a datetime variable.

QUARTER

quarter of the year of a datetime variable.

WEEK

week of a datetime variable.

WEEKEND

datetime is a weekend or not.

YEAR

year of a datetime variable.

discretize={discretizePhase}

specifies the parameters to use for the discretization phase.

The discretizePhase value can be one or more of the following:

arguments={discretizeArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The discretizeArguments value can be one or more of the following:

binEnds={double-1 <, double-2, ...>}

specifies the bin end values. If applicable, they override the data maximum values.

AliasbinEnd
binStarts={double-1 <, double-2, ...>}

specifies the bin start values. If applicable, they override the data minimum values.

AliasbinStart
binWidths={double-1 <, double-2, ...>}

specifies the bin width.

AliasbinWidth
contingencyTblOpts={contingencyTableOptions}

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

inputsNLevels=integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
inputsRawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
cutPoints={double-1 <, double-2, ...>}

specifies the user-provided cutpoints, for the CUTPTS binning technique.

AliascutPts
maxNBins=integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

Default5
Minimum value (exclusive)0
minNBins=integer

specifies the minimum number of bins.

Default1
Minimum value (exclusive)0
nBinsArray={integer-1 <, integer-2, ...>} | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

alpha=double

specifies the significance level.

Default0.05
binMapping="LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
binMissing=TRUE | FALSE

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFALSE
binOutliers=TRUE | FALSE

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

DefaultFALSE
emptyBins=TRUE | FALSE

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFALSE
enforceBinaryLevels=TRUE | FALSE

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTRUE
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=TRUE | FALSE

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTRUE
missingEvalNonEvent=TRUE | FALSE

when set to True, missing values of the target variables are considered as non-event values.

DefaultFALSE
noDataLowerUpperBound=TRUE | FALSE

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

DefaultFALSE
outlierBinsStats=TRUE | FALSE

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

DefaultTRUE
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
treeCrit="ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
method="BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"

specifies the binning technique to use.

BUCKET

creates equal-width bins.

CACC

creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.

CAIM

creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.

CHIMERGE

creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.

CUTPTS

creates bins according to the user-specified cutpoints.

DTREE

creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.

MDLP

creates bins based on the minimum description length. This is a top-down supervised discretization technique.

QUANTILE

creates equal-frequency bins.

RTREE

creates bins based on a one-level regression tree. This is a top-down supervised discretization technique.

WOE

creates bins based on WOE criterion. This is a top-down supervised discretization technique.

evaluationStats=TRUE | FALSE | {evaluationStatsOptions}

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats

The evaluationStatsOptions value can be one or more of the following:

chiSqGroup=TRUE | FALSE
when set to True, the Chi-Square, G2 and Cramer's V statistics are computed.
DefaultFALSE
ftTestGroup=TRUE | FALSE
when set to True, F-test and Welch's T-test statistics are computed.
DefaultFALSE
missIndicatorTarget=TRUE | FALSE
when specified, the target is transformed to a missing indicator binary target.
DefaultFALSE
nominalTarget=TRUE | FALSE
when set to True, the target variables are considered as nominal.
DefaultFALSE
woeGroup=TRUE | FALSE
when set to True, the WOE, IV and Gini index statistics are computed.
DefaultFALSE
events={"string-1" <, "string-2", ...>}

specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.

AliasevalVarsEvents
featureInteraction={featureInteraction}

options that control the generation of interaction features.

Aliasinteraction

The featureInteraction value can be one or more of the following:

coefficients={double-1 <, double-2, ...>}

specifies the coefficients for the linear interaction operator.

inputTransformations={"string-1" <, "string-2", ...>}

specifies the transformations that are to be used for generating the input component features of the interaction features.

Aliasinputs
method="CROSS" | "ORDERED"

specifies the method to use for feature generation. These include the cross and ordered component feature construction methods.

DefaultCROSS
CROSS

Cross product

ORDERED

Ordered grouping.

power=integer

specifies the number of inputs for the polynomial feature interaction operator.

Default2
Range1–4
synthesizer="DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"

specifies the operator to use for composing the values of the component feature interaction.

DIVISION

Division

LINEAR

Linear combination

MULTIPLICATION

Multiplication

NOMINAL

Nominal interaction

POLYNOMIAL

Polynomial feature generation

targetTransformation="string"

specifies the transformation that is to be used for generating the target features for interaction feature generation.

Aliasestargets
target
featureProbe={featureProbePhase}

specifies the parameters to use for the feature probe transformation phase.

The featureProbePhase value can be one or more of the following:

arguments={featureProbeArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The featureProbeArguments value can be one or more of the following:

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function.

Default0.001
Range1E-06–0.1
nProbes=integer

the number of feature probes.

Default1
Minimum value1
probeMissing=TRUE | FALSE

when set to True, generates missing values at the observed missing rate.

DefaultTRUE
rareThreshold=integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
rareThresholdPercent=double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Range(0, 100)
rawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
shrinkageFactor=double

specifies the shrinkage factor for level probability estimation.

Default0
Minimum value0
useRawLevel=TRUE | FALSE

specifies that the raw values be used as levels of the nominal variable.

DefaultFALSE
method="ECDF" | "FREQUENCY"

specifies the feature probe method to use.

ECDF

creates feature probes using the empirical cumulative distribution function.

FREQUENCY

creates feature probes using the empirical frequency distribution.

function={functionPhase}

specifies the parameters to use for the functional transformation phase.

The functionPhase value can be one or more of the following:

arguments={functionArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The functionArguments value can be one or more of the following:

aadLocationUseMean=TRUE | FALSE

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTRUE
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
lowerPercentile=double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Default10
Range(0, 50)
otherArguments={double-1 <, double-2, ...>}

specifies other values to use. The values depend on the type of functional transformation.

scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
scaleMultiplier=double

specifies the multiplying factor for the chosen scale estimator.

AliasscaleMulFac
shiftMax=double

specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value

AliasshiftNegative
shiftMin=double

specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value

AliasshiftPositive
symmetricPercentile=double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Default10
Range(0, 100)
upperPercentile=double

specifies the upper percentile threshold to use.

AliasupperPerc
Default90
Range(50, 100)
method="ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"

specifies the functional transformation.

ABS

returns the absolute value of the variable.

ARCSIN

returns the arcsine of the square root of a variable.

BOXCOX

returns the Box-Cox transformation of the variable.

CENTER

returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.

COS

returns the cosine of the variable.

COSH

returns the hyperbolic cosine of the variable.

EXP

returns the value of the e constant raised to the value of the variable.

IDENTITY

returns the value of a variable unmodified.

INVERSE

returns the inverse of a variable (1/x).

INVSQUARESHIFT

returns the inverse square shift value.

LOG

returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.

POWER

returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.

RANGE

returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.

SCALESHIFT

returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.

SIN

returns the sine of the variable.

SINH

returns the hyperbolic sine of the variable.

SQRT

returns the square root of the variable.

STANDARDIZE

returns the value minus the location determined from the loc parameter, and then scales it according to the scale parameter.

TAN

returns the tangent of the variable.

TANH

returns the hyperbolic tangent of the variable.

hash={hashPhase}

specifies the parameters to use for the hashing transformation phase.

The hashPhase value can be one or more of the following:

arguments={hashArguments}

specifies the arguments for this phase of the transform.

Aliasargs
nBuckets=integer

specifies the arguments for this phase of the transform.

method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"

specifies the hash function.

impute={imputePhase}

specifies the parameters to use for the imputation phase.

The imputePhase value can be one or more of the following:

maxRandom=double

specifies the maximum random number to generate.

method="MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"

specifies the binning technique to use.

MAX

replaces missing values with the maximum value. This technique applies to interval variables.

MEAN

replaces missing values with the mean. This technique applies to interval variables.

MEDIAN

replaces missing values with the median. This technique applies to interval variables.

MIDRANGE

replaces missing values with the mean of the maximum value and minimum value. This technique applies to interval variables.

MIN

replaces missing values with the minimum value. This technique applies to interval variables.

MODE

replaces missing values with the mode. This technique applies to nominal variables.

RANDOM

replaces missing values with uniform random numbers. This technique applies to interval variables.

VALUE

replaces missing values with the values specified in the valuesInterval and valuesNominal parameters.

minRandom=double

specifies the minimum random number to generate.

valuesInterval={double-1 <, double-2, ...>}

specifies a list of double values for imputation for the interval variables.

AliasvaluesNumeric
valuesNominal={"string-1" <, "string-2", ...>}

specifies a list of string values for imputation for the nominal variables.

AliasvaluesNonNumeric
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

inputsInheritFormats=TRUE | FALSE

specifies that the variables inherit formats from the underlying table.

DefaultFALSE
mapInterval={mapIntervalPhase}

specifies the parameters to use for the map to interval transformation phase.

The mapIntervalPhase value can be one or more of the following:

arguments={mapIntervalArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The mapIntervalArguments value can be one or more of the following:

descending=TRUE | FALSE

specifies that the label count encoding be performed.

DefaultTRUE
includeMissingLevel=TRUE | FALSE

when set to True, missing values are included in the distinct level analysis instead of being discarded.

DefaultFALSE
nLevels=integer

specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.

Default2
Range1–10
nMoments=integer

specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.

Default2
Range1–6
noise=double

specifies the parameter for the Laplace or uniform noise to be added to the level statistics.

Default0
Minimum value (exclusive)0
shrinkageFactor=double

specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.

Default0
Minimum value0
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
method="COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"

specifies the interval map criterion to use.

Aliastech
DefaultWOE
COUNTTARGET

maps to the counts of the levels of the nominal response.

AliasCOUNT
COUNTINPUT

count encoding of categorical variables.

EMPBAYES

maps to the empirical Bayes.

EVENTPROB

maps to the event probability.

FREQRATIO

maps to the frequency ratio of the levels of the nominal response.

LABELCOUNT

label count encoding of categorical variables.

MAX

maps to the max.

MIN

maps to the min.

MOMENTS

maps to the centralized moments.

WOE

maps to the weight of evidence (WOE).

name="string"

specifies a name for the request package.

outlier={outlierPhase}

specifies the parameters to use for the outlier determination and treatment phase.

The outlierPhase value can be one or more of the following:

arguments={outlierArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The outlierArguments value can be one or more of the following:

aadLocationUseMean=TRUE | FALSE

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTRUE
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

DefaultMEAN
BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
lowerPercentile=double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Range(0, 50)
max=double

specifies a global maximum value.

min=double

specifies a global minimum value.

replacements={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {double-1 <, double-2, ...>}

specifies the values to use as replacements for outliers. These can be user defined values or location estimates.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

DefaultSTD
AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
scaleMultiplier=double

specifies the multiplying factor for the chosen scale estimator.

symmetricPercentile=double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Range(0, 100)
upperPercentile=double

specifies the upper percentile threshold to use.

AliasupperPerc
Range(50, 100)
userDefinedLimits={double-1 <, double-2, ...>}

uses the specified user-defined limits as the lower and upper thresholds for each variable.

zScoreThreshold=double

specifies the Z threshold.

method="IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"

specifies the outlier definition.

IQR

uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.

MIQR

uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.

MZSCORE

uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.

PERC

uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.

UDFLIMITS

uses user-defined values to define outliers. Use the min, max, or userDefLims parameters to set the boundaries.

ZSCORE

uses the Z-score to define outliers. Uses the mean as the location, and the standard deviation as the scale estimates.

treatment="REPLACE" | "TRIM" | "WINSOR"

specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.

REPLACE

outliers are replaced with user defined values or location estimates.

TRIM

outliers are set to missing and discarded.

WINSOR

outliers are replaced with the lower or upper threshold and then binned.

output={outputPhase}

specifies the parameters to use for the output phase.

The outputPhase value can be one or more of the following:

noScoreCode=TRUE | FALSE

when set to True, no score code is sent to output.

DefaultFALSE
noScoreTable=TRUE | FALSE

when set to True, no scoring is sent to output.

AliasnoScoreTbl
DefaultFALSE
scoreWOE=TRUE | FALSE

when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.

DefaultFALSE
phaseOrder="FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"

specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.

DefaultIOF
FIO

specifies the phase order: function, impute and outlier.

FOI

specifies the phase order: function, outlier and impute.

IFO

specifies the phase order: impute, function and outlier.

IOF

specifies the phase order: impute, outlier and function.

OFI

specifies the phase order: outlier, function and impute.

OIF

specifies the phase order: outlier, impute and function.

targets={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies a list of target variables to use.

For more information about specifying the targets parameter, see the common casinvardesc parameter.

AliasevalVars
targetsInheritFormats=TRUE | FALSE

specifies that the variables inherit formats from the underlying table.

DefaultFALSE

sasVarNameLength=TRUE | FALSE

when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.

DefaultFALSE

saveState={casouttable}

specifies the settings for an output table that contains the transformation model table.

AliassaveModel
Long formsaveState={name="table-name"}
Shortcut formsaveState="table-name"

The casouttable value can be one or more of the following:

caslib="string"

specifies the name of the caslib for the output table.

indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data.

lifetime=64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
memoryFormat="DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

name="table-name"

specifies the name for the output table.

promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table that has the same name.

DefaultFALSE
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

* table={castable}

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

weight="variable-name"

specifies the weight variable.

transform Action

Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.

results, info = s:dataPreprocess_transform{
casOut
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=true | false,
replace=true | false,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
casOutBinDetails
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=true | false,
replace=true | false,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
casOutLevelBinMap
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=true | false,
replace=true | false,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
casOutVarTransInfo
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=true | false,
replace=true | false,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
code
={
casOut
={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
lifetime=64-bit-integer
maxMemSize=64-bit-integer
memoryFormat="DVR" | "INHERIT" | "STANDARD"
name="table-name"
promote=true | false
replace=true | false
replication=integer
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"
threadBlockSize=64-bit-integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
comment=true | false,
fmtWdth=integer,
indentSize=integer,
labelId=integer,
lineSize=integer,
noTrim=true | false,
tabForm=true | false
},
copyAllVars=true | false,
copyVars={"variable-name-1" <, "variable-name-2", ...>},
evaluationStats=true | false,
freq="variable-name",
fuzzyCompare=double,
includeInputVars=true | false,
includeMissingGroup=true | false,
maxIterations=integer,
misraGries=true | false,
outputTableOptions
={
forceTableReturn=true | false,
tableNames={"string-1" <, "string-2", ...>}
},
overrides
={
alpha=double,
binMapping="LEFT" | "RIGHT",
binMissing=true | false,
binOutliers=true | false,
emptyBins=true | false,
enforceBinaryLevels=true | false,
ivFactor=double,
minNObsInBin=64-bit-integer,
missingBinStats=true | false,
missingEvalNonEvent=true | false,
noDataLowerUpperBound=true | false,
outlierBinsStats=true | false,
woeAdjust=double,
woeDefinition="EVENT" | "NONEVENT"
},
requestPackages
={{
catTrans
={
arguments
={
maxNBins=integer
minNBins=integer
nBinsArray={integer-1 <, integer-2, ...>} | integer
overrides
={
binMissing=true | false
emptyBins=true | false
enforceBinaryLevels=true | false
ivFactor=double
minNObsInBin=64-bit-integer
missingBinStats=true | false
missingEvalNonEvent=true | false
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
}
preprocessRare=true | false
}
},
dateTime
={
method={"ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"}
},
discretize
={
arguments
={
binEnds={double-1 <, double-2, ...>}
binStarts={double-1 <, double-2, ...>}
binWidths={double-1 <, double-2, ...>}
cutPoints={double-1 <, double-2, ...>}
maxNBins=integer
minNBins=integer
nBinsArray={integer-1 <, integer-2, ...>} | integer
overrides
={
alpha=double
binMapping="LEFT" | "RIGHT"
binMissing=true | false
binOutliers=true | false
emptyBins=true | false
enforceBinaryLevels=true | false
ivFactor=double
minNObsInBin=64-bit-integer
missingBinStats=true | false
missingEvalNonEvent=true | false
outlierBinsStats=true | false
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
}
}
},
evaluationStats=true | false | {evaluationStatsOptions},
events={"string-1" <, "string-2", ...>},
hash
={
arguments
={
nBuckets=integer
}
method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
},
impute
={
maxRandom=double
minRandom=double
valuesInterval={double-1 <, double-2, ...>}
valuesNominal={"string-1" <, "string-2", ...>}
},
inputs
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
inputsInheritFormats=true | false,
name="string",
outlier
={
arguments
={
aadLocationUseMean=true | false
max=double
min=double
replacements={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {double-1 <, double-2, ...>}
userDefinedLimits={double-1 <, double-2, ...>}
}
},
output
={
noScoreCode=true | false
noScoreTable=true | false
scoreWOE=true | false
},
targets
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
}, {...}},
sasVarNameLength=true | false,
saveState
={
caslib="string",
indexVars={"variable-name-1" <, "variable-name-2", ...>},
lifetime=64-bit-integer,
name="table-name",
promote=true | false,
replace=true | false,
},
seed=integer,
required parameter table
={
caslib="string",
computedOnDemand=true | false,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
computedVarsProgram="string",
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
orderBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
tolerance=double,
weight="variable-name"
}
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOut

—

specifies the settings for an output table.

 casOutBinDetails

—

specifies the settings for an output table that includes information about the binning results.

 casOutLevelBinMap

—

specifies the settings for an output table that contains the nominal level bin mapping information.

 casOutVarTransInfo

—

specifies the settings for an output table that includes information for the variable transformations.

 code

casOut

specifies the settings for generating SAS DATA step scoring code.

 saveState

—

specifies the settings for an output table that contains the transformation model table.

Parameter Descriptions

casOut={casouttable}

specifies the settings for an output table.

For more information about specifying the casOut parameter, see the common casouttable parameter.

casOutBinDetails={casouttable}

specifies the settings for an output table that includes information about the binning results.

For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.

casOutLevelBinMap={casouttable}

specifies the settings for an output table that contains the nominal level bin mapping information.

For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.

casOutVarTransInfo={casouttable}

specifies the settings for an output table that includes information for the variable transformations.

For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.

code={codegen}

specifies the settings for generating SAS DATA step scoring code.

For more information about specifying the code parameter, see the common codegen parameter.

copyAllVars=true | false

when set to True, all the variables from the input table are copied to the scored output table.

AliasallIdVars
Defaultfalse

copyVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.

distinctCountLimit=integer

specifies the distinct count limit.

evaluationStats=true | false

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats
Defaultfalse

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05

includeInputVars=true | false

when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.

Defaultfalse

includeMissingGroup=true | false

when set to True, missing values are allowed as group-by keys.

Defaultfalse

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

misraGries=true | false

specifies that the Misra-Gries algorithm be used for most frequent estimation.

Defaultfalse

outputTableOptions={outputTableOptions}

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliastblOpts

The outputTableOptions value can be one or more of the following:

forceTableReturn=true | false

when set to True, result tables are returned to the client even if the output is also saved as an output table.

Defaultfalse
tableNames={"string-1" <, "string-2", ...>}

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

alpha=double

specifies the significance level.

Default0.05
binMapping="LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
binMissing=true | false

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
Defaultfalse
binOutliers=true | false

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

Defaultfalse
emptyBins=true | false

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

Defaultfalse
enforceBinaryLevels=true | false

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

Defaulttrue
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=true | false

when set to True, the missing bin is considered during the computation of the evaluation statistics.

Defaulttrue
missingEvalNonEvent=true | false

when set to True, missing values of the target variables are considered as non-event values.

Defaultfalse
noDataLowerUpperBound=true | false

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

Defaultfalse
outlierBinsStats=true | false

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

Defaulttrue
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch={quantileSketchOptions}

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

compressionFactor=double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
epsilon=double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

rank={rankOptions}

specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.

Long formrank={intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"}
Shortcut formrank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

The rankOptions value can be one or more of the following:

intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

AD

Anderson-Darling statistic.

AVGQUANKURT

average quantile skewness.

AVGQUANSKEW

average quantile kurtosis.

CLASSICALKURT

moment kurtosis.

CLASSICALSKEW

moment skewness.

CVM

Cramer-von Mises statistic.

KS

Kolmogorov-Smirnov statistic.

PEARSON

Pearson correlation.

VARIANCE

variance.

nominalStat="CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"

specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

CHISQ

chi-square statistic.

CRAMERSV

Cramer's V

FTEST

F-test statistic.

G2

g2 statistic.

GINI

gini Index.

IV

information value.

WELCHTTEST

Welch's t-test statistic.

WOE

weight of evidence.

topKInteractions=integer
Default10
Minimum value1
topKSave=integer
Default1
Minimum value1

requestPackages={{transformRequestPackage-1} <, {transformRequestPackage-2}, ...>}

specifies an array of transform request packages to be processed by the action.

Aliasespipelines
reqPacks

The transformRequestPackage value can be one or more of the following:

catTrans={catTransPhase}

specifies the parameters to use for the categorical transformation phase.

The catTransPhase value can be one or more of the following:

arguments={catTransArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The catTransArguments value can be one or more of the following:

contingencyTblOpts={contingencyTableOptions}

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

inputsNLevels=integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
inputsRawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
maxNBins=integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

minNBins=integer

specifies the minimum number of bins.

nBinsArray={integer-1 <, integer-2, ...>} | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

AliasesmiscellaneousOpts
opts

The globalOverrides value can be one or more of the following:

binMissing=true | false

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
Defaultfalse
emptyBins=true | false

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

Defaultfalse
enforceBinaryLevels=true | false

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

Defaulttrue
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=true | false

when set to True, the missing bin is considered during the computation of the evaluation statistics.

Defaulttrue
missingEvalNonEvent=true | false

when set to True, missing values of the target variables are considered as non-event values.

Defaultfalse
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
preprocessRare=true | false

when set to True, rare levels are grouped into a single group at the start of the grouping process.

Aliaspreprocess
Defaultfalse
rareThreshold=integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
rareThresholdPercent=double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Default5
Range(0, 100)
treeCrit="ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

DefaultGAINRATIO
ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
method="DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"

specifies the binning technique to use.

DTREE

groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.

GROUPRARE

groups rare levels of the analysis variable. This is an unsupervised technique. If you do not specify one of maxNLevels, rareFreqCutOff, or rareThresholdPer, then rareThresholdPer is set to 5.

ONEHOT

one hot encoding of categorical variables.

AliasLABEL
RTREE

groups based on a one-level regression tree. The criterion is the sum of squared error (SSE).

WOE

groups based on maximizing the information value (IV).

dateTime={dateTimePhase}

specifies the parameters to use for the date-time transformation phase.

The dateTimePhase value can be one or more of the following:

inputType="DATE" | "DATETIME" | "TIME"

input variable type as one of date, time or datetime.

DATE

date variable type

DATETIME

datetime variable type

TIME

time variable type

method={"ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"}

specifies the binning technique to use.

Aliastech
ALL

all (year, day, month, ...) of a datetime variable.

DAYMONTH

day of the month of a datetime variable.

DAYWEEK

day of the week of a datetime variable.

HOUR

hour of a datetime variable.

LEAPYEAR

datetime is a leap year or not.

MINUTE

minute of a datetime variable.

MONTH

month of a datetime variable.

QUARTER

quarter of the year of a datetime variable.

WEEK

week of a datetime variable.

WEEKEND

datetime is a weekend or not.

YEAR

year of a datetime variable.

discretize={discretizePhase}

specifies the parameters to use for the discretization phase.

The discretizePhase value can be one or more of the following:

arguments={discretizeArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The discretizeArguments value can be one or more of the following:

binEnds={double-1 <, double-2, ...>}

specifies the bin end values. If applicable, they override the data maximum values.

AliasbinEnd
binStarts={double-1 <, double-2, ...>}

specifies the bin start values. If applicable, they override the data minimum values.

AliasbinStart
binWidths={double-1 <, double-2, ...>}

specifies the bin width.

AliasbinWidth
contingencyTblOpts={contingencyTableOptions}

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

inputsNLevels=integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
inputsRawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
cutPoints={double-1 <, double-2, ...>}

specifies the user-provided cutpoints, for the CUTPTS binning technique.

AliascutPts
maxNBins=integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

Default5
Minimum value (exclusive)0
minNBins=integer

specifies the minimum number of bins.

Default1
Minimum value (exclusive)0
nBinsArray={integer-1 <, integer-2, ...>} | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

alpha=double

specifies the significance level.

Default0.05
binMapping="LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
binMissing=true | false

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
Defaultfalse
binOutliers=true | false

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

Defaultfalse
emptyBins=true | false

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

Defaultfalse
enforceBinaryLevels=true | false

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

Defaulttrue
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=true | false

when set to True, the missing bin is considered during the computation of the evaluation statistics.

Defaulttrue
missingEvalNonEvent=true | false

when set to True, missing values of the target variables are considered as non-event values.

Defaultfalse
noDataLowerUpperBound=true | false

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

Defaultfalse
outlierBinsStats=true | false

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

Defaulttrue
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
treeCrit="ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
method="BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"

specifies the binning technique to use.

BUCKET

creates equal-width bins.

CACC

creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.

CAIM

creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.

CHIMERGE

creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.

CUTPTS

creates bins according to the user-specified cutpoints.

DTREE

creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.

MDLP

creates bins based on the minimum description length. This is a top-down supervised discretization technique.

QUANTILE

creates equal-frequency bins.

RTREE

creates bins based on a one-level regression tree. This is a top-down supervised discretization technique.

WOE

creates bins based on WOE criterion. This is a top-down supervised discretization technique.

evaluationStats=true | false | {evaluationStatsOptions}

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats

The evaluationStatsOptions value can be one or more of the following:

chiSqGroup=true | false
when set to True, the Chi-Square, G2 and Cramer's V statistics are computed.
Defaultfalse
ftTestGroup=true | false
when set to True, F-test and Welch's T-test statistics are computed.
Defaultfalse
missIndicatorTarget=true | false
when specified, the target is transformed to a missing indicator binary target.
Defaultfalse
nominalTarget=true | false
when set to True, the target variables are considered as nominal.
Defaultfalse
woeGroup=true | false
when set to True, the WOE, IV and Gini index statistics are computed.
Defaultfalse
events={"string-1" <, "string-2", ...>}

specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.

AliasevalVarsEvents
featureInteraction={featureInteraction}

options that control the generation of interaction features.

Aliasinteraction

The featureInteraction value can be one or more of the following:

coefficients={double-1 <, double-2, ...>}

specifies the coefficients for the linear interaction operator.

inputTransformations={"string-1" <, "string-2", ...>}

specifies the transformations that are to be used for generating the input component features of the interaction features.

Aliasinputs
method="CROSS" | "ORDERED"

specifies the method to use for feature generation. These include the cross and ordered component feature construction methods.

DefaultCROSS
CROSS

Cross product

ORDERED

Ordered grouping.

power=integer

specifies the number of inputs for the polynomial feature interaction operator.

Default2
Range1–4
synthesizer="DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"

specifies the operator to use for composing the values of the component feature interaction.

DIVISION

Division

LINEAR

Linear combination

MULTIPLICATION

Multiplication

NOMINAL

Nominal interaction

POLYNOMIAL

Polynomial feature generation

targetTransformation="string"

specifies the transformation that is to be used for generating the target features for interaction feature generation.

Aliasestargets
target
featureProbe={featureProbePhase}

specifies the parameters to use for the feature probe transformation phase.

The featureProbePhase value can be one or more of the following:

arguments={featureProbeArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The featureProbeArguments value can be one or more of the following:

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function.

Default0.001
Range1E-06–0.1
nProbes=integer

the number of feature probes.

Default1
Minimum value1
probeMissing=true | false

when set to True, generates missing values at the observed missing rate.

Defaulttrue
rareThreshold=integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
rareThresholdPercent=double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Range(0, 100)
rawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
shrinkageFactor=double

specifies the shrinkage factor for level probability estimation.

Default0
Minimum value0
useRawLevel=true | false

specifies that the raw values be used as levels of the nominal variable.

Defaultfalse
method="ECDF" | "FREQUENCY"

specifies the feature probe method to use.

ECDF

creates feature probes using the empirical cumulative distribution function.

FREQUENCY

creates feature probes using the empirical frequency distribution.

function={functionPhase}

specifies the parameters to use for the functional transformation phase.

The functionPhase value can be one or more of the following:

arguments={functionArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The functionArguments value can be one or more of the following:

aadLocationUseMean=true | false

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
Defaulttrue
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
lowerPercentile=double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Default10
Range(0, 50)
otherArguments={double-1 <, double-2, ...>}

specifies other values to use. The values depend on the type of functional transformation.

scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
scaleMultiplier=double

specifies the multiplying factor for the chosen scale estimator.

AliasscaleMulFac
shiftMax=double

specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value

AliasshiftNegative
shiftMin=double

specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value

AliasshiftPositive
symmetricPercentile=double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Default10
Range(0, 100)
upperPercentile=double

specifies the upper percentile threshold to use.

AliasupperPerc
Default90
Range(50, 100)
method="ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"

specifies the functional transformation.

ABS

returns the absolute value of the variable.

ARCSIN

returns the arcsine of the square root of a variable.

BOXCOX

returns the Box-Cox transformation of the variable.

CENTER

returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.

COS

returns the cosine of the variable.

COSH

returns the hyperbolic cosine of the variable.

EXP

returns the value of the e constant raised to the value of the variable.

IDENTITY

returns the value of a variable unmodified.

INVERSE

returns the inverse of a variable (1/x).

INVSQUARESHIFT

returns the inverse square shift value.

LOG

returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.

POWER

returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.

RANGE

returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.

SCALESHIFT

returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.

SIN

returns the sine of the variable.

SINH

returns the hyperbolic sine of the variable.

SQRT

returns the square root of the variable.

STANDARDIZE

returns the value minus the location determined from the loc parameter, and then scales it according to the scale parameter.

TAN

returns the tangent of the variable.

TANH

returns the hyperbolic tangent of the variable.

hash={hashPhase}

specifies the parameters to use for the hashing transformation phase.

The hashPhase value can be one or more of the following:

arguments={hashArguments}

specifies the arguments for this phase of the transform.

Aliasargs
nBuckets=integer

specifies the arguments for this phase of the transform.

method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"

specifies the hash function.

impute={imputePhase}

specifies the parameters to use for the imputation phase.

The imputePhase value can be one or more of the following:

maxRandom=double

specifies the maximum random number to generate.

method="MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"

specifies the binning technique to use.

MAX

replaces missing values with the maximum value. This technique applies to interval variables.

MEAN

replaces missing values with the mean. This technique applies to interval variables.

MEDIAN

replaces missing values with the median. This technique applies to interval variables.

MIDRANGE

replaces missing values with the mean of the maximum value and minimum value. This technique applies to interval variables.

MIN

replaces missing values with the minimum value. This technique applies to interval variables.

MODE

replaces missing values with the mode. This technique applies to nominal variables.

RANDOM

replaces missing values with uniform random numbers. This technique applies to interval variables.

VALUE

replaces missing values with the values specified in the valuesInterval and valuesNominal parameters.

minRandom=double

specifies the minimum random number to generate.

valuesInterval={double-1 <, double-2, ...>}

specifies a list of double values for imputation for the interval variables.

AliasvaluesNumeric
valuesNominal={"string-1" <, "string-2", ...>}

specifies a list of string values for imputation for the nominal variables.

AliasvaluesNonNumeric
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

inputsInheritFormats=true | false

specifies that the variables inherit formats from the underlying table.

Defaultfalse
mapInterval={mapIntervalPhase}

specifies the parameters to use for the map to interval transformation phase.

The mapIntervalPhase value can be one or more of the following:

arguments={mapIntervalArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The mapIntervalArguments value can be one or more of the following:

descending=true | false

specifies that the label count encoding be performed.

Defaulttrue
includeMissingLevel=true | false

when set to True, missing values are included in the distinct level analysis instead of being discarded.

Defaultfalse
nLevels=integer

specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.

Default2
Range1–10
nMoments=integer

specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.

Default2
Range1–6
noise=double

specifies the parameter for the Laplace or uniform noise to be added to the level statistics.

Default0
Minimum value (exclusive)0
shrinkageFactor=double

specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.

Default0
Minimum value0
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
method="COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"

specifies the interval map criterion to use.

Aliastech
DefaultWOE
COUNTTARGET

maps to the counts of the levels of the nominal response.

AliasCOUNT
COUNTINPUT

count encoding of categorical variables.

EMPBAYES

maps to the empirical Bayes.

EVENTPROB

maps to the event probability.

FREQRATIO

maps to the frequency ratio of the levels of the nominal response.

LABELCOUNT

label count encoding of categorical variables.

MAX

maps to the max.

MIN

maps to the min.

MOMENTS

maps to the centralized moments.

WOE

maps to the weight of evidence (WOE).

name="string"

specifies a name for the request package.

outlier={outlierPhase}

specifies the parameters to use for the outlier determination and treatment phase.

The outlierPhase value can be one or more of the following:

arguments={outlierArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The outlierArguments value can be one or more of the following:

aadLocationUseMean=true | false

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
Defaulttrue
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

DefaultMEAN
BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
lowerPercentile=double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Range(0, 50)
max=double

specifies a global maximum value.

min=double

specifies a global minimum value.

replacements={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {double-1 <, double-2, ...>}

specifies the values to use as replacements for outliers. These can be user defined values or location estimates.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

DefaultSTD
AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
scaleMultiplier=double

specifies the multiplying factor for the chosen scale estimator.

symmetricPercentile=double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Range(0, 100)
upperPercentile=double

specifies the upper percentile threshold to use.

AliasupperPerc
Range(50, 100)
userDefinedLimits={double-1 <, double-2, ...>}

uses the specified user-defined limits as the lower and upper thresholds for each variable.

zScoreThreshold=double

specifies the Z threshold.

method="IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"

specifies the outlier definition.

IQR

uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.

MIQR

uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.

MZSCORE

uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.

PERC

uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.

UDFLIMITS

uses user-defined values to define outliers. Use the min, max, or userDefLims parameters to set the boundaries.

ZSCORE

uses the Z-score to define outliers. Uses the mean as the location, and the standard deviation as the scale estimates.

treatment="REPLACE" | "TRIM" | "WINSOR"

specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.

REPLACE

outliers are replaced with user defined values or location estimates.

TRIM

outliers are set to missing and discarded.

WINSOR

outliers are replaced with the lower or upper threshold and then binned.

output={outputPhase}

specifies the parameters to use for the output phase.

The outputPhase value can be one or more of the following:

noScoreCode=true | false

when set to True, no score code is sent to output.

Defaultfalse
noScoreTable=true | false

when set to True, no scoring is sent to output.

AliasnoScoreTbl
Defaultfalse
scoreWOE=true | false

when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.

Defaultfalse
phaseOrder="FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"

specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.

DefaultIOF
FIO

specifies the phase order: function, impute and outlier.

FOI

specifies the phase order: function, outlier and impute.

IFO

specifies the phase order: impute, function and outlier.

IOF

specifies the phase order: impute, outlier and function.

OFI

specifies the phase order: outlier, function and impute.

OIF

specifies the phase order: outlier, impute and function.

targets={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies a list of target variables to use.

For more information about specifying the targets parameter, see the common casinvardesc parameter.

AliasevalVars
targetsInheritFormats=true | false

specifies that the variables inherit formats from the underlying table.

Defaultfalse

sasVarNameLength=true | false

when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.

Defaultfalse

saveState={casouttable}

specifies the settings for an output table that contains the transformation model table.

AliassaveModel
Long formsaveState={name="table-name"}
Shortcut formsaveState="table-name"

The casouttable value can be one or more of the following:

caslib="string"

specifies the name of the caslib for the output table.

indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data.

lifetime=64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
memoryFormat="DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

name="table-name"

specifies the name for the output table.

promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table that has the same name.

Defaultfalse
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

* table={castable}

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

weight="variable-name"

specifies the weight variable.

transform Action

Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.

results=s.dataPreprocess.transform(
casOut
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"promote":True | False,
"replace":True | False,
"replication":integer,
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE",
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
casOutBinDetails
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"promote":True | False,
"replace":True | False,
"replication":integer,
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE",
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
casOutLevelBinMap
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"promote":True | False,
"replace":True | False,
"replication":integer,
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE",
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
casOutVarTransInfo
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"promote":True | False,
"replace":True | False,
"replication":integer,
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE",
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
code
={
"casOut"
:{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"lifetime":64-bit-integer
"maxMemSize":64-bit-integer
"memoryFormat":"DVR" | "INHERIT" | "STANDARD"
"name":"table-name"
"promote":True | False
"replace":True | False
"replication":integer
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE"
"threadBlockSize":64-bit-integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"comment":True | False,
"fmtWdth":integer,
"indentSize":integer,
"labelId":integer,
"lineSize":integer,
"noTrim":True | False,
"tabForm":True | False
},
copyAllVars=True | False,
copyVars=["variable-name-1" <, "variable-name-2", ...>],
evaluationStats=True | False,
freq="variable-name",
fuzzyCompare=double,
includeInputVars=True | False,
includeMissingGroup=True | False,
maxIterations=integer,
misraGries=True | False,
outputTableOptions
={
"forceTableReturn":True | False,
"tableNames":["string-1" <, "string-2", ...>]
},
overrides
={
"alpha":double,
"binMapping":"LEFT" | "RIGHT",
"binMissing":True | False,
"binOutliers":True | False,
"emptyBins":True | False,
"enforceBinaryLevels":True | False,
"ivFactor":double,
"minNObsInBin":64-bit-integer,
"minPerNObsInBin":double,
"missingBinStats":True | False,
"missingEvalNonEvent":True | False,
"noDataLowerUpperBound":True | False,
"outlierBinsStats":True | False,
"woeAdjust":double,
"woeDefinition":"EVENT" | "NONEVENT"
},
quantileSketch
={
"epsilon":double
},
requestPackages
=[{
"catTrans"
:{
"arguments"
:{
"maxNBins":integer
"minNBins":integer
"nBinsArray":[integer-1 <, integer-2, ...>] | integer
"overrides"
:{
"binMissing":True | False
"emptyBins":True | False
"enforceBinaryLevels":True | False
"ivFactor":double
"minNObsInBin":64-bit-integer
"missingBinStats":True | False
"missingEvalNonEvent":True | False
"woeAdjust":double
"woeDefinition":"EVENT" | "NONEVENT"
}
"preprocessRare":True | False
"rareThreshold":integer
}
},
"dateTime"
:{
"method":["ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"]
},
"discretize"
:{
"arguments"
:{
"binEnds":[double-1 <, double-2, ...>]
"binStarts":[double-1 <, double-2, ...>]
"binWidths":[double-1 <, double-2, ...>]
"cutPoints":[double-1 <, double-2, ...>]
"maxNBins":integer
"minNBins":integer
"nBinsArray":[integer-1 <, integer-2, ...>] | integer
"overrides"
:{
"alpha":double
"binMapping":"LEFT" | "RIGHT"
"binMissing":True | False
"binOutliers":True | False
"emptyBins":True | False
"enforceBinaryLevels":True | False
"ivFactor":double
"minNObsInBin":64-bit-integer
"missingBinStats":True | False
"missingEvalNonEvent":True | False
"noDataLowerUpperBound":True | False
"outlierBinsStats":True | False
"woeAdjust":double
"woeDefinition":"EVENT" | "NONEVENT"
}
}
},
"evaluationStats":True | False | {evaluationStatsOptions},
"events":["string-1" <, "string-2", ...>],
"featureInteraction"
:{
"coefficients":[double-1 <, double-2, ...>]
"inputTransformations":["string-1" <, "string-2", ...>]
"power":integer
},
"featureProbe"
:{
"arguments"
:{
"ecdfTolerance":double
"nProbes":integer
"probeMissing":True | False
"rareThreshold":integer
"useRawLevel":True | False
}
},
"hash"
:{
"arguments"
:{
"nBuckets":integer
}
"method":"BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
},
"impute"
:{
"maxRandom":double
"minRandom":double
"valuesInterval":[double-1 <, double-2, ...>]
"valuesNominal":["string-1" <, "string-2", ...>]
},
"inputs"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"inputsInheritFormats":True | False,
"mapInterval"
:{
"arguments"
:{
"descending":True | False
"includeMissingLevel":True | False
"nLevels":integer
"nMoments":integer
"noise":double
"woeAdjust":double
"woeDefinition":"EVENT" | "NONEVENT"
}
},
"name":"string",
"outlier"
:{
"arguments"
:{
"aadLocationUseMean":True | False
"max":double
"min":double
"replacements":["BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"] | [double-1 <, double-2, ...>]
"userDefinedLimits":[double-1 <, double-2, ...>]
}
},
"output"
:{
"noScoreCode":True | False
"noScoreTable":True | False
"scoreWOE":True | False
},
"targets"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"targetsInheritFormats":True | False
}<, {...}>],
sasVarNameLength=True | False,
saveState
={
"caslib":"string",
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"lifetime":64-bit-integer,
"name":"table-name",
"promote":True | False,
"replace":True | False,
},
seed=integer,
required parameter table
={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"groupByMode":"NOSORT" | "REDISTRIBUTE",
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"orderBy"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression",
"whereTable"
:{
"casLib":"string"
"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter "name":"table-name"
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>]
"where":"where-expression"
}
},
tolerance=double,
weight="variable-name"
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOut

—

specifies the settings for an output table.

 casOutBinDetails

—

specifies the settings for an output table that includes information about the binning results.

 casOutLevelBinMap

—

specifies the settings for an output table that contains the nominal level bin mapping information.

 casOutVarTransInfo

—

specifies the settings for an output table that includes information for the variable transformations.

 code

casOut

specifies the settings for generating SAS DATA step scoring code.

 saveState

—

specifies the settings for an output table that contains the transformation model table.

Parameter Descriptions

casOut={casouttable}

specifies the settings for an output table.

For more information about specifying the casOut parameter, see the common casouttable parameter.

casOutBinDetails={casouttable}

specifies the settings for an output table that includes information about the binning results.

For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.

casOutLevelBinMap={casouttable}

specifies the settings for an output table that contains the nominal level bin mapping information.

For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.

casOutVarTransInfo={casouttable}

specifies the settings for an output table that includes information for the variable transformations.

For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.

code={codegen}

specifies the settings for generating SAS DATA step scoring code.

For more information about specifying the code parameter, see the common codegen parameter.

copyAllVars=True | False

when set to True, all the variables from the input table are copied to the scored output table.

AliasallIdVars
DefaultFalse

copyVars=["variable-name-1" <, "variable-name-2", ...>]

specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.

distinctCountLimit=integer

specifies the distinct count limit.

evaluationStats=True | False

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats
DefaultFalse

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05

includeInputVars=True | False

when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.

DefaultFalse

includeMissingGroup=True | False

when set to True, missing values are allowed as group-by keys.

DefaultFalse

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

misraGries=True | False

specifies that the Misra-Gries algorithm be used for most frequent estimation.

DefaultFalse

outputTableOptions={outputTableOptions}

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliastblOpts

The outputTableOptions value can be one or more of the following:

"forceTableReturn":True | False

when set to True, result tables are returned to the client even if the output is also saved as an output table.

DefaultFalse
"tableNames":["string-1" <, "string-2", ...>]

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

overrides={globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

"alpha":double

specifies the significance level.

Default0.05
"binMapping":"LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
"binMissing":True | False

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFalse
"binOutliers":True | False

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

DefaultFalse
"emptyBins":True | False

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFalse
"enforceBinaryLevels":True | False

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTrue
"ivFactor":double

specifies the information value adjustment factor.

Default2
"minNObsInBin":64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
"minPerNObsInBin":double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
"missingBinStats":True | False

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTrue
"missingEvalNonEvent":True | False

when set to True, missing values of the target variables are considered as non-event values.

DefaultFalse
"noDataLowerUpperBound":True | False

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

DefaultFalse
"outlierBinsStats":True | False

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

DefaultTrue
"woeAdjust":double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
"woeDefinition":"EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch={quantileSketchOptions}

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

"compressionFactor":double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
"epsilon":double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

rank={rankOptions}

specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.

Long formrank={"intervalStat":"AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"}
Shortcut formrank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

The rankOptions value can be one or more of the following:

"intervalStat":"AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

AD

Anderson-Darling statistic.

AVGQUANKURT

average quantile skewness.

AVGQUANSKEW

average quantile kurtosis.

CLASSICALKURT

moment kurtosis.

CLASSICALSKEW

moment skewness.

CVM

Cramer-von Mises statistic.

KS

Kolmogorov-Smirnov statistic.

PEARSON

Pearson correlation.

VARIANCE

variance.

"nominalStat":"CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"

specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

CHISQ

chi-square statistic.

CRAMERSV

Cramer's V

FTEST

F-test statistic.

G2

g2 statistic.

GINI

gini Index.

IV

information value.

WELCHTTEST

Welch's t-test statistic.

WOE

weight of evidence.

"topKInteractions":integer
Default10
Minimum value1
"topKSave":integer
Default1
Minimum value1

requestPackages=[{transformRequestPackage-1} <, {transformRequestPackage-2}, ...>]

specifies an array of transform request packages to be processed by the action.

Aliasespipelines
reqPacks

The transformRequestPackage value can be one or more of the following:

"catTrans":{catTransPhase}

specifies the parameters to use for the categorical transformation phase.

The catTransPhase value can be one or more of the following:

"arguments":{catTransArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The catTransArguments value can be one or more of the following:

"contingencyTblOpts":{contingencyTableOptions}

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

"inputsMethod":"BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

"inputsNLevels":integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
"inputsRawLevelStartingValue":integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
"maxNBins":integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

"minNBins":integer

specifies the minimum number of bins.

"nBinsArray":[integer-1 <, integer-2, ...>] | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
"overrides":{globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

AliasesmiscellaneousOpts
opts

The globalOverrides value can be one or more of the following:

"binMissing":True | False

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFalse
"emptyBins":True | False

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFalse
"enforceBinaryLevels":True | False

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTrue
"ivFactor":double

specifies the information value adjustment factor.

Default2
"minNObsInBin":64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
"minPerNObsInBin":double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
"missingBinStats":True | False

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTrue
"missingEvalNonEvent":True | False

when set to True, missing values of the target variables are considered as non-event values.

DefaultFalse
"woeAdjust":double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
"woeDefinition":"EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
"preprocessRare":True | False

when set to True, rare levels are grouped into a single group at the start of the grouping process.

Aliaspreprocess
DefaultFalse
"rareThreshold":integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
"rareThresholdPercent":double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Default5
Range(0, 100)
"treeCrit":"ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

DefaultGAINRATIO
ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
"method":"DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"

specifies the binning technique to use.

DTREE

groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.

GROUPRARE

groups rare levels of the analysis variable. This is an unsupervised technique. If you do not specify one of maxNLevels, rareFreqCutOff, or rareThresholdPer, then rareThresholdPer is set to 5.

ONEHOT

one hot encoding of categorical variables.

AliasLABEL
RTREE

groups based on a one-level regression tree. The criterion is the sum of squared error (SSE).

WOE

groups based on maximizing the information value (IV).

"dateTime":{dateTimePhase}

specifies the parameters to use for the date-time transformation phase.

The dateTimePhase value can be one or more of the following:

"inputType":"DATE" | "DATETIME" | "TIME"

input variable type as one of date, time or datetime.

DATE

date variable type

DATETIME

datetime variable type

TIME

time variable type

"method":["ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"]

specifies the binning technique to use.

Aliastech
ALL

all (year, day, month, ...) of a datetime variable.

DAYMONTH

day of the month of a datetime variable.

DAYWEEK

day of the week of a datetime variable.

HOUR

hour of a datetime variable.

LEAPYEAR

datetime is a leap year or not.

MINUTE

minute of a datetime variable.

MONTH

month of a datetime variable.

QUARTER

quarter of the year of a datetime variable.

WEEK

week of a datetime variable.

WEEKEND

datetime is a weekend or not.

YEAR

year of a datetime variable.

"discretize":{discretizePhase}

specifies the parameters to use for the discretization phase.

The discretizePhase value can be one or more of the following:

"arguments":{discretizeArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The discretizeArguments value can be one or more of the following:

"binEnds":[double-1 <, double-2, ...>]

specifies the bin end values. If applicable, they override the data maximum values.

AliasbinEnd
"binStarts":[double-1 <, double-2, ...>]

specifies the bin start values. If applicable, they override the data minimum values.

AliasbinStart
"binWidths":[double-1 <, double-2, ...>]

specifies the bin width.

AliasbinWidth
"contingencyTblOpts":{contingencyTableOptions}

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

"inputsMethod":"BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

"inputsNLevels":integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
"inputsRawLevelStartingValue":integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
"cutPoints":[double-1 <, double-2, ...>]

specifies the user-provided cutpoints, for the CUTPTS binning technique.

AliascutPts
"maxNBins":integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

Default5
Minimum value (exclusive)0
"minNBins":integer

specifies the minimum number of bins.

Default1
Minimum value (exclusive)0
"nBinsArray":[integer-1 <, integer-2, ...>] | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
"overrides":{globalOverrides}

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

"alpha":double

specifies the significance level.

Default0.05
"binMapping":"LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
"binMissing":True | False

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFalse
"binOutliers":True | False

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

DefaultFalse
"emptyBins":True | False

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFalse
"enforceBinaryLevels":True | False

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTrue
"ivFactor":double

specifies the information value adjustment factor.

Default2
"minNObsInBin":64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
"minPerNObsInBin":double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
"missingBinStats":True | False

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTrue
"missingEvalNonEvent":True | False

when set to True, missing values of the target variables are considered as non-event values.

DefaultFalse
"noDataLowerUpperBound":True | False

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

DefaultFalse
"outlierBinsStats":True | False

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

DefaultTrue
"woeAdjust":double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
"woeDefinition":"EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
"treeCrit":"ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
"method":"BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"

specifies the binning technique to use.

BUCKET

creates equal-width bins.

CACC

creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.

CAIM

creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.

CHIMERGE

creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.

CUTPTS

creates bins according to the user-specified cutpoints.

DTREE

creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.

MDLP

creates bins based on the minimum description length. This is a top-down supervised discretization technique.

QUANTILE

creates equal-frequency bins.

RTREE

creates bins based on a one-level regression tree. This is a top-down supervised discretization technique.

WOE

creates bins based on WOE criterion. This is a top-down supervised discretization technique.

"evaluationStats":True | False | {evaluationStatsOptions}

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats

The evaluationStatsOptions value can be one or more of the following:

chiSqGroup=True | False
when set to True, the Chi-Square, G2 and Cramer's V statistics are computed.
DefaultFalse
ftTestGroup=True | False
when set to True, F-test and Welch's T-test statistics are computed.
DefaultFalse
missIndicatorTarget=True | False
when specified, the target is transformed to a missing indicator binary target.
DefaultFalse
nominalTarget=True | False
when set to True, the target variables are considered as nominal.
DefaultFalse
woeGroup=True | False
when set to True, the WOE, IV and Gini index statistics are computed.
DefaultFalse
"events":["string-1" <, "string-2", ...>]

specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.

AliasevalVarsEvents
"featureInteraction":{featureInteraction}

options that control the generation of interaction features.

Aliasinteraction

The featureInteraction value can be one or more of the following:

"coefficients":[double-1 <, double-2, ...>]

specifies the coefficients for the linear interaction operator.

"inputTransformations":["string-1" <, "string-2", ...>]

specifies the transformations that are to be used for generating the input component features of the interaction features.

Aliasinputs
"method":"CROSS" | "ORDERED"

specifies the method to use for feature generation. These include the cross and ordered component feature construction methods.

DefaultCROSS
CROSS

Cross product

ORDERED

Ordered grouping.

"power":integer

specifies the number of inputs for the polynomial feature interaction operator.

Default2
Range1–4
"synthesizer":"DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"

specifies the operator to use for composing the values of the component feature interaction.

DIVISION

Division

LINEAR

Linear combination

MULTIPLICATION

Multiplication

NOMINAL

Nominal interaction

POLYNOMIAL

Polynomial feature generation

"targetTransformation":"string"

specifies the transformation that is to be used for generating the target features for interaction feature generation.

Aliasestargets
target
"featureProbe":{featureProbePhase}

specifies the parameters to use for the feature probe transformation phase.

The featureProbePhase value can be one or more of the following:

"arguments":{featureProbeArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The featureProbeArguments value can be one or more of the following:

"ecdfTolerance":double

specifies the tolerance value for the empirical cumulative distribution function.

Default0.001
Range1E-06–0.1
"nProbes":integer

the number of feature probes.

Default1
Minimum value1
"probeMissing":True | False

when set to True, generates missing values at the observed missing rate.

DefaultTrue
"rareThreshold":integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
"rareThresholdPercent":double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Range(0, 100)
"rawLevelStartingValue":integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
"shrinkageFactor":double

specifies the shrinkage factor for level probability estimation.

Default0
Minimum value0
"useRawLevel":True | False

specifies that the raw values be used as levels of the nominal variable.

DefaultFalse
"method":"ECDF" | "FREQUENCY"

specifies the feature probe method to use.

ECDF

creates feature probes using the empirical cumulative distribution function.

FREQUENCY

creates feature probes using the empirical frequency distribution.

"function":{functionPhase}

specifies the parameters to use for the functional transformation phase.

The functionPhase value can be one or more of the following:

"arguments":{functionArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The functionArguments value can be one or more of the following:

"aadLocationUseMean":True | False

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTrue
"location":"BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

"locationBiweightTuning":double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
"lowerPercentile":double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Default10
Range(0, 50)
"otherArguments":[double-1 <, double-2, ...>]

specifies other values to use. The values depend on the type of functional transformation.

"scale":"AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

"scaleBiweightTuning":double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
"scaleMultiplier":double

specifies the multiplying factor for the chosen scale estimator.

AliasscaleMulFac
"shiftMax":double

specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value

AliasshiftNegative
"shiftMin":double

specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value

AliasshiftPositive
"symmetricPercentile":double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Default10
Range(0, 100)
"upperPercentile":double

specifies the upper percentile threshold to use.

AliasupperPerc
Default90
Range(50, 100)
"method":"ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"

specifies the functional transformation.

ABS

returns the absolute value of the variable.

ARCSIN

returns the arcsine of the square root of a variable.

BOXCOX

returns the Box-Cox transformation of the variable.

CENTER

returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.

COS

returns the cosine of the variable.

COSH

returns the hyperbolic cosine of the variable.

EXP

returns the value of the e constant raised to the value of the variable.

IDENTITY

returns the value of a variable unmodified.

INVERSE

returns the inverse of a variable (1/x).

INVSQUARESHIFT

returns the inverse square shift value.

LOG

returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.

POWER

returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.

RANGE

returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.

SCALESHIFT

returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.

SIN

returns the sine of the variable.

SINH

returns the hyperbolic sine of the variable.

SQRT

returns the square root of the variable.

STANDARDIZE

returns the value minus the location determined from the loc parameter, and then scales it according to the scale parameter.

TAN

returns the tangent of the variable.

TANH

returns the hyperbolic tangent of the variable.

"hash":{hashPhase}

specifies the parameters to use for the hashing transformation phase.

The hashPhase value can be one or more of the following:

"arguments":{hashArguments}

specifies the arguments for this phase of the transform.

Aliasargs
"nBuckets":integer

specifies the arguments for this phase of the transform.

"method":"BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"

specifies the hash function.

"impute":{imputePhase}

specifies the parameters to use for the imputation phase.

The imputePhase value can be one or more of the following:

"maxRandom":double

specifies the maximum random number to generate.

"method":"MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"

specifies the binning technique to use.

MAX

replaces missing values with the maximum value. This technique applies to interval variables.

MEAN

replaces missing values with the mean. This technique applies to interval variables.

MEDIAN

replaces missing values with the median. This technique applies to interval variables.

MIDRANGE

replaces missing values with the mean of the maximum value and minimum value. This technique applies to interval variables.

MIN

replaces missing values with the minimum value. This technique applies to interval variables.

MODE

replaces missing values with the mode. This technique applies to nominal variables.

RANDOM

replaces missing values with uniform random numbers. This technique applies to interval variables.

VALUE

replaces missing values with the values specified in the valuesInterval and valuesNominal parameters.

"minRandom":double

specifies the minimum random number to generate.

"valuesInterval":[double-1 <, double-2, ...>]

specifies a list of double values for imputation for the interval variables.

AliasvaluesNumeric
"valuesNominal":["string-1" <, "string-2", ...>]

specifies a list of string values for imputation for the nominal variables.

AliasvaluesNonNumeric
"inputs":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

"inputsInheritFormats":True | False

specifies that the variables inherit formats from the underlying table.

DefaultFalse
"mapInterval":{mapIntervalPhase}

specifies the parameters to use for the map to interval transformation phase.

The mapIntervalPhase value can be one or more of the following:

"arguments":{mapIntervalArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The mapIntervalArguments value can be one or more of the following:

"descending":True | False

specifies that the label count encoding be performed.

DefaultTrue
"includeMissingLevel":True | False

when set to True, missing values are included in the distinct level analysis instead of being discarded.

DefaultFalse
"nLevels":integer

specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.

Default2
Range1–10
"nMoments":integer

specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.

Default2
Range1–6
"noise":double

specifies the parameter for the Laplace or uniform noise to be added to the level statistics.

Default0
Minimum value (exclusive)0
"shrinkageFactor":double

specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.

Default0
Minimum value0
"woeAdjust":double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
"woeDefinition":"EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
"method":"COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"

specifies the interval map criterion to use.

Aliastech
DefaultWOE
COUNTTARGET

maps to the counts of the levels of the nominal response.

AliasCOUNT
COUNTINPUT

count encoding of categorical variables.

EMPBAYES

maps to the empirical Bayes.

EVENTPROB

maps to the event probability.

FREQRATIO

maps to the frequency ratio of the levels of the nominal response.

LABELCOUNT

label count encoding of categorical variables.

MAX

maps to the max.

MIN

maps to the min.

MOMENTS

maps to the centralized moments.

WOE

maps to the weight of evidence (WOE).

"name":"string"

specifies a name for the request package.

"outlier":{outlierPhase}

specifies the parameters to use for the outlier determination and treatment phase.

The outlierPhase value can be one or more of the following:

"arguments":{outlierArguments}

specifies the arguments for this phase of the transform.

Aliasargs

The outlierArguments value can be one or more of the following:

"aadLocationUseMean":True | False

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTrue
"location":"BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

DefaultMEAN
BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

"locationBiweightTuning":double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
"lowerPercentile":double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Range(0, 50)
"max":double

specifies a global maximum value.

"min":double

specifies a global minimum value.

"replacements":["BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"] | [double-1 <, double-2, ...>]

specifies the values to use as replacements for outliers. These can be user defined values or location estimates.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
"scale":"AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

DefaultSTD
AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

"scaleBiweightTuning":double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
"scaleMultiplier":double

specifies the multiplying factor for the chosen scale estimator.

"symmetricPercentile":double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Range(0, 100)
"upperPercentile":double

specifies the upper percentile threshold to use.

AliasupperPerc
Range(50, 100)
"userDefinedLimits":[double-1 <, double-2, ...>]

uses the specified user-defined limits as the lower and upper thresholds for each variable.

"zScoreThreshold":double

specifies the Z threshold.

"method":"IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"

specifies the outlier definition.

IQR

uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.

MIQR

uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.

MZSCORE

uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.

PERC

uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.

UDFLIMITS

uses user-defined values to define outliers. Use the min, max, or userDefLims parameters to set the boundaries.

ZSCORE

uses the Z-score to define outliers. Uses the mean as the location, and the standard deviation as the scale estimates.

"treatment":"REPLACE" | "TRIM" | "WINSOR"

specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.

REPLACE

outliers are replaced with user defined values or location estimates.

TRIM

outliers are set to missing and discarded.

WINSOR

outliers are replaced with the lower or upper threshold and then binned.

"output":{outputPhase}

specifies the parameters to use for the output phase.

The outputPhase value can be one or more of the following:

"noScoreCode":True | False

when set to True, no score code is sent to output.

DefaultFalse
"noScoreTable":True | False

when set to True, no scoring is sent to output.

AliasnoScoreTbl
DefaultFalse
"scoreWOE":True | False

when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.

DefaultFalse
"phaseOrder":"FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"

specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.

DefaultIOF
FIO

specifies the phase order: function, impute and outlier.

FOI

specifies the phase order: function, outlier and impute.

IFO

specifies the phase order: impute, function and outlier.

IOF

specifies the phase order: impute, outlier and function.

OFI

specifies the phase order: outlier, function and impute.

OIF

specifies the phase order: outlier, impute and function.

"targets":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies a list of target variables to use.

For more information about specifying the targets parameter, see the common casinvardesc parameter.

AliasevalVars
"targetsInheritFormats":True | False

specifies that the variables inherit formats from the underlying table.

DefaultFalse

sasVarNameLength=True | False

when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.

DefaultFalse

saveState={casouttable}

specifies the settings for an output table that contains the transformation model table.

AliassaveModel
Long formsaveState={"name":"table-name"}
Shortcut formsaveState="table-name"

The casouttable value can be one or more of the following:

"caslib":"string"

specifies the name of the caslib for the output table.

"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data.

"lifetime":64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
"memoryFormat":"DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

"name":"table-name"

specifies the name for the output table.

"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table that has the same name.

DefaultFalse
"tableRedistUpPolicy":"DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

* table={castable}

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

weight="variable-name"

specifies the weight variable.

transform Action

Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.

results <– cas.dataPreprocess.transform(s,
casOut
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
casOutBinDetails
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
casOutLevelBinMap
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
casOutVarTransInfo
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE",
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
code
=list(
casOut
=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
lifetime=64-bit-integer
maxMemSize=64-bit-integer
memoryFormat="DVR" | "INHERIT" | "STANDARD"
name="table-name"
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"
threadBlockSize=64-bit-integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
comment=TRUE | FALSE,
fmtWdth=integer,
indentSize=integer,
labelId=integer,
lineSize=integer,
noTrim=TRUE | FALSE,
tabForm=TRUE | FALSE
),
copyAllVars=TRUE | FALSE,
copyVars=list("variable-name-1" <, "variable-name-2", ...>),
evaluationStats=TRUE | FALSE,
freq="variable-name",
fuzzyCompare=double,
includeInputVars=TRUE | FALSE,
includeMissingGroup=TRUE | FALSE,
maxIterations=integer,
misraGries=TRUE | FALSE,
outputTableOptions
=list(
forceTableReturn=TRUE | FALSE,
tableNames=list("string-1" <, "string-2", ...>)
),
overrides
=list(
alpha=double,
binMapping="LEFT" | "RIGHT",
binMissing=TRUE | FALSE,
binOutliers=TRUE | FALSE,
emptyBins=TRUE | FALSE,
enforceBinaryLevels=TRUE | FALSE,
ivFactor=double,
minNObsInBin=64-bit-integer,
missingBinStats=TRUE | FALSE,
missingEvalNonEvent=TRUE | FALSE,
noDataLowerUpperBound=TRUE | FALSE,
outlierBinsStats=TRUE | FALSE,
woeAdjust=double,
woeDefinition="EVENT" | "NONEVENT"
),
quantileSketch
=list(
epsilon=double
),
requestPackages
=list( list(
catTrans
=list(
arguments
=list(
maxNBins=integer
minNBins=integer
nBinsArray=list(integer-1 <, integer-2, ...>) | integer
overrides
=list(
binMissing=TRUE | FALSE
emptyBins=TRUE | FALSE
enforceBinaryLevels=TRUE | FALSE
ivFactor=double
minNObsInBin=64-bit-integer
missingBinStats=TRUE | FALSE
missingEvalNonEvent=TRUE | FALSE
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
)
preprocessRare=TRUE | FALSE
)
),
dateTime
=list(
method=list("ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR")
),
discretize
=list(
arguments
=list(
binEnds=list(double-1 <, double-2, ...>)
binStarts=list(double-1 <, double-2, ...>)
binWidths=list(double-1 <, double-2, ...>)
cutPoints=list(double-1 <, double-2, ...>)
maxNBins=integer
minNBins=integer
nBinsArray=list(integer-1 <, integer-2, ...>) | integer
overrides
=list(
alpha=double
binMapping="LEFT" | "RIGHT"
binMissing=TRUE | FALSE
binOutliers=TRUE | FALSE
emptyBins=TRUE | FALSE
enforceBinaryLevels=TRUE | FALSE
ivFactor=double
minNObsInBin=64-bit-integer
missingBinStats=TRUE | FALSE
missingEvalNonEvent=TRUE | FALSE
outlierBinsStats=TRUE | FALSE
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
)
)
),
evaluationStats=TRUE | FALSE | {evaluationStatsOptions},
events=list("string-1" <, "string-2", ...>),
featureInteraction
=list(
coefficients=list(double-1 <, double-2, ...>)
inputTransformations=list("string-1" <, "string-2", ...>)
power=integer
),
hash
=list(
arguments
=list(
nBuckets=integer
)
method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
),
impute
=list(
maxRandom=double
minRandom=double
valuesInterval=list(double-1 <, double-2, ...>)
valuesNominal=list("string-1" <, "string-2", ...>)
),
inputs
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
inputsInheritFormats=TRUE | FALSE,
mapInterval
=list(
arguments
=list(
descending=TRUE | FALSE
includeMissingLevel=TRUE | FALSE
nLevels=integer
nMoments=integer
noise=double
woeAdjust=double
woeDefinition="EVENT" | "NONEVENT"
)
),
name="string",
outlier
=list(
arguments
=list(
aadLocationUseMean=TRUE | FALSE
max=double
min=double
replacements=list("BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN") | list(double-1 <, double-2, ...>)
userDefinedLimits=list(double-1 <, double-2, ...>)
)
),
output
=list(
noScoreCode=TRUE | FALSE
noScoreTable=TRUE | FALSE
scoreWOE=TRUE | FALSE
),
targets
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
) <, list(...)>),
sasVarNameLength=TRUE | FALSE,
saveState
=list(
caslib="string",
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
lifetime=64-bit-integer,
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE,
),
seed=integer,
required parameter table
=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
computedVarsProgram="string",
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
orderBy
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression",
whereTable
=list(
casLib="string"
dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | informix-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
required parameter name="table-name"
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>)
where="where-expression"
)
),
tolerance=double,
weight="variable-name"
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOut

—

specifies the settings for an output table.

 casOutBinDetails

—

specifies the settings for an output table that includes information about the binning results.

 casOutLevelBinMap

—

specifies the settings for an output table that contains the nominal level bin mapping information.

 casOutVarTransInfo

—

specifies the settings for an output table that includes information for the variable transformations.

 code

casOut

specifies the settings for generating SAS DATA step scoring code.

 saveState

—

specifies the settings for an output table that contains the transformation model table.

Parameter Descriptions

casOut=list(casouttable)

specifies the settings for an output table.

For more information about specifying the casOut parameter, see the common casouttable parameter.

casOutBinDetails=list(casouttable)

specifies the settings for an output table that includes information about the binning results.

For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.

casOutLevelBinMap=list(casouttable)

specifies the settings for an output table that contains the nominal level bin mapping information.

For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.

casOutVarTransInfo=list(casouttable)

specifies the settings for an output table that includes information for the variable transformations.

For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.

code=list(codegen)

specifies the settings for generating SAS DATA step scoring code.

For more information about specifying the code parameter, see the common codegen parameter.

copyAllVars=TRUE | FALSE

when set to True, all the variables from the input table are copied to the scored output table.

AliasallIdVars
DefaultFALSE

copyVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.

distinctCountLimit=integer

specifies the distinct count limit.

evaluationStats=TRUE | FALSE

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats
DefaultFALSE

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05

includeInputVars=TRUE | FALSE

when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.

DefaultFALSE

includeMissingGroup=TRUE | FALSE

when set to True, missing values are allowed as group-by keys.

DefaultFALSE

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

misraGries=TRUE | FALSE

specifies that the Misra-Gries algorithm be used for most frequent estimation.

DefaultFALSE

outputTableOptions=list(outputTableOptions)

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliastblOpts

The outputTableOptions value can be one or more of the following:

forceTableReturn=TRUE | FALSE

when set to True, result tables are returned to the client even if the output is also saved as an output table.

DefaultFALSE
tableNames=list("string-1" <, "string-2", ...>)

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

overrides=list(globalOverrides)

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

alpha=double

specifies the significance level.

Default0.05
binMapping="LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
binMissing=TRUE | FALSE

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFALSE
binOutliers=TRUE | FALSE

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

DefaultFALSE
emptyBins=TRUE | FALSE

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFALSE
enforceBinaryLevels=TRUE | FALSE

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTRUE
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=TRUE | FALSE

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTRUE
missingEvalNonEvent=TRUE | FALSE

when set to True, missing values of the target variables are considered as non-event values.

DefaultFALSE
noDataLowerUpperBound=TRUE | FALSE

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

DefaultFALSE
outlierBinsStats=TRUE | FALSE

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

DefaultTRUE
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch=list(quantileSketchOptions)

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

compressionFactor=double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
epsilon=double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

rank=list(rankOptions)

specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.

Long formrank=list(intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE")
Shortcut formrank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

The rankOptions value can be one or more of the following:

intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"

specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

AD

Anderson-Darling statistic.

AVGQUANKURT

average quantile skewness.

AVGQUANSKEW

average quantile kurtosis.

CLASSICALKURT

moment kurtosis.

CLASSICALSKEW

moment skewness.

CVM

Cramer-von Mises statistic.

KS

Kolmogorov-Smirnov statistic.

PEARSON

Pearson correlation.

VARIANCE

variance.

nominalStat="CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"

specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.

CHISQ

chi-square statistic.

CRAMERSV

Cramer's V

FTEST

F-test statistic.

G2

g2 statistic.

GINI

gini Index.

IV

information value.

WELCHTTEST

Welch's t-test statistic.

WOE

weight of evidence.

topKInteractions=integer
Default10
Minimum value1
topKSave=integer
Default1
Minimum value1

requestPackages=list( list(transformRequestPackage-1) <, list(transformRequestPackage-2), ...>)

specifies an array of transform request packages to be processed by the action.

Aliasespipelines
reqPacks

The transformRequestPackage value can be one or more of the following:

catTrans=list(catTransPhase)

specifies the parameters to use for the categorical transformation phase.

The catTransPhase value can be one or more of the following:

arguments=list(catTransArguments)

specifies the arguments for this phase of the transform.

Aliasargs

The catTransArguments value can be one or more of the following:

contingencyTblOpts=list(contingencyTableOptions)

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

inputsNLevels=integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
inputsRawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
maxNBins=integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

minNBins=integer

specifies the minimum number of bins.

nBinsArray=list(integer-1 <, integer-2, ...>) | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
overrides=list(globalOverrides)

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

AliasesmiscellaneousOpts
opts

The globalOverrides value can be one or more of the following:

binMissing=TRUE | FALSE

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFALSE
emptyBins=TRUE | FALSE

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFALSE
enforceBinaryLevels=TRUE | FALSE

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTRUE
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=TRUE | FALSE

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTRUE
missingEvalNonEvent=TRUE | FALSE

when set to True, missing values of the target variables are considered as non-event values.

DefaultFALSE
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
preprocessRare=TRUE | FALSE

when set to True, rare levels are grouped into a single group at the start of the grouping process.

Aliaspreprocess
DefaultFALSE
rareThreshold=integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
rareThresholdPercent=double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Default5
Range(0, 100)
treeCrit="ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

DefaultGAINRATIO
ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
method="DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"

specifies the binning technique to use.

DTREE

groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.

GROUPRARE

groups rare levels of the analysis variable. This is an unsupervised technique. If you do not specify one of maxNLevels, rareFreqCutOff, or rareThresholdPer, then rareThresholdPer is set to 5.

ONEHOT

one hot encoding of categorical variables.

AliasLABEL
RTREE

groups based on a one-level regression tree. The criterion is the sum of squared error (SSE).

WOE

groups based on maximizing the information value (IV).

dateTime=list(dateTimePhase)

specifies the parameters to use for the date-time transformation phase.

The dateTimePhase value can be one or more of the following:

inputType="DATE" | "DATETIME" | "TIME"

input variable type as one of date, time or datetime.

DATE

date variable type

DATETIME

datetime variable type

TIME

time variable type

method=list("ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR")

specifies the binning technique to use.

Aliastech
ALL

all (year, day, month, ...) of a datetime variable.

DAYMONTH

day of the month of a datetime variable.

DAYWEEK

day of the week of a datetime variable.

HOUR

hour of a datetime variable.

LEAPYEAR

datetime is a leap year or not.

MINUTE

minute of a datetime variable.

MONTH

month of a datetime variable.

QUARTER

quarter of the year of a datetime variable.

WEEK

week of a datetime variable.

WEEKEND

datetime is a weekend or not.

YEAR

year of a datetime variable.

discretize=list(discretizePhase)

specifies the parameters to use for the discretization phase.

The discretizePhase value can be one or more of the following:

arguments=list(discretizeArguments)

specifies the arguments for this phase of the transform.

Aliasargs

The discretizeArguments value can be one or more of the following:

binEnds=list(double-1 <, double-2, ...>)

specifies the bin end values. If applicable, they override the data maximum values.

AliasbinEnd
binStarts=list(double-1 <, double-2, ...>)

specifies the bin start values. If applicable, they override the data minimum values.

AliasbinStart
binWidths=list(double-1 <, double-2, ...>)

specifies the bin width.

AliasbinWidth
contingencyTblOpts=list(contingencyTableOptions)

controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.

AliascTblOpts

The contingencyTableOptions value can be one or more of the following:

inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"

specifies the method for determining the levels of the transformation variable.

BUCKET

generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.

DISTINCTLEVELS

generates levels that map to the distinct values of the target variable.

AliasesCLASS
LEVEL
QUANTILE

generates levels that map to the bins of quantile binning. Specify the number of bins with the evalNLevels parameter.

RAW

generates levels directly from the input value. The target variable must be numeric. For nonzero starting values, specify a starting value with the evalRawLevelStartingVal parameter. This is appropriate for ordinal variables.

inputsNLevels=integer

specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.

AliasnInitBins
inputsRawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
cutPoints=list(double-1 <, double-2, ...>)

specifies the user-provided cutpoints, for the CUTPTS binning technique.

AliascutPts
maxNBins=integer

specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.

Default5
Minimum value (exclusive)0
minNBins=integer

specifies the minimum number of bins.

Default1
Minimum value (exclusive)0
nBinsArray=list(integer-1 <, integer-2, ...>) | integer

specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.

AliasnBins
overrides=list(globalOverrides)

specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.

The globalOverrides value can be one or more of the following:

alpha=double

specifies the significance level.

Default0.05
binMapping="LEFT" | "RIGHT"

controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.

DefaultRIGHT
binMissing=TRUE | FALSE

when set to True, bins missing values into a separate bin. The ID for this bin is 0.

AliasmapMissing
DefaultFALSE
binOutliers=TRUE | FALSE

when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.

DefaultFALSE
emptyBins=TRUE | FALSE

when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.

DefaultFALSE
enforceBinaryLevels=TRUE | FALSE

when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.

DefaultTRUE
ivFactor=double

specifies the information value adjustment factor.

Default2
minNObsInBin=64-bit-integer

specifies the minimum number of observations to include in a bin.

AliasleafSize
minPerNObsInBin=double

specifies the minimum percentage of all observations to include in a bin.

Default5
Range(0, 50)
missingBinStats=TRUE | FALSE

when set to True, the missing bin is considered during the computation of the evaluation statistics.

DefaultTRUE
missingEvalNonEvent=TRUE | FALSE

when set to True, missing values of the target variables are considered as non-event values.

DefaultFALSE
noDataLowerUpperBound=TRUE | FALSE

when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.

DefaultFALSE
outlierBinsStats=TRUE | FALSE

when set to True, the outlier bins are considered during the computation of the evaluation statistics.

DefaultTRUE
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
treeCrit="ENTROPY" | "GAINRATIO" | "GINI" | "RSS"

specifies the tree criterion to use.

ENTROPY

information gain.

AliasGAIN
GAINRATIO

information gain ratio.

AliasIGR
GINI

average gini index.

RSS

sum of squared error (SSE).

AliasSSE
method="BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"

specifies the binning technique to use.

BUCKET

creates equal-width bins.

CACC

creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.

CAIM

creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.

CHIMERGE

creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.

CUTPTS

creates bins according to the user-specified cutpoints.

DTREE

creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.

MDLP

creates bins based on the minimum description length. This is a top-down supervised discretization technique.

QUANTILE

creates equal-frequency bins.

RTREE

creates bins based on a one-level regression tree. This is a top-down supervised discretization technique.

WOE

creates bins based on WOE criterion. This is a top-down supervised discretization technique.

evaluationStats=TRUE | FALSE | {evaluationStatsOptions}

when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.

AliasevalStats

The evaluationStatsOptions value can be one or more of the following:

chiSqGroup=TRUE | FALSE
when set to True, the Chi-Square, G2 and Cramer's V statistics are computed.
DefaultFALSE
ftTestGroup=TRUE | FALSE
when set to True, F-test and Welch's T-test statistics are computed.
DefaultFALSE
missIndicatorTarget=TRUE | FALSE
when specified, the target is transformed to a missing indicator binary target.
DefaultFALSE
nominalTarget=TRUE | FALSE
when set to True, the target variables are considered as nominal.
DefaultFALSE
woeGroup=TRUE | FALSE
when set to True, the WOE, IV and Gini index statistics are computed.
DefaultFALSE
events=list("string-1" <, "string-2", ...>)

specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.

AliasevalVarsEvents
featureInteraction=list(featureInteraction)

options that control the generation of interaction features.

Aliasinteraction

The featureInteraction value can be one or more of the following:

coefficients=list(double-1 <, double-2, ...>)

specifies the coefficients for the linear interaction operator.

inputTransformations=list("string-1" <, "string-2", ...>)

specifies the transformations that are to be used for generating the input component features of the interaction features.

Aliasinputs
method="CROSS" | "ORDERED"

specifies the method to use for feature generation. These include the cross and ordered component feature construction methods.

DefaultCROSS
CROSS

Cross product

ORDERED

Ordered grouping.

power=integer

specifies the number of inputs for the polynomial feature interaction operator.

Default2
Range1–4
synthesizer="DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"

specifies the operator to use for composing the values of the component feature interaction.

DIVISION

Division

LINEAR

Linear combination

MULTIPLICATION

Multiplication

NOMINAL

Nominal interaction

POLYNOMIAL

Polynomial feature generation

targetTransformation="string"

specifies the transformation that is to be used for generating the target features for interaction feature generation.

Aliasestargets
target
featureProbe=list(featureProbePhase)

specifies the parameters to use for the feature probe transformation phase.

The featureProbePhase value can be one or more of the following:

arguments=list(featureProbeArguments)

specifies the arguments for this phase of the transform.

Aliasargs

The featureProbeArguments value can be one or more of the following:

ecdfTolerance=double

specifies the tolerance value for the empirical cumulative distribution function.

Default0.001
Range1E-06–0.1
nProbes=integer

the number of feature probes.

Default1
Minimum value1
probeMissing=TRUE | FALSE

when set to True, generates missing values at the observed missing rate.

DefaultTRUE
rareThreshold=integer

specifies the rare frequency threshold.

AliasrareFreqCutOff
Minimum value (exclusive)0
rareThresholdPercent=double

specifies the rare threshold percentile. Levels less than the threshold are grouped together.

Range(0, 100)
rawLevelStartingValue=integer

specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.

Default0
Minimum value0
shrinkageFactor=double

specifies the shrinkage factor for level probability estimation.

Default0
Minimum value0
useRawLevel=TRUE | FALSE

specifies that the raw values be used as levels of the nominal variable.

DefaultFALSE
method="ECDF" | "FREQUENCY"

specifies the feature probe method to use.

ECDF

creates feature probes using the empirical cumulative distribution function.

FREQUENCY

creates feature probes using the empirical frequency distribution.

function=list(functionPhase)

specifies the parameters to use for the functional transformation phase.

The functionPhase value can be one or more of the following:

arguments=list(functionArguments)

specifies the arguments for this phase of the transform.

Aliasargs

The functionArguments value can be one or more of the following:

aadLocationUseMean=TRUE | FALSE

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTRUE
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
lowerPercentile=double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Default10
Range(0, 50)
otherArguments=list(double-1 <, double-2, ...>)

specifies other values to use. The values depend on the type of functional transformation.

scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
scaleMultiplier=double

specifies the multiplying factor for the chosen scale estimator.

AliasscaleMulFac
shiftMax=double

specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value

AliasshiftNegative
shiftMin=double

specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value

AliasshiftPositive
symmetricPercentile=double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Default10
Range(0, 100)
upperPercentile=double

specifies the upper percentile threshold to use.

AliasupperPerc
Default90
Range(50, 100)
method="ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"

specifies the functional transformation.

ABS

returns the absolute value of the variable.

ARCSIN

returns the arcsine of the square root of a variable.

BOXCOX

returns the Box-Cox transformation of the variable.

CENTER

returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.

COS

returns the cosine of the variable.

COSH

returns the hyperbolic cosine of the variable.

EXP

returns the value of the e constant raised to the value of the variable.

IDENTITY

returns the value of a variable unmodified.

INVERSE

returns the inverse of a variable (1/x).

INVSQUARESHIFT

returns the inverse square shift value.

LOG

returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.

POWER

returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.

RANGE

returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.

SCALESHIFT

returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.

SIN

returns the sine of the variable.

SINH

returns the hyperbolic sine of the variable.

SQRT

returns the square root of the variable.

STANDARDIZE

returns the value minus the location determined from the loc parameter, and then scales it according to the scale parameter.

TAN

returns the tangent of the variable.

TANH

returns the hyperbolic tangent of the variable.

hash=list(hashPhase)

specifies the parameters to use for the hashing transformation phase.

The hashPhase value can be one or more of the following:

arguments=list(hashArguments)

specifies the arguments for this phase of the transform.

Aliasargs
nBuckets=integer

specifies the arguments for this phase of the transform.

method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"

specifies the hash function.

impute=list(imputePhase)

specifies the parameters to use for the imputation phase.

The imputePhase value can be one or more of the following:

maxRandom=double

specifies the maximum random number to generate.

method="MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"

specifies the binning technique to use.

MAX

replaces missing values with the maximum value. This technique applies to interval variables.

MEAN

replaces missing values with the mean. This technique applies to interval variables.

MEDIAN

replaces missing values with the median. This technique applies to interval variables.

MIDRANGE

replaces missing values with the mean of the maximum value and minimum value. This technique applies to interval variables.

MIN

replaces missing values with the minimum value. This technique applies to interval variables.

MODE

replaces missing values with the mode. This technique applies to nominal variables.

RANDOM

replaces missing values with uniform random numbers. This technique applies to interval variables.

VALUE

replaces missing values with the values specified in the valuesInterval and valuesNominal parameters.

minRandom=double

specifies the minimum random number to generate.

valuesInterval=list(double-1 <, double-2, ...>)

specifies a list of double values for imputation for the interval variables.

AliasvaluesNumeric
valuesNominal=list("string-1" <, "string-2", ...>)

specifies a list of string values for imputation for the nominal variables.

AliasvaluesNonNumeric
inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

inputsInheritFormats=TRUE | FALSE

specifies that the variables inherit formats from the underlying table.

DefaultFALSE
mapInterval=list(mapIntervalPhase)

specifies the parameters to use for the map to interval transformation phase.

The mapIntervalPhase value can be one or more of the following:

arguments=list(mapIntervalArguments)

specifies the arguments for this phase of the transform.

Aliasargs

The mapIntervalArguments value can be one or more of the following:

descending=TRUE | FALSE

specifies that the label count encoding be performed.

DefaultTRUE
includeMissingLevel=TRUE | FALSE

when set to True, missing values are included in the distinct level analysis instead of being discarded.

DefaultFALSE
nLevels=integer

specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.

Default2
Range1–10
nMoments=integer

specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.

Default2
Range1–6
noise=double

specifies the parameter for the Laplace or uniform noise to be added to the level statistics.

Default0
Minimum value (exclusive)0
shrinkageFactor=double

specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.

Default0
Minimum value0
woeAdjust=double

specifies the weight of evidence (WOE) adjustment factor.

Default0.5
woeDefinition="EVENT" | "NONEVENT"

specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.

DefaultNONEVENT
method="COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"

specifies the interval map criterion to use.

Aliastech
DefaultWOE
COUNTTARGET

maps to the counts of the levels of the nominal response.

AliasCOUNT
COUNTINPUT

count encoding of categorical variables.

EMPBAYES

maps to the empirical Bayes.

EVENTPROB

maps to the event probability.

FREQRATIO

maps to the frequency ratio of the levels of the nominal response.

LABELCOUNT

label count encoding of categorical variables.

MAX

maps to the max.

MIN

maps to the min.

MOMENTS

maps to the centralized moments.

WOE

maps to the weight of evidence (WOE).

name="string"

specifies a name for the request package.

outlier=list(outlierPhase)

specifies the parameters to use for the outlier determination and treatment phase.

The outlierPhase value can be one or more of the following:

arguments=list(outlierArguments)

specifies the arguments for this phase of the transform.

Aliasargs

The outlierArguments value can be one or more of the following:

aadLocationUseMean=TRUE | FALSE

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTRUE
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"

specifies the estimation method of location.

DefaultMEAN
BIWEIGHT

uses Tukey biweight based estimate for location.

GEOMETRICMEAN

uses the geometric mean for location.

HARMONICMEAN

uses the harmonic mean for location.

MEAN

uses the arithmetic mean for location.

MEDIAN

uses the median value for location.

TRIMMEDMEAN

uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

WINSORIZEDMEAN

uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.

locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Minimum value (exclusive)0
lowerPercentile=double

specifies the lower percentile threshold (PERC outlier definition).

AliaslowerPerc
Range(0, 50)
max=double

specifies a global maximum value.

min=double

specifies a global minimum value.

replacements=list("BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN") | list(double-1 <, double-2, ...>)

specifies the values to use as replacements for outliers. These can be user defined values or location estimates.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"

specifies the scale method to use.

DefaultSTD
AAD

uses the absolute deviation about the mean or median as scale.

BIWEIGHT

uses Tukey biweight based estimate for scale.

GINI

uses the Gini scale for scale.

IQR

uses the inter-quartile range for scale.

MAD

uses the median absolute deviation about the median for scale.

STD

uses the standard deviation for scale.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Minimum value (exclusive)0
scaleMultiplier=double

specifies the multiplying factor for the chosen scale estimator.

symmetricPercentile=double

specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliassymPerc
Range(0, 100)
upperPercentile=double

specifies the upper percentile threshold to use.

AliasupperPerc
Range(50, 100)
userDefinedLimits=list(double-1 <, double-2, ...>)

uses the specified user-defined limits as the lower and upper thresholds for each variable.

zScoreThreshold=double

specifies the Z threshold.

method="IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"

specifies the outlier definition.

IQR

uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.

MIQR

uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.

MZSCORE

uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.

PERC

uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.

UDFLIMITS

uses user-defined values to define outliers. Use the min, max, or userDefLims parameters to set the boundaries.

ZSCORE

uses the Z-score to define outliers. Uses the mean as the location, and the standard deviation as the scale estimates.

treatment="REPLACE" | "TRIM" | "WINSOR"

specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.

REPLACE

outliers are replaced with user defined values or location estimates.

TRIM

outliers are set to missing and discarded.

WINSOR

outliers are replaced with the lower or upper threshold and then binned.

output=list(outputPhase)

specifies the parameters to use for the output phase.

The outputPhase value can be one or more of the following:

noScoreCode=TRUE | FALSE

when set to True, no score code is sent to output.

DefaultFALSE
noScoreTable=TRUE | FALSE

when set to True, no scoring is sent to output.

AliasnoScoreTbl
DefaultFALSE
scoreWOE=TRUE | FALSE

when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.

DefaultFALSE
phaseOrder="FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"

specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.

DefaultIOF
FIO

specifies the phase order: function, impute and outlier.

FOI

specifies the phase order: function, outlier and impute.

IFO

specifies the phase order: impute, function and outlier.

IOF

specifies the phase order: impute, outlier and function.

OFI

specifies the phase order: outlier, function and impute.

OIF

specifies the phase order: outlier, impute and function.

targets=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies a list of target variables to use.

For more information about specifying the targets parameter, see the common casinvardesc parameter.

AliasevalVars
targetsInheritFormats=TRUE | FALSE

specifies that the variables inherit formats from the underlying table.

DefaultFALSE

sasVarNameLength=TRUE | FALSE

when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.

DefaultFALSE

saveState=list(casouttable)

specifies the settings for an output table that contains the transformation model table.

AliassaveModel
Long formsaveState=list(name="table-name")
Shortcut formsaveState="table-name"

The casouttable value can be one or more of the following:

caslib="string"

specifies the name of the caslib for the output table.

indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data.

lifetime=64-bit-integer

specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.

Default0
Minimum value0
memoryFormat="DVR" | "INHERIT" | "STANDARD"

specifies the memory format for the output table.

DefaultINHERIT
DVR

use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.

INHERIT

use the default memory format that is set for the server. By default, the server uses the standard memory format. If an administrator sets the CAS_DEFAULT_MEMORY_FORMAT environment variable to DVR, then the DVR memory format is set as the default for the server.

STANDARD

use the standard memory format.

name="table-name"

specifies the name for the output table.

promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table that has the same name.

DefaultFALSE
tableRedistUpPolicy="DEFER" | "NOREDIST" | "REBALANCE"

Specifies the Table Redistribution Policy when the number of worker pods increases on a running CAS server.

DEFER

Defer redistribution policy selection to higher-level entity.

NOREDIST

Do not redistribute table data when the number of worker pods changes on a running CAS server.

REBALANCE

Rebalance table data when the number of worker pods changes on a running CAS server.

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

* table=list(castable)

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

weight="variable-name"

specifies the weight variable.

Last updated: March 19, 2025