Data Preprocess Action Set: Syntax
Provides actions for data preprocessing and transformation
transform Action
Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. | |
|
— |
specifies the settings for an output table that includes information about the binning results. | |
|
— |
specifies the settings for an output table that contains the nominal level bin mapping information. | |
|
— |
specifies the settings for an output table that includes information for the variable transformations. | |
|
casOut |
specifies the settings for generating SAS DATA step scoring code. | |
|
— |
specifies the settings for an output table that contains the transformation model table. |
Parameter Descriptions
casOut={casouttable}
specifies the settings for an output table.
For more information about specifying the casOut parameter, see the common casouttable parameter.
casOutBinDetails={casouttable}
specifies the settings for an output table that includes information about the binning results.
For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.
casOutLevelBinMap={casouttable}
specifies the settings for an output table that contains the nominal level bin mapping information.
For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.
casOutVarTransInfo={casouttable}
specifies the settings for an output table that includes information for the variable transformations.
For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.
code={codegen}
specifies the settings for generating SAS DATA step scoring code.
For more information about specifying the code parameter, see the common codegen parameter.
copyAllVars=TRUE | FALSE
when set to True, all the variables from the input table are copied to the scored output table.
| Alias | allIdVars |
|---|---|
| Default | FALSE |
copyVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.
distinctCountLimit=integer
specifies the distinct count limit.
evaluationStats=TRUE | FALSE
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|---|
| Default | FALSE |
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
includeInputVars=TRUE | FALSE
when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.
| Default | FALSE |
|---|
includeMissingGroup=TRUE | FALSE
when set to True, missing values are allowed as group-by keys.
| Default | FALSE |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
misraGries=TRUE | FALSE
specifies that the Misra-Gries algorithm be used for most frequent estimation.
| Default | FALSE |
|---|
outputTableOptions={outputTableOptions}
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | tblOpts |
|---|
The outputTableOptions value can be one or more of the following:
forceTableReturn=TRUE | FALSE
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | FALSE |
|---|
tableNames={"string-1" <, "string-2", ...>}
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
alpha=double
specifies the significance level.
| Default | 0.05 |
|---|
binMapping="LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
binMissing=TRUE | FALSE
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | FALSE |
binOutliers=TRUE | FALSE
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | FALSE |
|---|
emptyBins=TRUE | FALSE
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | FALSE |
|---|
enforceBinaryLevels=TRUE | FALSE
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | TRUE |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=TRUE | FALSE
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
missingEvalNonEvent=TRUE | FALSE
when set to True, missing values of the target variables are considered as non-event values.
| Default | FALSE |
|---|
noDataLowerUpperBound=TRUE | FALSE
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | FALSE |
|---|
outlierBinsStats=TRUE | FALSE
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch={quantileSketchOptions}
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
compressionFactor=double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
epsilon=double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
rank={rankOptions}
specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.
| Long form | rank={intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"} |
|---|---|
| Shortcut form | rank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE" |
The rankOptions value can be one or more of the following:
intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"
specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
nominalStat="CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"
specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
topKInteractions=integer
| Default | 10 |
|---|---|
| Minimum value | 1 |
topKSave=integer
| Default | 1 |
|---|---|
| Minimum value | 1 |
requestPackages={{transformRequestPackage-1} <, {transformRequestPackage-2}, ...>}
specifies an array of transform request packages to be processed by the action.
| Aliases | pipelines |
|---|---|
| reqPacks |
The transformRequestPackage value can be one or more of the following:
catTrans={catTransPhase}
specifies the parameters to use for the categorical transformation phase.
The catTransPhase value can be one or more of the following:
arguments={catTransArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The catTransArguments value can be one or more of the following:
contingencyTblOpts={contingencyTableOptions}
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
inputsNLevels=integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
inputsRawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
maxNBins=integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
minNBins=integer
specifies the minimum number of bins.
nBinsArray={integer-1 <, integer-2, ...>} | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
| Aliases | miscellaneousOpts |
|---|---|
| opts |
The globalOverrides value can be one or more of the following:
binMissing=TRUE | FALSE
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | FALSE |
emptyBins=TRUE | FALSE
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | FALSE |
|---|
enforceBinaryLevels=TRUE | FALSE
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | TRUE |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=TRUE | FALSE
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
missingEvalNonEvent=TRUE | FALSE
when set to True, missing values of the target variables are considered as non-event values.
| Default | FALSE |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
preprocessRare=TRUE | FALSE
when set to True, rare levels are grouped into a single group at the start of the grouping process.
| Alias | preprocess |
|---|---|
| Default | FALSE |
rareThreshold=integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
rareThresholdPercent=double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Default | 5 |
|---|---|
| Range | (0, 100) |
method="DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"
specifies the binning technique to use.
DTREE
groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.
dateTime={dateTimePhase}
specifies the parameters to use for the date-time transformation phase.
The dateTimePhase value can be one or more of the following:
inputType="DATE" | "DATETIME" | "TIME"
method={"ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"}
specifies the binning technique to use.
| Alias | tech |
|---|
discretize={discretizePhase}
specifies the parameters to use for the discretization phase.
The discretizePhase value can be one or more of the following:
arguments={discretizeArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The discretizeArguments value can be one or more of the following:
binEnds={double-1 <, double-2, ...>}
specifies the bin end values. If applicable, they override the data maximum values.
| Alias | binEnd |
|---|
binStarts={double-1 <, double-2, ...>}
specifies the bin start values. If applicable, they override the data minimum values.
| Alias | binStart |
|---|
binWidths={double-1 <, double-2, ...>}
specifies the bin width.
| Alias | binWidth |
|---|
contingencyTblOpts={contingencyTableOptions}
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
inputsNLevels=integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
inputsRawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
cutPoints={double-1 <, double-2, ...>}
specifies the user-provided cutpoints, for the CUTPTS binning technique.
| Alias | cutPts |
|---|
maxNBins=integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
| Default | 5 |
|---|---|
| Minimum value (exclusive) | 0 |
minNBins=integer
specifies the minimum number of bins.
| Default | 1 |
|---|---|
| Minimum value (exclusive) | 0 |
nBinsArray={integer-1 <, integer-2, ...>} | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
alpha=double
specifies the significance level.
| Default | 0.05 |
|---|
binMapping="LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
binMissing=TRUE | FALSE
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | FALSE |
binOutliers=TRUE | FALSE
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | FALSE |
|---|
emptyBins=TRUE | FALSE
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | FALSE |
|---|
enforceBinaryLevels=TRUE | FALSE
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | TRUE |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=TRUE | FALSE
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
missingEvalNonEvent=TRUE | FALSE
when set to True, missing values of the target variables are considered as non-event values.
| Default | FALSE |
|---|
noDataLowerUpperBound=TRUE | FALSE
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | FALSE |
|---|
outlierBinsStats=TRUE | FALSE
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
method="BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"
specifies the binning technique to use.
CACC
creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.
CAIM
creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.
CHIMERGE
creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.
DTREE
creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.
MDLP
creates bins based on the minimum description length. This is a top-down supervised discretization technique.
evaluationStats=TRUE | FALSE | {evaluationStatsOptions}
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|
The evaluationStatsOptions value can be one or more of the following:
chiSqGroup=TRUE | FALSE
| Default | FALSE |
|---|
ftTestGroup=TRUE | FALSE
| Default | FALSE |
|---|
missIndicatorTarget=TRUE | FALSE
| Default | FALSE |
|---|
nominalTarget=TRUE | FALSE
| Default | FALSE |
|---|
woeGroup=TRUE | FALSE
| Default | FALSE |
|---|
events={"string-1" <, "string-2", ...>}
specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.
| Alias | evalVarsEvents |
|---|
featureInteraction={featureInteraction}
options that control the generation of interaction features.
| Alias | interaction |
|---|
The featureInteraction value can be one or more of the following:
coefficients={double-1 <, double-2, ...>}
specifies the coefficients for the linear interaction operator.
inputTransformations={"string-1" <, "string-2", ...>}
specifies the transformations that are to be used for generating the input component features of the interaction features.
| Alias | inputs |
|---|
method="CROSS" | "ORDERED"
power=integer
specifies the number of inputs for the polynomial feature interaction operator.
| Default | 2 |
|---|---|
| Range | 1–4 |
synthesizer="DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"
targetTransformation="string"
specifies the transformation that is to be used for generating the target features for interaction feature generation.
| Aliases | targets |
|---|---|
| target |
featureProbe={featureProbePhase}
specifies the parameters to use for the feature probe transformation phase.
The featureProbePhase value can be one or more of the following:
arguments={featureProbeArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The featureProbeArguments value can be one or more of the following:
ecdfTolerance=double
specifies the tolerance value for the empirical cumulative distribution function.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
nProbes=integer
the number of feature probes.
| Default | 1 |
|---|---|
| Minimum value | 1 |
probeMissing=TRUE | FALSE
when set to True, generates missing values at the observed missing rate.
| Default | TRUE |
|---|
rareThreshold=integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
rareThresholdPercent=double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Range | (0, 100) |
|---|
rawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
shrinkageFactor=double
specifies the shrinkage factor for level probability estimation.
| Default | 0 |
|---|---|
| Minimum value | 0 |
useRawLevel=TRUE | FALSE
specifies that the raw values be used as levels of the nominal variable.
| Default | FALSE |
|---|
function={functionPhase}
specifies the parameters to use for the functional transformation phase.
The functionPhase value can be one or more of the following:
arguments={functionArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The functionArguments value can be one or more of the following:
aadLocationUseMean=TRUE | FALSE
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | TRUE |
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
lowerPercentile=double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Default | 10 |
| Range | (0, 50) |
otherArguments={double-1 <, double-2, ...>}
specifies other values to use. The values depend on the type of functional transformation.
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
scaleMultiplier=double
specifies the multiplying factor for the chosen scale estimator.
| Alias | scaleMulFac |
|---|
shiftMax=double
specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value
| Alias | shiftNegative |
|---|
shiftMin=double
specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value
| Alias | shiftPositive |
|---|
symmetricPercentile=double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
upperPercentile=double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Default | 90 |
| Range | (50, 100) |
method="ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"
specifies the functional transformation.
CENTER
returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.
LOG
returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.
POWER
returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.
RANGE
returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.
SCALESHIFT
returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.
hash={hashPhase}
specifies the parameters to use for the hashing transformation phase.
The hashPhase value can be one or more of the following:
arguments={hashArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
nBuckets=integer
specifies the arguments for this phase of the transform.
method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
specifies the hash function.
impute={imputePhase}
specifies the parameters to use for the imputation phase.
The imputePhase value can be one or more of the following:
maxRandom=double
specifies the maximum random number to generate.
method="MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"
minRandom=double
specifies the minimum random number to generate.
valuesInterval={double-1 <, double-2, ...>}
specifies a list of double values for imputation for the interval variables.
| Alias | valuesNumeric |
|---|
valuesNominal={"string-1" <, "string-2", ...>}
specifies a list of string values for imputation for the nominal variables.
| Alias | valuesNonNumeric |
|---|
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
inputsInheritFormats=TRUE | FALSE
specifies that the variables inherit formats from the underlying table.
| Default | FALSE |
|---|
mapInterval={mapIntervalPhase}
specifies the parameters to use for the map to interval transformation phase.
The mapIntervalPhase value can be one or more of the following:
arguments={mapIntervalArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The mapIntervalArguments value can be one or more of the following:
descending=TRUE | FALSE
specifies that the label count encoding be performed.
| Default | TRUE |
|---|
includeMissingLevel=TRUE | FALSE
when set to True, missing values are included in the distinct level analysis instead of being discarded.
| Default | FALSE |
|---|
nLevels=integer
specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.
| Default | 2 |
|---|---|
| Range | 1–10 |
nMoments=integer
specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.
| Default | 2 |
|---|---|
| Range | 1–6 |
noise=double
specifies the parameter for the Laplace or uniform noise to be added to the level statistics.
| Default | 0 |
|---|---|
| Minimum value (exclusive) | 0 |
shrinkageFactor=double
specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.
| Default | 0 |
|---|---|
| Minimum value | 0 |
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
method="COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"
specifies the interval map criterion to use.
| Alias | tech |
|---|---|
| Default | WOE |
name="string"
specifies a name for the request package.
outlier={outlierPhase}
specifies the parameters to use for the outlier determination and treatment phase.
The outlierPhase value can be one or more of the following:
arguments={outlierArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The outlierArguments value can be one or more of the following:
aadLocationUseMean=TRUE | FALSE
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | TRUE |
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
lowerPercentile=double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Range | (0, 50) |
max=double
specifies a global maximum value.
min=double
specifies a global minimum value.
replacements={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {double-1 <, double-2, ...>}
specifies the values to use as replacements for outliers. These can be user defined values or location estimates.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
| Default | STD |
|---|
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
scaleMultiplier=double
specifies the multiplying factor for the chosen scale estimator.
symmetricPercentile=double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Range | (0, 100) |
upperPercentile=double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Range | (50, 100) |
userDefinedLimits={double-1 <, double-2, ...>}
uses the specified user-defined limits as the lower and upper thresholds for each variable.
zScoreThreshold=double
specifies the Z threshold.
method="IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"
specifies the outlier definition.
IQR
uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.
MIQR
uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.
MZSCORE
uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.
PERC
uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.
treatment="REPLACE" | "TRIM" | "WINSOR"
specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.
output={outputPhase}
specifies the parameters to use for the output phase.
The outputPhase value can be one or more of the following:
noScoreCode=TRUE | FALSE
when set to True, no score code is sent to output.
| Default | FALSE |
|---|
noScoreTable=TRUE | FALSE
when set to True, no scoring is sent to output.
| Alias | noScoreTbl |
|---|---|
| Default | FALSE |
scoreWOE=TRUE | FALSE
when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.
| Default | FALSE |
|---|
phaseOrder="FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"
specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.
| Default | IOF |
|---|
targets={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies a list of target variables to use.
For more information about specifying the targets parameter, see the common casinvardesc parameter.
| Alias | evalVars |
|---|
targetsInheritFormats=TRUE | FALSE
specifies that the variables inherit formats from the underlying table.
| Default | FALSE |
|---|
sasVarNameLength=TRUE | FALSE
when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.
| Default | FALSE |
|---|
saveState={casouttable}
specifies the settings for an output table that contains the transformation model table.
| Alias | saveModel |
|---|
| Long form | saveState={name="table-name"} |
|---|---|
| Shortcut form | saveState="table-name" |
The casouttable value can be one or more of the following:
caslib="string"
specifies the name of the caslib for the output table.
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data.
lifetime=64-bit-integer
specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.
| Default | 0 |
|---|---|
| Minimum value | 0 |
memoryFormat="DVR" | "INHERIT" | "STANDARD"
specifies the memory format for the output table.
| Default | INHERIT |
|---|
DVR
use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.
name="table-name"
specifies the name for the output table.
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
|---|
replace=TRUE | FALSE
when set to True, overwrites an existing table that has the same name.
| Default | FALSE |
|---|
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
* table={castable}
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
weight="variable-name"
specifies the weight variable.
transform Action
Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. | |
|
— |
specifies the settings for an output table that includes information about the binning results. | |
|
— |
specifies the settings for an output table that contains the nominal level bin mapping information. | |
|
— |
specifies the settings for an output table that includes information for the variable transformations. | |
|
casOut |
specifies the settings for generating SAS DATA step scoring code. | |
|
— |
specifies the settings for an output table that contains the transformation model table. |
Parameter Descriptions
casOut={casouttable}
specifies the settings for an output table.
For more information about specifying the casOut parameter, see the common casouttable parameter.
casOutBinDetails={casouttable}
specifies the settings for an output table that includes information about the binning results.
For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.
casOutLevelBinMap={casouttable}
specifies the settings for an output table that contains the nominal level bin mapping information.
For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.
casOutVarTransInfo={casouttable}
specifies the settings for an output table that includes information for the variable transformations.
For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.
code={codegen}
specifies the settings for generating SAS DATA step scoring code.
For more information about specifying the code parameter, see the common codegen parameter.
copyAllVars=true | false
when set to True, all the variables from the input table are copied to the scored output table.
| Alias | allIdVars |
|---|---|
| Default | false |
copyVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.
distinctCountLimit=integer
specifies the distinct count limit.
evaluationStats=true | false
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|---|
| Default | false |
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
includeInputVars=true | false
when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.
| Default | false |
|---|
includeMissingGroup=true | false
when set to True, missing values are allowed as group-by keys.
| Default | false |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
misraGries=true | false
specifies that the Misra-Gries algorithm be used for most frequent estimation.
| Default | false |
|---|
outputTableOptions={outputTableOptions}
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | tblOpts |
|---|
The outputTableOptions value can be one or more of the following:
forceTableReturn=true | false
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | false |
|---|
tableNames={"string-1" <, "string-2", ...>}
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
alpha=double
specifies the significance level.
| Default | 0.05 |
|---|
binMapping="LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
binMissing=true | false
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | false |
binOutliers=true | false
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | false |
|---|
emptyBins=true | false
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | false |
|---|
enforceBinaryLevels=true | false
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | true |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=true | false
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | true |
|---|
missingEvalNonEvent=true | false
when set to True, missing values of the target variables are considered as non-event values.
| Default | false |
|---|
noDataLowerUpperBound=true | false
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | false |
|---|
outlierBinsStats=true | false
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | true |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch={quantileSketchOptions}
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
compressionFactor=double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
epsilon=double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
rank={rankOptions}
specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.
| Long form | rank={intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"} |
|---|---|
| Shortcut form | rank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE" |
The rankOptions value can be one or more of the following:
intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"
specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
nominalStat="CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"
specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
topKInteractions=integer
| Default | 10 |
|---|---|
| Minimum value | 1 |
topKSave=integer
| Default | 1 |
|---|---|
| Minimum value | 1 |
requestPackages={{transformRequestPackage-1} <, {transformRequestPackage-2}, ...>}
specifies an array of transform request packages to be processed by the action.
| Aliases | pipelines |
|---|---|
| reqPacks |
The transformRequestPackage value can be one or more of the following:
catTrans={catTransPhase}
specifies the parameters to use for the categorical transformation phase.
The catTransPhase value can be one or more of the following:
arguments={catTransArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The catTransArguments value can be one or more of the following:
contingencyTblOpts={contingencyTableOptions}
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
inputsNLevels=integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
inputsRawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
maxNBins=integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
minNBins=integer
specifies the minimum number of bins.
nBinsArray={integer-1 <, integer-2, ...>} | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
| Aliases | miscellaneousOpts |
|---|---|
| opts |
The globalOverrides value can be one or more of the following:
binMissing=true | false
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | false |
emptyBins=true | false
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | false |
|---|
enforceBinaryLevels=true | false
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | true |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=true | false
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | true |
|---|
missingEvalNonEvent=true | false
when set to True, missing values of the target variables are considered as non-event values.
| Default | false |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
preprocessRare=true | false
when set to True, rare levels are grouped into a single group at the start of the grouping process.
| Alias | preprocess |
|---|---|
| Default | false |
rareThreshold=integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
rareThresholdPercent=double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Default | 5 |
|---|---|
| Range | (0, 100) |
method="DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"
specifies the binning technique to use.
DTREE
groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.
dateTime={dateTimePhase}
specifies the parameters to use for the date-time transformation phase.
The dateTimePhase value can be one or more of the following:
inputType="DATE" | "DATETIME" | "TIME"
method={"ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"}
specifies the binning technique to use.
| Alias | tech |
|---|
discretize={discretizePhase}
specifies the parameters to use for the discretization phase.
The discretizePhase value can be one or more of the following:
arguments={discretizeArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The discretizeArguments value can be one or more of the following:
binEnds={double-1 <, double-2, ...>}
specifies the bin end values. If applicable, they override the data maximum values.
| Alias | binEnd |
|---|
binStarts={double-1 <, double-2, ...>}
specifies the bin start values. If applicable, they override the data minimum values.
| Alias | binStart |
|---|
binWidths={double-1 <, double-2, ...>}
specifies the bin width.
| Alias | binWidth |
|---|
contingencyTblOpts={contingencyTableOptions}
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
inputsNLevels=integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
inputsRawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
cutPoints={double-1 <, double-2, ...>}
specifies the user-provided cutpoints, for the CUTPTS binning technique.
| Alias | cutPts |
|---|
maxNBins=integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
| Default | 5 |
|---|---|
| Minimum value (exclusive) | 0 |
minNBins=integer
specifies the minimum number of bins.
| Default | 1 |
|---|---|
| Minimum value (exclusive) | 0 |
nBinsArray={integer-1 <, integer-2, ...>} | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
alpha=double
specifies the significance level.
| Default | 0.05 |
|---|
binMapping="LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
binMissing=true | false
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | false |
binOutliers=true | false
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | false |
|---|
emptyBins=true | false
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | false |
|---|
enforceBinaryLevels=true | false
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | true |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=true | false
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | true |
|---|
missingEvalNonEvent=true | false
when set to True, missing values of the target variables are considered as non-event values.
| Default | false |
|---|
noDataLowerUpperBound=true | false
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | false |
|---|
outlierBinsStats=true | false
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | true |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
method="BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"
specifies the binning technique to use.
CACC
creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.
CAIM
creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.
CHIMERGE
creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.
DTREE
creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.
MDLP
creates bins based on the minimum description length. This is a top-down supervised discretization technique.
evaluationStats=true | false | {evaluationStatsOptions}
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|
The evaluationStatsOptions value can be one or more of the following:
chiSqGroup=true | false
| Default | false |
|---|
ftTestGroup=true | false
| Default | false |
|---|
missIndicatorTarget=true | false
| Default | false |
|---|
nominalTarget=true | false
| Default | false |
|---|
woeGroup=true | false
| Default | false |
|---|
events={"string-1" <, "string-2", ...>}
specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.
| Alias | evalVarsEvents |
|---|
featureInteraction={featureInteraction}
options that control the generation of interaction features.
| Alias | interaction |
|---|
The featureInteraction value can be one or more of the following:
coefficients={double-1 <, double-2, ...>}
specifies the coefficients for the linear interaction operator.
inputTransformations={"string-1" <, "string-2", ...>}
specifies the transformations that are to be used for generating the input component features of the interaction features.
| Alias | inputs |
|---|
method="CROSS" | "ORDERED"
power=integer
specifies the number of inputs for the polynomial feature interaction operator.
| Default | 2 |
|---|---|
| Range | 1–4 |
synthesizer="DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"
targetTransformation="string"
specifies the transformation that is to be used for generating the target features for interaction feature generation.
| Aliases | targets |
|---|---|
| target |
featureProbe={featureProbePhase}
specifies the parameters to use for the feature probe transformation phase.
The featureProbePhase value can be one or more of the following:
arguments={featureProbeArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The featureProbeArguments value can be one or more of the following:
ecdfTolerance=double
specifies the tolerance value for the empirical cumulative distribution function.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
nProbes=integer
the number of feature probes.
| Default | 1 |
|---|---|
| Minimum value | 1 |
probeMissing=true | false
when set to True, generates missing values at the observed missing rate.
| Default | true |
|---|
rareThreshold=integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
rareThresholdPercent=double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Range | (0, 100) |
|---|
rawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
shrinkageFactor=double
specifies the shrinkage factor for level probability estimation.
| Default | 0 |
|---|---|
| Minimum value | 0 |
useRawLevel=true | false
specifies that the raw values be used as levels of the nominal variable.
| Default | false |
|---|
function={functionPhase}
specifies the parameters to use for the functional transformation phase.
The functionPhase value can be one or more of the following:
arguments={functionArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The functionArguments value can be one or more of the following:
aadLocationUseMean=true | false
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | true |
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
lowerPercentile=double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Default | 10 |
| Range | (0, 50) |
otherArguments={double-1 <, double-2, ...>}
specifies other values to use. The values depend on the type of functional transformation.
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
scaleMultiplier=double
specifies the multiplying factor for the chosen scale estimator.
| Alias | scaleMulFac |
|---|
shiftMax=double
specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value
| Alias | shiftNegative |
|---|
shiftMin=double
specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value
| Alias | shiftPositive |
|---|
symmetricPercentile=double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
upperPercentile=double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Default | 90 |
| Range | (50, 100) |
method="ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"
specifies the functional transformation.
CENTER
returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.
LOG
returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.
POWER
returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.
RANGE
returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.
SCALESHIFT
returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.
hash={hashPhase}
specifies the parameters to use for the hashing transformation phase.
The hashPhase value can be one or more of the following:
arguments={hashArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
nBuckets=integer
specifies the arguments for this phase of the transform.
method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
specifies the hash function.
impute={imputePhase}
specifies the parameters to use for the imputation phase.
The imputePhase value can be one or more of the following:
maxRandom=double
specifies the maximum random number to generate.
method="MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"
minRandom=double
specifies the minimum random number to generate.
valuesInterval={double-1 <, double-2, ...>}
specifies a list of double values for imputation for the interval variables.
| Alias | valuesNumeric |
|---|
valuesNominal={"string-1" <, "string-2", ...>}
specifies a list of string values for imputation for the nominal variables.
| Alias | valuesNonNumeric |
|---|
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
inputsInheritFormats=true | false
specifies that the variables inherit formats from the underlying table.
| Default | false |
|---|
mapInterval={mapIntervalPhase}
specifies the parameters to use for the map to interval transformation phase.
The mapIntervalPhase value can be one or more of the following:
arguments={mapIntervalArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The mapIntervalArguments value can be one or more of the following:
descending=true | false
specifies that the label count encoding be performed.
| Default | true |
|---|
includeMissingLevel=true | false
when set to True, missing values are included in the distinct level analysis instead of being discarded.
| Default | false |
|---|
nLevels=integer
specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.
| Default | 2 |
|---|---|
| Range | 1–10 |
nMoments=integer
specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.
| Default | 2 |
|---|---|
| Range | 1–6 |
noise=double
specifies the parameter for the Laplace or uniform noise to be added to the level statistics.
| Default | 0 |
|---|---|
| Minimum value (exclusive) | 0 |
shrinkageFactor=double
specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.
| Default | 0 |
|---|---|
| Minimum value | 0 |
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
method="COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"
specifies the interval map criterion to use.
| Alias | tech |
|---|---|
| Default | WOE |
name="string"
specifies a name for the request package.
outlier={outlierPhase}
specifies the parameters to use for the outlier determination and treatment phase.
The outlierPhase value can be one or more of the following:
arguments={outlierArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The outlierArguments value can be one or more of the following:
aadLocationUseMean=true | false
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | true |
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
lowerPercentile=double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Range | (0, 50) |
max=double
specifies a global maximum value.
min=double
specifies a global minimum value.
replacements={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {double-1 <, double-2, ...>}
specifies the values to use as replacements for outliers. These can be user defined values or location estimates.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
| Default | STD |
|---|
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
scaleMultiplier=double
specifies the multiplying factor for the chosen scale estimator.
symmetricPercentile=double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Range | (0, 100) |
upperPercentile=double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Range | (50, 100) |
userDefinedLimits={double-1 <, double-2, ...>}
uses the specified user-defined limits as the lower and upper thresholds for each variable.
zScoreThreshold=double
specifies the Z threshold.
method="IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"
specifies the outlier definition.
IQR
uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.
MIQR
uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.
MZSCORE
uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.
PERC
uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.
treatment="REPLACE" | "TRIM" | "WINSOR"
specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.
output={outputPhase}
specifies the parameters to use for the output phase.
The outputPhase value can be one or more of the following:
noScoreCode=true | false
when set to True, no score code is sent to output.
| Default | false |
|---|
noScoreTable=true | false
when set to True, no scoring is sent to output.
| Alias | noScoreTbl |
|---|---|
| Default | false |
scoreWOE=true | false
when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.
| Default | false |
|---|
phaseOrder="FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"
specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.
| Default | IOF |
|---|
targets={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies a list of target variables to use.
For more information about specifying the targets parameter, see the common casinvardesc parameter.
| Alias | evalVars |
|---|
targetsInheritFormats=true | false
specifies that the variables inherit formats from the underlying table.
| Default | false |
|---|
sasVarNameLength=true | false
when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.
| Default | false |
|---|
saveState={casouttable}
specifies the settings for an output table that contains the transformation model table.
| Alias | saveModel |
|---|
| Long form | saveState={name="table-name"} |
|---|---|
| Shortcut form | saveState="table-name" |
The casouttable value can be one or more of the following:
caslib="string"
specifies the name of the caslib for the output table.
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data.
lifetime=64-bit-integer
specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.
| Default | 0 |
|---|---|
| Minimum value | 0 |
memoryFormat="DVR" | "INHERIT" | "STANDARD"
specifies the memory format for the output table.
| Default | INHERIT |
|---|
DVR
use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.
name="table-name"
specifies the name for the output table.
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
|---|
replace=true | false
when set to True, overwrites an existing table that has the same name.
| Default | false |
|---|
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
* table={castable}
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
weight="variable-name"
specifies the weight variable.
transform Action
Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. | |
|
— |
specifies the settings for an output table that includes information about the binning results. | |
|
— |
specifies the settings for an output table that contains the nominal level bin mapping information. | |
|
— |
specifies the settings for an output table that includes information for the variable transformations. | |
|
casOut |
specifies the settings for generating SAS DATA step scoring code. | |
|
— |
specifies the settings for an output table that contains the transformation model table. |
Parameter Descriptions
casOut={casouttable}
specifies the settings for an output table.
For more information about specifying the casOut parameter, see the common casouttable parameter.
casOutBinDetails={casouttable}
specifies the settings for an output table that includes information about the binning results.
For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.
casOutLevelBinMap={casouttable}
specifies the settings for an output table that contains the nominal level bin mapping information.
For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.
casOutVarTransInfo={casouttable}
specifies the settings for an output table that includes information for the variable transformations.
For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.
code={codegen}
specifies the settings for generating SAS DATA step scoring code.
For more information about specifying the code parameter, see the common codegen parameter.
copyAllVars=True | False
when set to True, all the variables from the input table are copied to the scored output table.
| Alias | allIdVars |
|---|---|
| Default | False |
copyVars=["variable-name-1" <, "variable-name-2", ...>]
specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.
distinctCountLimit=integer
specifies the distinct count limit.
evaluationStats=True | False
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|---|
| Default | False |
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
includeInputVars=True | False
when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.
| Default | False |
|---|
includeMissingGroup=True | False
when set to True, missing values are allowed as group-by keys.
| Default | False |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
misraGries=True | False
specifies that the Misra-Gries algorithm be used for most frequent estimation.
| Default | False |
|---|
outputTableOptions={outputTableOptions}
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | tblOpts |
|---|
The outputTableOptions value can be one or more of the following:
"forceTableReturn":True | False
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | False |
|---|
"tableNames":["string-1" <, "string-2", ...>]
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
overrides={globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
"alpha":double
specifies the significance level.
| Default | 0.05 |
|---|
"binMapping":"LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
"binMissing":True | False
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | False |
"binOutliers":True | False
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | False |
|---|
"emptyBins":True | False
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | False |
|---|
"enforceBinaryLevels":True | False
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | True |
|---|
"ivFactor":double
specifies the information value adjustment factor.
| Default | 2 |
|---|
"minNObsInBin":64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
"minPerNObsInBin":double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
"missingBinStats":True | False
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | True |
|---|
"missingEvalNonEvent":True | False
when set to True, missing values of the target variables are considered as non-event values.
| Default | False |
|---|
"noDataLowerUpperBound":True | False
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | False |
|---|
"outlierBinsStats":True | False
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | True |
|---|
"woeAdjust":double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
"woeDefinition":"EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch={quantileSketchOptions}
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
"compressionFactor":double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
"epsilon":double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
rank={rankOptions}
specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.
| Long form | rank={"intervalStat":"AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"} |
|---|---|
| Shortcut form | rank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE" |
The rankOptions value can be one or more of the following:
"intervalStat":"AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"
specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
"nominalStat":"CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"
specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
"topKInteractions":integer
| Default | 10 |
|---|---|
| Minimum value | 1 |
"topKSave":integer
| Default | 1 |
|---|---|
| Minimum value | 1 |
requestPackages=[{transformRequestPackage-1} <, {transformRequestPackage-2}, ...>]
specifies an array of transform request packages to be processed by the action.
| Aliases | pipelines |
|---|---|
| reqPacks |
The transformRequestPackage value can be one or more of the following:
"catTrans":{catTransPhase}
specifies the parameters to use for the categorical transformation phase.
The catTransPhase value can be one or more of the following:
"arguments":{catTransArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The catTransArguments value can be one or more of the following:
"contingencyTblOpts":{contingencyTableOptions}
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
"inputsMethod":"BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
"inputsNLevels":integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
"inputsRawLevelStartingValue":integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"maxNBins":integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
"minNBins":integer
specifies the minimum number of bins.
"nBinsArray":[integer-1 <, integer-2, ...>] | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
"overrides":{globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
| Aliases | miscellaneousOpts |
|---|---|
| opts |
The globalOverrides value can be one or more of the following:
"binMissing":True | False
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | False |
"emptyBins":True | False
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | False |
|---|
"enforceBinaryLevels":True | False
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | True |
|---|
"ivFactor":double
specifies the information value adjustment factor.
| Default | 2 |
|---|
"minNObsInBin":64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
"minPerNObsInBin":double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
"missingBinStats":True | False
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | True |
|---|
"missingEvalNonEvent":True | False
when set to True, missing values of the target variables are considered as non-event values.
| Default | False |
|---|
"woeAdjust":double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
"woeDefinition":"EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
"preprocessRare":True | False
when set to True, rare levels are grouped into a single group at the start of the grouping process.
| Alias | preprocess |
|---|---|
| Default | False |
"rareThreshold":integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
"rareThresholdPercent":double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Default | 5 |
|---|---|
| Range | (0, 100) |
"method":"DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"
specifies the binning technique to use.
DTREE
groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.
"dateTime":{dateTimePhase}
specifies the parameters to use for the date-time transformation phase.
The dateTimePhase value can be one or more of the following:
"inputType":"DATE" | "DATETIME" | "TIME"
"method":["ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR"]
specifies the binning technique to use.
| Alias | tech |
|---|
"discretize":{discretizePhase}
specifies the parameters to use for the discretization phase.
The discretizePhase value can be one or more of the following:
"arguments":{discretizeArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The discretizeArguments value can be one or more of the following:
"binEnds":[double-1 <, double-2, ...>]
specifies the bin end values. If applicable, they override the data maximum values.
| Alias | binEnd |
|---|
"binStarts":[double-1 <, double-2, ...>]
specifies the bin start values. If applicable, they override the data minimum values.
| Alias | binStart |
|---|
"binWidths":[double-1 <, double-2, ...>]
specifies the bin width.
| Alias | binWidth |
|---|
"contingencyTblOpts":{contingencyTableOptions}
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
"inputsMethod":"BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
"inputsNLevels":integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
"inputsRawLevelStartingValue":integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"cutPoints":[double-1 <, double-2, ...>]
specifies the user-provided cutpoints, for the CUTPTS binning technique.
| Alias | cutPts |
|---|
"maxNBins":integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
| Default | 5 |
|---|---|
| Minimum value (exclusive) | 0 |
"minNBins":integer
specifies the minimum number of bins.
| Default | 1 |
|---|---|
| Minimum value (exclusive) | 0 |
"nBinsArray":[integer-1 <, integer-2, ...>] | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
"overrides":{globalOverrides}
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
"alpha":double
specifies the significance level.
| Default | 0.05 |
|---|
"binMapping":"LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
"binMissing":True | False
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | False |
"binOutliers":True | False
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | False |
|---|
"emptyBins":True | False
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | False |
|---|
"enforceBinaryLevels":True | False
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | True |
|---|
"ivFactor":double
specifies the information value adjustment factor.
| Default | 2 |
|---|
"minNObsInBin":64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
"minPerNObsInBin":double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
"missingBinStats":True | False
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | True |
|---|
"missingEvalNonEvent":True | False
when set to True, missing values of the target variables are considered as non-event values.
| Default | False |
|---|
"noDataLowerUpperBound":True | False
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | False |
|---|
"outlierBinsStats":True | False
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | True |
|---|
"woeAdjust":double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
"woeDefinition":"EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
"method":"BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"
specifies the binning technique to use.
CACC
creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.
CAIM
creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.
CHIMERGE
creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.
DTREE
creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.
MDLP
creates bins based on the minimum description length. This is a top-down supervised discretization technique.
"evaluationStats":True | False | {evaluationStatsOptions}
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|
The evaluationStatsOptions value can be one or more of the following:
chiSqGroup=True | False
| Default | False |
|---|
ftTestGroup=True | False
| Default | False |
|---|
missIndicatorTarget=True | False
| Default | False |
|---|
nominalTarget=True | False
| Default | False |
|---|
woeGroup=True | False
| Default | False |
|---|
"events":["string-1" <, "string-2", ...>]
specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.
| Alias | evalVarsEvents |
|---|
"featureInteraction":{featureInteraction}
options that control the generation of interaction features.
| Alias | interaction |
|---|
The featureInteraction value can be one or more of the following:
"coefficients":[double-1 <, double-2, ...>]
specifies the coefficients for the linear interaction operator.
"inputTransformations":["string-1" <, "string-2", ...>]
specifies the transformations that are to be used for generating the input component features of the interaction features.
| Alias | inputs |
|---|
"method":"CROSS" | "ORDERED"
"power":integer
specifies the number of inputs for the polynomial feature interaction operator.
| Default | 2 |
|---|---|
| Range | 1–4 |
"synthesizer":"DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"
"targetTransformation":"string"
specifies the transformation that is to be used for generating the target features for interaction feature generation.
| Aliases | targets |
|---|---|
| target |
"featureProbe":{featureProbePhase}
specifies the parameters to use for the feature probe transformation phase.
The featureProbePhase value can be one or more of the following:
"arguments":{featureProbeArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The featureProbeArguments value can be one or more of the following:
"ecdfTolerance":double
specifies the tolerance value for the empirical cumulative distribution function.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
"nProbes":integer
the number of feature probes.
| Default | 1 |
|---|---|
| Minimum value | 1 |
"probeMissing":True | False
when set to True, generates missing values at the observed missing rate.
| Default | True |
|---|
"rareThreshold":integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
"rareThresholdPercent":double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Range | (0, 100) |
|---|
"rawLevelStartingValue":integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"shrinkageFactor":double
specifies the shrinkage factor for level probability estimation.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"useRawLevel":True | False
specifies that the raw values be used as levels of the nominal variable.
| Default | False |
|---|
"function":{functionPhase}
specifies the parameters to use for the functional transformation phase.
The functionPhase value can be one or more of the following:
"arguments":{functionArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The functionArguments value can be one or more of the following:
"aadLocationUseMean":True | False
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | True |
"location":"BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
"locationBiweightTuning":double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
"lowerPercentile":double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Default | 10 |
| Range | (0, 50) |
"otherArguments":[double-1 <, double-2, ...>]
specifies other values to use. The values depend on the type of functional transformation.
"scale":"AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
"scaleBiweightTuning":double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
"scaleMultiplier":double
specifies the multiplying factor for the chosen scale estimator.
| Alias | scaleMulFac |
|---|
"shiftMax":double
specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value
| Alias | shiftNegative |
|---|
"shiftMin":double
specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value
| Alias | shiftPositive |
|---|
"symmetricPercentile":double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
"upperPercentile":double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Default | 90 |
| Range | (50, 100) |
"method":"ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"
specifies the functional transformation.
CENTER
returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.
LOG
returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.
POWER
returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.
RANGE
returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.
SCALESHIFT
returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.
"hash":{hashPhase}
specifies the parameters to use for the hashing transformation phase.
The hashPhase value can be one or more of the following:
"arguments":{hashArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
"nBuckets":integer
specifies the arguments for this phase of the transform.
"method":"BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
specifies the hash function.
"impute":{imputePhase}
specifies the parameters to use for the imputation phase.
The imputePhase value can be one or more of the following:
"maxRandom":double
specifies the maximum random number to generate.
"method":"MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"
"minRandom":double
specifies the minimum random number to generate.
"valuesInterval":[double-1 <, double-2, ...>]
specifies a list of double values for imputation for the interval variables.
| Alias | valuesNumeric |
|---|
"valuesNominal":["string-1" <, "string-2", ...>]
specifies a list of string values for imputation for the nominal variables.
| Alias | valuesNonNumeric |
|---|
"inputs":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
"inputsInheritFormats":True | False
specifies that the variables inherit formats from the underlying table.
| Default | False |
|---|
"mapInterval":{mapIntervalPhase}
specifies the parameters to use for the map to interval transformation phase.
The mapIntervalPhase value can be one or more of the following:
"arguments":{mapIntervalArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The mapIntervalArguments value can be one or more of the following:
"descending":True | False
specifies that the label count encoding be performed.
| Default | True |
|---|
"includeMissingLevel":True | False
when set to True, missing values are included in the distinct level analysis instead of being discarded.
| Default | False |
|---|
"nLevels":integer
specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.
| Default | 2 |
|---|---|
| Range | 1–10 |
"nMoments":integer
specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.
| Default | 2 |
|---|---|
| Range | 1–6 |
"noise":double
specifies the parameter for the Laplace or uniform noise to be added to the level statistics.
| Default | 0 |
|---|---|
| Minimum value (exclusive) | 0 |
"shrinkageFactor":double
specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"woeAdjust":double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
"woeDefinition":"EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
"method":"COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"
specifies the interval map criterion to use.
| Alias | tech |
|---|---|
| Default | WOE |
"name":"string"
specifies a name for the request package.
"outlier":{outlierPhase}
specifies the parameters to use for the outlier determination and treatment phase.
The outlierPhase value can be one or more of the following:
"arguments":{outlierArguments}
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The outlierArguments value can be one or more of the following:
"aadLocationUseMean":True | False
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | True |
"location":"BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
"locationBiweightTuning":double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
"lowerPercentile":double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Range | (0, 50) |
"max":double
specifies a global maximum value.
"min":double
specifies a global minimum value.
"replacements":["BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"] | [double-1 <, double-2, ...>]
specifies the values to use as replacements for outliers. These can be user defined values or location estimates.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
"scale":"AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
| Default | STD |
|---|
"scaleBiweightTuning":double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
"scaleMultiplier":double
specifies the multiplying factor for the chosen scale estimator.
"symmetricPercentile":double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Range | (0, 100) |
"upperPercentile":double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Range | (50, 100) |
"userDefinedLimits":[double-1 <, double-2, ...>]
uses the specified user-defined limits as the lower and upper thresholds for each variable.
"zScoreThreshold":double
specifies the Z threshold.
"method":"IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"
specifies the outlier definition.
IQR
uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.
MIQR
uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.
MZSCORE
uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.
PERC
uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.
"treatment":"REPLACE" | "TRIM" | "WINSOR"
specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.
"output":{outputPhase}
specifies the parameters to use for the output phase.
The outputPhase value can be one or more of the following:
"noScoreCode":True | False
when set to True, no score code is sent to output.
| Default | False |
|---|
"noScoreTable":True | False
when set to True, no scoring is sent to output.
| Alias | noScoreTbl |
|---|---|
| Default | False |
"scoreWOE":True | False
when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.
| Default | False |
|---|
"phaseOrder":"FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"
specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.
| Default | IOF |
|---|
"targets":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies a list of target variables to use.
For more information about specifying the targets parameter, see the common casinvardesc parameter.
| Alias | evalVars |
|---|
"targetsInheritFormats":True | False
specifies that the variables inherit formats from the underlying table.
| Default | False |
|---|
sasVarNameLength=True | False
when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.
| Default | False |
|---|
saveState={casouttable}
specifies the settings for an output table that contains the transformation model table.
| Alias | saveModel |
|---|
| Long form | saveState={"name":"table-name"} |
|---|---|
| Shortcut form | saveState="table-name" |
The casouttable value can be one or more of the following:
"caslib":"string"
specifies the name of the caslib for the output table.
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data.
"lifetime":64-bit-integer
specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"memoryFormat":"DVR" | "INHERIT" | "STANDARD"
specifies the memory format for the output table.
| Default | INHERIT |
|---|
DVR
use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.
"name":"table-name"
specifies the name for the output table.
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
|---|
"replace":True | False
when set to True, overwrites an existing table that has the same name.
| Default | False |
|---|
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
* table={castable}
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
weight="variable-name"
specifies the weight variable.
transform Action
Performs pipelined variable imputation, outlier detection and treatment, functional transformation, binning, and robust univariate statistics to evaluate the quality of the transformation.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. | |
|
— |
specifies the settings for an output table that includes information about the binning results. | |
|
— |
specifies the settings for an output table that contains the nominal level bin mapping information. | |
|
— |
specifies the settings for an output table that includes information for the variable transformations. | |
|
casOut |
specifies the settings for generating SAS DATA step scoring code. | |
|
— |
specifies the settings for an output table that contains the transformation model table. |
Parameter Descriptions
casOut=list(casouttable)
specifies the settings for an output table.
For more information about specifying the casOut parameter, see the common casouttable parameter.
casOutBinDetails=list(casouttable)
specifies the settings for an output table that includes information about the binning results.
For more information about specifying the casOutBinDetails parameter, see the common casouttable parameter.
casOutLevelBinMap=list(casouttable)
specifies the settings for an output table that contains the nominal level bin mapping information.
For more information about specifying the casOutLevelBinMap parameter, see the common casouttable parameter.
casOutVarTransInfo=list(casouttable)
specifies the settings for an output table that includes information for the variable transformations.
For more information about specifying the casOutVarTransInfo parameter, see the common casouttable parameter.
code=list(codegen)
specifies the settings for generating SAS DATA step scoring code.
For more information about specifying the code parameter, see the common codegen parameter.
copyAllVars=TRUE | FALSE
when set to True, all the variables from the input table are copied to the scored output table.
| Alias | allIdVars |
|---|---|
| Default | FALSE |
copyVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the names of variables in the input table to use for identifying scored observations in the output table. The specified variables are copied to the output table.
distinctCountLimit=integer
specifies the distinct count limit.
evaluationStats=TRUE | FALSE
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|---|
| Default | FALSE |
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
includeInputVars=TRUE | FALSE
when set to True, the analysis variables from the input table that are specified in the vars parameter are copied to the output table.
| Default | FALSE |
|---|
includeMissingGroup=TRUE | FALSE
when set to True, missing values are allowed as group-by keys.
| Default | FALSE |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
misraGries=TRUE | FALSE
specifies that the Misra-Gries algorithm be used for most frequent estimation.
| Default | FALSE |
|---|
outputTableOptions=list(outputTableOptions)
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | tblOpts |
|---|
The outputTableOptions value can be one or more of the following:
forceTableReturn=TRUE | FALSE
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | FALSE |
|---|
tableNames=list("string-1" <, "string-2", ...>)
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
overrides=list(globalOverrides)
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
alpha=double
specifies the significance level.
| Default | 0.05 |
|---|
binMapping="LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
binMissing=TRUE | FALSE
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | FALSE |
binOutliers=TRUE | FALSE
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | FALSE |
|---|
emptyBins=TRUE | FALSE
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | FALSE |
|---|
enforceBinaryLevels=TRUE | FALSE
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | TRUE |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=TRUE | FALSE
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
missingEvalNonEvent=TRUE | FALSE
when set to True, missing values of the target variables are considered as non-event values.
| Default | FALSE |
|---|
noDataLowerUpperBound=TRUE | FALSE
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | FALSE |
|---|
outlierBinsStats=TRUE | FALSE
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch=list(quantileSketchOptions)
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
compressionFactor=double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
epsilon=double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
rank=list(rankOptions)
specifies options for ranking the transformations. The ranking includes both local ranking, among the transformations of a variable, and global ranking across all transformations of all variables.
| Long form | rank=list(intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE") |
|---|---|
| Shortcut form | rank="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE" |
The rankOptions value can be one or more of the following:
intervalStat="AD" | "AVGQUANKURT" | "AVGQUANSKEW" | "CLASSICALKURT" | "CLASSICALSKEW" | "CVM" | "KS" | "PEARSON" | "VARIANCE"
specifies the interval transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
nominalStat="CHISQ" | "CRAMERSV" | "FTEST" | "G2" | "GINI" | "IV" | "WELCHTTEST" | "WOE"
specifies the nominal transformation ranking statistic when evaluation statistic method is chosen as the ranking method.
topKInteractions=integer
| Default | 10 |
|---|---|
| Minimum value | 1 |
topKSave=integer
| Default | 1 |
|---|---|
| Minimum value | 1 |
requestPackages=list( list(transformRequestPackage-1) <, list(transformRequestPackage-2), ...>)
specifies an array of transform request packages to be processed by the action.
| Aliases | pipelines |
|---|---|
| reqPacks |
The transformRequestPackage value can be one or more of the following:
catTrans=list(catTransPhase)
specifies the parameters to use for the categorical transformation phase.
The catTransPhase value can be one or more of the following:
arguments=list(catTransArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The catTransArguments value can be one or more of the following:
contingencyTblOpts=list(contingencyTableOptions)
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
inputsNLevels=integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
inputsRawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
maxNBins=integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
minNBins=integer
specifies the minimum number of bins.
nBinsArray=list(integer-1 <, integer-2, ...>) | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
overrides=list(globalOverrides)
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
| Aliases | miscellaneousOpts |
|---|---|
| opts |
The globalOverrides value can be one or more of the following:
binMissing=TRUE | FALSE
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | FALSE |
emptyBins=TRUE | FALSE
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | FALSE |
|---|
enforceBinaryLevels=TRUE | FALSE
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | TRUE |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=TRUE | FALSE
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
missingEvalNonEvent=TRUE | FALSE
when set to True, missing values of the target variables are considered as non-event values.
| Default | FALSE |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
preprocessRare=TRUE | FALSE
when set to True, rare levels are grouped into a single group at the start of the grouping process.
| Alias | preprocess |
|---|---|
| Default | FALSE |
rareThreshold=integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
rareThresholdPercent=double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Default | 5 |
|---|---|
| Range | (0, 100) |
method="DTREE" | "GROUPRARE" | "ONEHOT" | "RTREE" | "WOE"
specifies the binning technique to use.
DTREE
groups based on a one-level decision tree. The criterion is controlled with the crit parameter. This is a supervised technique.
dateTime=list(dateTimePhase)
specifies the parameters to use for the date-time transformation phase.
The dateTimePhase value can be one or more of the following:
inputType="DATE" | "DATETIME" | "TIME"
method=list("ALL", "DAYMONTH", "DAYWEEK", "HOUR", "LEAPYEAR", "MINUTE", "MONTH", "QUARTER", "WEEK", "WEEKEND", "YEAR")
specifies the binning technique to use.
| Alias | tech |
|---|
discretize=list(discretizePhase)
specifies the parameters to use for the discretization phase.
The discretizePhase value can be one or more of the following:
arguments=list(discretizeArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The discretizeArguments value can be one or more of the following:
binEnds=list(double-1 <, double-2, ...>)
specifies the bin end values. If applicable, they override the data maximum values.
| Alias | binEnd |
|---|
binStarts=list(double-1 <, double-2, ...>)
specifies the bin start values. If applicable, they override the data minimum values.
| Alias | binStart |
|---|
binWidths=list(double-1 <, double-2, ...>)
specifies the bin width.
| Alias | binWidth |
|---|
contingencyTblOpts=list(contingencyTableOptions)
controls the number of rows for the X axis transformation variable, the number of columns for the Y axis target variable, and the location of the row cutpoints.
| Alias | cTblOpts |
|---|
The contingencyTableOptions value can be one or more of the following:
inputsMethod="BUCKET" | "DISTINCTLEVELS" | "QUANTILE" | "RAW"
specifies the method for determining the levels of the transformation variable.
BUCKET
generates levels that map to the bins of bucket binning. Specify the number of bins with the evalNLevels parameter.
DISTINCTLEVELS
generates levels that map to the distinct values of the target variable.
| Aliases | CLASS |
|---|---|
| LEVEL |
inputsNLevels=integer
specifies the number of levels to use for the transformation variable. This parameter applies to the BUCKET and QUANTILE methods.
| Alias | nInitBins |
|---|
inputsRawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
cutPoints=list(double-1 <, double-2, ...>)
specifies the user-provided cutpoints, for the CUTPTS binning technique.
| Alias | cutPts |
|---|
maxNBins=integer
specifies the maximum number of bins for supervised techniques. The default for discretization is five, while for nominal grouping, a default value is computed as log2 of the number of distinct values.
| Default | 5 |
|---|---|
| Minimum value (exclusive) | 0 |
minNBins=integer
specifies the minimum number of bins.
| Default | 1 |
|---|---|
| Minimum value (exclusive) | 0 |
nBinsArray=list(integer-1 <, integer-2, ...>) | integer
specifies a list of the number of bins to create for each variable. If there are more variables than specified bins, the last value is used for the remaining variables. Extra values are discarded. By default, five bins are used for each variable.
| Alias | nBins |
|---|
overrides=list(globalOverrides)
specifies the global options that apply across request packages. Each request package can override these parameters by setting the corresponding parameter.
The globalOverrides value can be one or more of the following:
alpha=double
specifies the significance level.
| Default | 0.05 |
|---|
binMapping="LEFT" | "RIGHT"
controls how to map values that fall at the boundary between consecutive bins. LEFT enables you to express the bins with [], (], ..., (] notation. RIGHT enables [), [), ..., [] notation.
| Default | RIGHT |
|---|
binMissing=TRUE | FALSE
when set to True, bins missing values into a separate bin. The ID for this bin is 0.
| Alias | mapMissing |
|---|---|
| Default | FALSE |
binOutliers=TRUE | FALSE
when set to True, outliers are binned into distinct bins. If n bins are generated for non-outlier values, then the lower and upper outlier bins correspond to bin IDs n+1 and n+2, respectively.
| Default | FALSE |
|---|
emptyBins=TRUE | FALSE
when set to True, bins with zero observations are permitted. By default, leading and trailing empty bins are removed. Other empty bins are combined with the first non-empty bin to the right.
| Default | FALSE |
|---|
enforceBinaryLevels=TRUE | FALSE
when set to True, enforces binary levels during the computation of WOE, IV, and Gini evaluation statistics. If set to False and the number of levels is greater than two, then binary evaluation statistics are ignored, even when they are requested.
| Default | TRUE |
|---|
ivFactor=double
specifies the information value adjustment factor.
| Default | 2 |
|---|
minNObsInBin=64-bit-integer
specifies the minimum number of observations to include in a bin.
| Alias | leafSize |
|---|
minPerNObsInBin=double
specifies the minimum percentage of all observations to include in a bin.
| Default | 5 |
|---|---|
| Range | (0, 50) |
missingBinStats=TRUE | FALSE
when set to True, the missing bin is considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
missingEvalNonEvent=TRUE | FALSE
when set to True, missing values of the target variables are considered as non-event values.
| Default | FALSE |
|---|
noDataLowerUpperBound=TRUE | FALSE
when set to True, during the score code generation, the binset global lower and upper bounds are unlimited instead of set to the values obtained from the data.
| Default | FALSE |
|---|
outlierBinsStats=TRUE | FALSE
when set to True, the outlier bins are considered during the computation of the evaluation statistics.
| Default | TRUE |
|---|
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
method="BUCKET" | "CACC" | "CAIM" | "CHIMERGE" | "CUTPTS" | "DTREE" | "MDLP" | "QUANTILE" | "RTREE" | "WOE"
specifies the binning technique to use.
CACC
creates bins based on class-attribute contingency coefficient. This is a top-down supervised discretization technique.
CAIM
creates bins based on class-attribute independence maximization. This is a top-down supervised discretization technique.
CHIMERGE
creates bins based on chi-square merging of neighboring bins. This is a bottom-up supervised discretization technique.
DTREE
creates bins based on a one-level decision tree. This is a top-down supervised discretization technique.
MDLP
creates bins based on the minimum description length. This is a top-down supervised discretization technique.
evaluationStats=TRUE | FALSE | {evaluationStatsOptions}
when set to True, requests that the default set of evaluation statistics be computed for the transformed variables.
| Alias | evalStats |
|---|
The evaluationStatsOptions value can be one or more of the following:
chiSqGroup=TRUE | FALSE
| Default | FALSE |
|---|
ftTestGroup=TRUE | FALSE
| Default | FALSE |
|---|
missIndicatorTarget=TRUE | FALSE
| Default | FALSE |
|---|
nominalTarget=TRUE | FALSE
| Default | FALSE |
|---|
woeGroup=TRUE | FALSE
| Default | FALSE |
|---|
events=list("string-1" <, "string-2", ...>)
specifies a list of events that correspond to the list of target variables. These values are matched one-to-one with the target variables from the evalVars parameter.
| Alias | evalVarsEvents |
|---|
featureInteraction=list(featureInteraction)
options that control the generation of interaction features.
| Alias | interaction |
|---|
The featureInteraction value can be one or more of the following:
coefficients=list(double-1 <, double-2, ...>)
specifies the coefficients for the linear interaction operator.
inputTransformations=list("string-1" <, "string-2", ...>)
specifies the transformations that are to be used for generating the input component features of the interaction features.
| Alias | inputs |
|---|
method="CROSS" | "ORDERED"
power=integer
specifies the number of inputs for the polynomial feature interaction operator.
| Default | 2 |
|---|---|
| Range | 1–4 |
synthesizer="DIVISION" | "LINEAR" | "MULTIPLICATION" | "NOMINAL" | "POLYNOMIAL"
targetTransformation="string"
specifies the transformation that is to be used for generating the target features for interaction feature generation.
| Aliases | targets |
|---|---|
| target |
featureProbe=list(featureProbePhase)
specifies the parameters to use for the feature probe transformation phase.
The featureProbePhase value can be one or more of the following:
arguments=list(featureProbeArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The featureProbeArguments value can be one or more of the following:
ecdfTolerance=double
specifies the tolerance value for the empirical cumulative distribution function.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
nProbes=integer
the number of feature probes.
| Default | 1 |
|---|---|
| Minimum value | 1 |
probeMissing=TRUE | FALSE
when set to True, generates missing values at the observed missing rate.
| Default | TRUE |
|---|
rareThreshold=integer
specifies the rare frequency threshold.
| Alias | rareFreqCutOff |
|---|---|
| Minimum value (exclusive) | 0 |
rareThresholdPercent=double
specifies the rare threshold percentile. Levels less than the threshold are grouped together.
| Range | (0, 100) |
|---|
rawLevelStartingValue=integer
specifies a starting integer value for creating levels of the transformation variable. This parameter applies to the RAW method.
| Default | 0 |
|---|---|
| Minimum value | 0 |
shrinkageFactor=double
specifies the shrinkage factor for level probability estimation.
| Default | 0 |
|---|---|
| Minimum value | 0 |
useRawLevel=TRUE | FALSE
specifies that the raw values be used as levels of the nominal variable.
| Default | FALSE |
|---|
function=list(functionPhase)
specifies the parameters to use for the functional transformation phase.
The functionPhase value can be one or more of the following:
arguments=list(functionArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The functionArguments value can be one or more of the following:
aadLocationUseMean=TRUE | FALSE
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | TRUE |
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
lowerPercentile=double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Default | 10 |
| Range | (0, 50) |
otherArguments=list(double-1 <, double-2, ...>)
specifies other values to use. The values depend on the type of functional transformation.
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
scaleMultiplier=double
specifies the multiplying factor for the chosen scale estimator.
| Alias | scaleMulFac |
|---|
shiftMax=double
specifies that the argument be shifted to negative value by subtracting the maximum and adding the shiftMax value
| Alias | shiftNegative |
|---|
shiftMin=double
specifies that the argument be shifted to positive value by subtracting the minimum and adding the shiftMin value
| Alias | shiftPositive |
|---|
symmetricPercentile=double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
upperPercentile=double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Default | 90 |
| Range | (50, 100) |
method="ABS" | "ARCSIN" | "BOXCOX" | "CENTER" | "COS" | "COSH" | "EXP" | "IDENTITY" | "INVERSE" | "INVSQUARESHIFT" | "LOG" | "POWER" | "RANGE" | "SCALESHIFT" | "SIN" | "SINH" | "SQRT" | "STANDARDIZE" | "TAN" | "TANH"
specifies the functional transformation.
CENTER
returns the value minus the location determined from the loc parameter. If you do not specify the loc parameter, then the mean is used.
LOG
returns the log of the variable. Specify a base in the otherArgs parameter. The default is to compute the natural log.
POWER
returns the value of the variable raised to a specified power. Specify the power in the otherArgs parameter. The default power is 2.
RANGE
returns the value of the variable, bounded by the range. Specify the minimum and maximum values for the range in the otherArgs parameter. If both values are not specified, then the default range, [0, 1], is used.
SCALESHIFT
returns a scaled and shifted value of the variable. Specify the scale value and the shift value in the otherArgs parameter. If both values are not specified, then the default values (1, 0) are used and perform an identity transformation.
hash=list(hashPhase)
specifies the parameters to use for the hashing transformation phase.
The hashPhase value can be one or more of the following:
arguments=list(hashArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
nBuckets=integer
specifies the arguments for this phase of the transform.
method="BUCKET" | "CAS" | "MISSINDICATOR" | "MURMUR3" | "QUANTILE" | "SUPERFAST"
specifies the hash function.
impute=list(imputePhase)
specifies the parameters to use for the imputation phase.
The imputePhase value can be one or more of the following:
maxRandom=double
specifies the maximum random number to generate.
method="MAX" | "MEAN" | "MEDIAN" | "MIDRANGE" | "MIN" | "MODE" | "RANDOM" | "VALUE"
minRandom=double
specifies the minimum random number to generate.
valuesInterval=list(double-1 <, double-2, ...>)
specifies a list of double values for imputation for the interval variables.
| Alias | valuesNumeric |
|---|
valuesNominal=list("string-1" <, "string-2", ...>)
specifies a list of string values for imputation for the nominal variables.
| Alias | valuesNonNumeric |
|---|
inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies a list of transformation variables. If you do not specify the variables, all numeric variables from the input table are used.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
inputsInheritFormats=TRUE | FALSE
specifies that the variables inherit formats from the underlying table.
| Default | FALSE |
|---|
mapInterval=list(mapIntervalPhase)
specifies the parameters to use for the map to interval transformation phase.
The mapIntervalPhase value can be one or more of the following:
arguments=list(mapIntervalArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The mapIntervalArguments value can be one or more of the following:
descending=TRUE | FALSE
specifies that the label count encoding be performed.
| Default | TRUE |
|---|
includeMissingLevel=TRUE | FALSE
when set to True, missing values are included in the distinct level analysis instead of being discarded.
| Default | FALSE |
|---|
nLevels=integer
specifies the number of target levels to consider for map-interval transformation. If the target has more levels than specified, the extra levels are ignored. If the target has less number of levels, missing values are generated.
| Default | 2 |
|---|---|
| Range | 1–10 |
nMoments=integer
specifies the number of centralized moments that replace the nominal value. The moments are, in order, the mean, the second, third and fourth order centralized moments.
| Default | 2 |
|---|---|
| Range | 1–6 |
noise=double
specifies the parameter for the Laplace or uniform noise to be added to the level statistics.
| Default | 0 |
|---|---|
| Minimum value (exclusive) | 0 |
shrinkageFactor=double
specifies the shrinkage factor for mapping the nominal values into interval value using the specified mapping criterion.
| Default | 0 |
|---|---|
| Minimum value | 0 |
woeAdjust=double
specifies the weight of evidence (WOE) adjustment factor.
| Default | 0.5 |
|---|
woeDefinition="EVENT" | "NONEVENT"
specifies the definition of WOE to use. If EVENT, then WOE is defined as event/nonevent. If NONEVENT, then WOE is defined as nonevent/event.
| Default | NONEVENT |
|---|
method="COUNTINPUT" | "COUNTTARGET" | "EMPBAYES" | "EVENTPROB" | "FREQRATIO" | "LABELCOUNT" | "MAX" | "MIN" | "MOMENTS" | "WOE"
specifies the interval map criterion to use.
| Alias | tech |
|---|---|
| Default | WOE |
name="string"
specifies a name for the request package.
outlier=list(outlierPhase)
specifies the parameters to use for the outlier determination and treatment phase.
The outlierPhase value can be one or more of the following:
arguments=list(outlierArguments)
specifies the arguments for this phase of the transform.
| Alias | args |
|---|
The outlierArguments value can be one or more of the following:
aadLocationUseMean=TRUE | FALSE
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | TRUE |
location="BIWEIGHT" | "GEOMETRICMEAN" | "HARMONICMEAN" | "MEAN" | "MEDIAN" | "TRIMMEDMEAN" | "WINSORIZEDMEAN"
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
lowerPercentile=double
specifies the lower percentile threshold (PERC outlier definition).
| Alias | lowerPerc |
|---|---|
| Range | (0, 50) |
max=double
specifies a global maximum value.
min=double
specifies a global minimum value.
replacements=list("BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN") | list(double-1 <, double-2, ...>)
specifies the values to use as replacements for outliers. These can be user defined values or location estimates.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
scale="AAD" | "BIWEIGHT" | "GINI" | "IQR" | "MAD" | "STD"
specifies the scale method to use.
| Default | STD |
|---|
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Minimum value (exclusive) | 0 |
scaleMultiplier=double
specifies the multiplying factor for the chosen scale estimator.
symmetricPercentile=double
specifies the symmetric percentile threshold to use. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | symPerc |
|---|---|
| Range | (0, 100) |
upperPercentile=double
specifies the upper percentile threshold to use.
| Alias | upperPerc |
|---|---|
| Range | (50, 100) |
userDefinedLimits=list(double-1 <, double-2, ...>)
uses the specified user-defined limits as the lower and upper thresholds for each variable.
zScoreThreshold=double
specifies the Z threshold.
method="IQR" | "MIQR" | "MZSCORE" | "PERC" | "UDFLIMITS" | "ZSCORE"
specifies the outlier definition.
IQR
uses the interquartile range to define outliers. Use the scaleMulFac parameter to set a multiplying factor.
MIQR
uses a robust interquartile range to define outliers. The robustification is accomplished by making the lower and upper thresholds depend exponentially a quantile skewness measure.
MZSCORE
uses the modified Z-score to define outliers. Use the scale, loc, locBiweightTuning, scaleBiweightTuning, aadLocUseMean, or scaleMulFac parameters to control the outlier definition.
PERC
uses percentiles to define outliers. Use the lowerPerc, upperPerc, or symPerc parameters to set the boundaries.
treatment="REPLACE" | "TRIM" | "WINSOR"
specifies the outlier treatment. If you specify a univariate technique for outDef, then you can choose a univariate treatment: TRIM or WINSOR.
output=list(outputPhase)
specifies the parameters to use for the output phase.
The outputPhase value can be one or more of the following:
noScoreCode=TRUE | FALSE
when set to True, no score code is sent to output.
| Default | FALSE |
|---|
noScoreTable=TRUE | FALSE
when set to True, no scoring is sent to output.
| Alias | noScoreTbl |
|---|---|
| Default | FALSE |
scoreWOE=TRUE | FALSE
when set to True, the weight of evidence (WOE) of the bin is used as the score value, instead of the bin id.
| Default | FALSE |
|---|
phaseOrder="FIO" | "FOI" | "IFO" | "IOF" | "OFI" | "OIF"
specifies the order for running the specified transformation phases. A phase must be specified for it to be included in the pipelining.
| Default | IOF |
|---|
targets=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies a list of target variables to use.
For more information about specifying the targets parameter, see the common casinvardesc parameter.
| Alias | evalVars |
|---|
targetsInheritFormats=TRUE | FALSE
specifies that the variables inherit formats from the underlying table.
| Default | FALSE |
|---|
sasVarNameLength=TRUE | FALSE
when set to True, the lengths of the names of the output variables are constrained to be less than or equal 32 characters.
| Default | FALSE |
|---|
saveState=list(casouttable)
specifies the settings for an output table that contains the transformation model table.
| Alias | saveModel |
|---|
| Long form | saveState=list(name="table-name") |
|---|---|
| Shortcut form | saveState="table-name" |
The casouttable value can be one or more of the following:
caslib="string"
specifies the name of the caslib for the output table.
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data.
lifetime=64-bit-integer
specifies the number of seconds to keep the table in memory after it is last accessed. The table is dropped if it is not accessed for the specified number of seconds.
| Default | 0 |
|---|---|
| Minimum value | 0 |
memoryFormat="DVR" | "INHERIT" | "STANDARD"
specifies the memory format for the output table.
| Default | INHERIT |
|---|
DVR
use the duplicate value reduction memory format. This memory format can reduce the memory consumption and file size when the input data contains duplicate values.
name="table-name"
specifies the name for the output table.
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
|---|
replace=TRUE | FALSE
when set to True, overwrites an existing table that has the same name.
| Default | FALSE |
|---|
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
* table=list(castable)
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
weight="variable-name"
specifies the weight variable.