Data Science Pilot Action Set
Automated Machine Learning (AutoML) Using the dsAutoMl Action
This section contains PROC CAS code.
Note: Input data must be accessible in your CAS session, either as a CAS table or as a transient-scope table. A CAS table has a two-level name: the first level is your CAS engine libref, and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 2, Shared Concepts (SAS Visual Data Mining and Machine Learning: Procedures). A transient-scope table is called directly from the action and exists in memory for the duration of the action. For more information about accessing data, see SAS Viya: System Programming Guide. For more information about PROC CAS and programming in CASL, see SAS Cloud Analytic Services: CASL Programmer’s Guide and SAS Cloud Analytic Services: CASL Reference.
The following DATA step creates the reference data table mycas.dmagecr in your CAS session. These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.
data mycas.dmagecr;
set sampsio.dmagecr;
run;
The following statements run the dsAutoMl action to automatically explore effective machine learning pipelines. The sampleSize parameter controls the total number of machine learning pipelines to construct, execute, and rank. The default policy settings are used for the explorationPolicy, screenPolicy, and selectionPolicy parameters. In contrast, the transformationPolicy parameter activates only the missing, cardinality, and skewness data-quality issues by default. This means that feature transformation and generation pipelines that are expected to alleviate these data-quality issues are executed by default. The transformationPolicy parameter in this example activates all data-quality issues except interaction (polynomial). The action produces four CAS output tables. The transformation_out table contains metadata about the feature transformation and generation pipelines. The feature_out table contains metadata about the generated features. The pipeline_out table contains metadata about the best-performing pipelines out of all the explored and executed pipelines. You can control the number of best-performing pipelines by using the topKPipelines parameter. The final output CAS table is the astore_out table, which contains the analytic store scoring object for the generated features.
proc cas;
loadactionset "dataSciencePilot";
dataSciencePilot.dsAutoMl
/ table = "DMAGECR"
target = "good_bad"
event = "good"
explorationPolicy = {}
screenPolicy = {}
selectionPolicy = {}
transformationPolicy = {missing = True,
cardinality = True,
entropy = True,
iqv = True,
skewness = True,
kurtosis = True,
Outlier = True
}
modelTypes = {"decisionTree"}
objective = "AUC"
sampleSize = 20
topKPipelines = 20
kFolds = 5
transformationOut = {name = "TRANSFORMATION_OUT",
replace = True}
featureOut = {name = "FEATURE_OUT",
replace = True}
pipelineOut = {name = "PIPELINE_OUT",
replace = True}
saveState = {name = "ASTORE_OUT",
replace = True}
;
run;
quit;
proc cas;
fetch / table = "PIPELINE_OUT";
run;
quit;
proc cas;
fetch / table = "FEATURE_OUT";
run;
quit;
proc cas;
fetch / table = "TRANSFORMATION_OUT";
run;
quit;
The table parameter names the input reference data table. The target parameter names the target variable. The explorationPolicy parameter names the data exploration policy; this example uses the default setting. The screenPolicy parameter names the variable screening policy; this example uses the default setting. The selectionPolicy parameter names the variable selection policy; this example uses the default setting. The transformationPolicy parameter names the feature transformation and generation policy. By selectively switching specific data-quality issues, you can control the amount of feature transformation exploration and generation that is done. The modelTypes parameter specifies the model types to consider in the pipeline sampling. The objective parameter specifies the objective metric to use during the hyperparameter tuning. The sampleSize parameter specifies the number of pipelines to sample. The topKPipelines parameter specifies the number of best-performing pipelines to include in the pipelines CAS output table. The kFolds parameter specifies the number of folds to use for cross validation to assess model fit error as the tuning objective. The transformationOut parameter names the transformation CAS output table. The featureOut parameter names the features CAS output table. The pipelineOut parameter names the pipelines CAS output table. The saveState parameter names the CAS output table to store the analytic store scoring object.
A sample of the pipelines table is shown in Output 10.3.1.
Output 10.3.1: Best-Performing Pipelines
| Selected Rows from Table PIPELINE_OUT | |||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| _Index_ | PipelineId | ModelType | MLType | Objective | ObjectiveType | Target | NFeatures | Feat1Id | Feat1IsNom | Feat2Id | Feat2IsNom | Feat3Id | Feat3IsNom | Feat4Id | Feat4IsNom | Feat5Id | Feat5IsNom | Feat6Id | Feat6IsNom | Feat7Id | Feat7IsNom | Feat8Id | Feat8IsNom | Feat9Id | Feat9IsNom | Feat10Id | Feat10IsNom | Feat11Id | Feat11IsNom | Feat12Id | Feat12IsNom | Feat13Id | Feat13IsNom | Feat14Id | Feat14IsNom | Feat15Id | Feat15IsNom | Feat16Id | Feat16IsNom | Feat17Id | Feat17IsNom | Feat18Id | Feat18IsNom | Feat19Id | Feat19IsNom | Feat20Id | Feat20IsNom |
| 1 | 9 | binary classification | dtree | 0.9710952381 | AUC | good_bad | 15 | 9 | 1 | 19 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 28 | 1 | 7 | 1 | 14 | 1 | 11 | 1 | 27 | 1 | 22 | 1 | . | . | . | . | . | . | . | . | . | . |
| 2 | 5 | binary classification | dtree | 0.94375 | AUC | good_bad | 10 | 9 | 1 | 19 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 3 | 13 | binary classification | dtree | 0.941 | AUC | good_bad | 14 | 9 | 1 | 19 | 1 | 1 | 0 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | 7 | 1 | 14 | 1 | 11 | 1 | 27 | 1 | . | . | . | . | . | . | . | . | . | . | . | . |
| 4 | 3 | binary classification | dtree | 0.9355714286 | AUC | good_bad | 10 | 9 | 1 | 19 | 1 | 3 | 1 | 34 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 30 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 5 | 16 | binary classification | dtree | 0.9328809524 | AUC | good_bad | 10 | 9 | 1 | 19 | 1 | 6 | 0 | 35 | 1 | 12 | 1 | 37 | 1 | 20 | 1 | 30 | 1 | 17 | 1 | 28 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 6 | 11 | binary classification | dtree | 0.9325595238 | AUC | good_bad | 10 | 9 | 1 | 18 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 7 | 15 | binary classification | dtree | 0.9305595238 | AUC | good_bad | 12 | 9 | 1 | 19 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | 7 | 1 | 14 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 8 | 2 | binary classification | dtree | 0.9290714286 | AUC | good_bad | 16 | 9 | 1 | 19 | 1 | 6 | 0 | 35 | 1 | 12 | 1 | 38 | 1 | 20 | 1 | 30 | 1 | 17 | 1 | 29 | 1 | 7 | 1 | 13 | 1 | 11 | 1 | 27 | 1 | 23 | 1 | 16 | 1 | . | . | . | . | . | . | . | . |
| 9 | 4 | binary classification | dtree | 0.9209404762 | AUC | good_bad | 10 | 9 | 1 | 19 | 1 | 1 | 0 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 30 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 10 | 10 | binary classification | dtree | 0.9141190476 | AUC | good_bad | 11 | 9 | 1 | 18 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | 7 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 11 | 18 | binary classification | dtree | 0.9039404762 | AUC | good_bad | 10 | 8 | 1 | 19 | 1 | 1 | 0 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 30 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 12 | 1 | binary classification | dtree | 0.9034761905 | AUC | good_bad | 20 | 8 | 1 | 18 | 1 | 3 | 1 | 6 | 0 | 34 | 1 | 12 | 1 | 37 | 1 | 20 | 1 | 30 | 1 | 17 | 1 | 28 | 1 | 7 | 1 | 13 | 1 | 10 | 1 | 26 | 1 | 22 | 1 | 15 | 1 | 24 | 1 | 36 | 1 | 32 | 1 |
| 13 | 19 | binary classification | dtree | 0.9027142857 | AUC | good_bad | 17 | 9 | 1 | 18 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | 7 | 1 | 14 | 1 | 11 | 1 | 27 | 1 | 23 | 1 | 16 | 1 | 25 | 1 | . | . | . | . | . | . |
| 14 | 17 | binary classification | dtree | 0.9011547619 | AUC | good_bad | 13 | 9 | 1 | 19 | 1 | 6 | 0 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 30 | 1 | 17 | 1 | 28 | 1 | 7 | 1 | 14 | 1 | 10 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 15 | 14 | binary classification | dtree | 0.8979285714 | AUC | good_bad | 10 | 9 | 1 | 19 | 1 | 1 | 0 | 35 | 1 | 12 | 1 | 38 | 1 | 20 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 16 | 6 | binary classification | dtree | 0.8905595238 | AUC | good_bad | 15 | 9 | 1 | 19 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 30 | 1 | 17 | 1 | 28 | 1 | 7 | 1 | 14 | 1 | 10 | 1 | 27 | 1 | 23 | 1 | . | . | . | . | . | . | . | . | . | . |
| 17 | 20 | binary classification | dtree | 0.8886309524 | AUC | good_bad | 10 | 9 | 1 | 18 | 1 | 2 | 1 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 18 | 7 | binary classification | dtree | 0.8798214286 | AUC | good_bad | 10 | 9 | 1 | 19 | 1 | 1 | 0 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 30 | 1 | 17 | 1 | 29 | 1 | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . | . |
| 19 | 12 | binary classification | dtree | 0.8415357143 | AUC | good_bad | 16 | 8 | 1 | 18 | 1 | 6 | 0 | 34 | 1 | 12 | 1 | 37 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 28 | 1 | 7 | 1 | 14 | 1 | 11 | 1 | 27 | 1 | 23 | 1 | 16 | 1 | . | . | . | . | . | . | . | . |
| 20 | 8 | binary classification | dtree | 0.8201547619 | AUC | good_bad | 18 | 9 | 1 | 19 | 1 | 1 | 0 | 35 | 1 | 12 | 1 | 38 | 1 | 21 | 1 | 31 | 1 | 17 | 1 | 29 | 1 | 7 | 1 | 14 | 1 | 11 | 1 | 27 | 1 | 22 | 1 | 16 | 1 | 25 | 1 | 36 | 1 | . | . | . | . |
Automated Machine Learning (AutoML) Using the dsAutoMl Action
This section contains Lua code for the analysis in the CASL version of this example, which contains details about the results.
Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:
s:loadtable{casLib="casuser", path="dmagecr.csv"}
For more information about coding in Lua, see Getting Started with SAS Viya for Lua and SAS Viya: System Programming Guide.
The following statements run the dsAutoMl action to automatically explore effective machine learning pipelines. The sampleSize parameter controls the total number of machine learning pipelines to construct, execute, and rank. The default policy settings are used for the explorationPolicy, screenPolicy, and selectionPolicy parameters. In contrast, the transformationPolicy parameter activates only the missing, cardinality, and skewness data-quality issues by default. This means that feature transformation and generation pipelines that are expected to alleviate these data-quality issues are executed by default. The transformationPolicy parameter in this example activates all data-quality issues except interaction (polynomial). The action produces four CAS output tables. The transformation_out table contains metadata about the feature transformation and generation pipelines. The feature_out table contains metadata about the generated features. The pipeline_out table contains metadata about the best-performing pipelines out of all the explored and executed pipelines. You can control the number of best-performing pipelines by using the topKPipelines parameter. The final output CAS table is the astore_out table, which contains the analytic store scoring object for the generated features.
s:loadactionset{actionset="dataSciencePilot"}
s:dataSciencePilot_dsAutoMl {
table = {name="DMAGECR"},
target = "good_bad"
explorationPolicy = {},
screenPolicy = {},
selectionPolicy = {},
transformationPolicy = {missing = True,
cardinality = True,
entropy = True,
iqv = True,
skewness = True,
kurtosis = True,
Outlier = True
},
modelTypes = {"decisionTree"},
objective = "AUC",
sampleSize = 20,
topKPipelines = 20,
kFolds = 5,
transformationOut = {name = "TRANSFORMATION_OUT",
replace = True},
featureOut = {name = "FEATURE_OUT",
replace = True},
pipelineOut = {name = "PIPELINE_OUT",
replace = True},
saveState = {name = "ASTORE_OUT",
replace = True}
}
s:fetch{table = {name = "PIPELINE_OUT"}}
s:fetch{table = {name = "FEATURE_OUT"}}
s:fetch{table = {name = "TRANSFORMATION_OUT"}}
The table parameter names the input reference data table. The target parameter names the target variable. The explorationPolicy parameter names the data exploration policy; this example uses the default setting. The screenPolicy parameter names the variable screening policy; this example uses the default setting. The selectionPolicy parameter names the variable selection policy; this example uses the default setting. The transformationPolicy parameter names the feature transformation and generation policy. By selectively switching specific data-quality issues, you can control the amount of feature transformation exploration and generation that is done. The modelTypes parameter specifies the model types to consider in the pipeline sampling. The objective parameter specifies the objective metric to use during the hyperparameter tuning. The sampleSize parameter specifies the number of pipelines to sample. The topKPipelines parameter specifies the number of best-performing pipelines to include in the pipelines CAS output table. The kFolds parameter specifies the number of folds to use for cross validation to assess model fit error as the tuning objective. The transformationOut parameter names the transformation CAS output table. The featureOut parameter names the features CAS output table. The pipelineOut parameter names the pipelines CAS output table. The saveState parameter names the CAS output table to store the analytic store scoring object.
For example output from the action, see the CASL code examples.
Automated Machine Learning (AutoML) Using the dsAutoMl Action
This section contains R code for the analysis in the CASL version of this example, which contains details about the results.
Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:
m <- cas.read.csv(s, "dmagecr.csv", casOut=list(name="dmagecr"))
For more information about coding in R, see Getting Started with SAS Viya for R and SAS Viya: System Programming Guide.
The following statements run the dsAutoMl action to automatically explore effective machine learning pipelines. The sampleSize parameter controls the total number of machine learning pipelines to construct, execute, and rank. The default policy settings are used for the explorationPolicy, screenPolicy, and selectionPolicy parameters. In contrast, the transformationPolicy parameter activates only the missing, cardinality, and skewness data-quality issues by default. This means that feature transformation and generation pipelines that are expected to alleviate these data-quality issues are executed by default. The transformationPolicy parameter in this example activates all data-quality issues except interaction (polynomial). The action produces four CAS output tables. The transformation_out table contains metadata about the feature transformation and generation pipelines. The feature_out table contains metadata about the generated features. The pipeline_out table contains metadata about the best-performing pipelines out of all the explored and executed pipelines. You can control the number of best-performing pipelines by using the topKPipelines parameter. The final output CAS table is the astore_out table, which contains the analytic store scoring object for the generated features.
cas.loadActionSet(s,"dataSciencePilot")
cas.dataSciencePilot.dsAutoMl(
s,
table = list(name ="DMAGECR"),
target = "good_bad",
explorationPolicy = list(),
screenPolicy = list(),
selectionPolicy = list(),
transformationPolicy = list(missing = True, cardinality = True,
entropy = True, iqv = True,
skewness = True, kurtosis = True, Outlier = True),
modelTypes = list("decisionTree"),
objective = "AUC",
sampleSize = 20,
topKPipelines = 20,
kFolds = 5,
transformationOut = list(name= "TRANSFORMATION_OUT", replace = True),
featureOut = list(name= "FEATURE_OUT", replace = True),
pipelineOut = list(name= "PIPELINE_OUT", replace = True),
saveState = list(name= "ASTORE_OUT", replace = True)
)
cas.fetch(s, table = list(name = "PIPELINE_OUT"))
cas.fetch(s, table = list(name = "FEATURE_OUT"))
cas.fetch(s, table = list(name = "TRANSFORMATION_OUT"))
The table parameter names the input reference data table. The target parameter names the target variable. The explorationPolicy parameter names the data exploration policy; this example uses the default setting. The screenPolicy parameter names the variable screening policy; this example uses the default setting. The selectionPolicy parameter names the variable selection policy; this example uses the default setting. The transformationPolicy parameter names the feature transformation and generation policy. By selectively switching specific data-quality issues, you can control the amount of feature transformation exploration and generation that is done. The modelTypes parameter specifies the model types to consider in the pipeline sampling. The objective parameter specifies the objective metric to use during the hyperparameter tuning. The sampleSize parameter specifies the number of pipelines to sample. The topKPipelines parameter specifies the number of best-performing pipelines to include in the pipelines CAS output table. The kFolds parameter specifies the number of folds to use for cross validation to assess model fit error as the tuning objective. The transformationOut parameter names the transformation CAS output table. The featureOut parameter names the features CAS output table. The pipelineOut parameter names the pipelines CAS output table. The saveState parameter names the CAS output table to store the analytic store scoring object.
For example output from the action, see the CASL code examples.
Automated Machine Learning (AutoML) Using the dsAutoMl Action
This section contains Python code for the analysis in the CASL version of this example, which contains details about the results.
Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:
s.upload_file('dmagecr.csv')
For more information about coding in Python, see Getting Started with SAS Viya for Python and SAS Viya: System Programming Guide.
The following statements run the dsAutoMl action to automatically explore effective machine learning pipelines. The sampleSize parameter controls the total number of machine learning pipelines to construct, execute, and rank. The default policy settings are used for the explorationPolicy, screenPolicy, and selectionPolicy parameters. In contrast, the transformationPolicy parameter activates only the missing, cardinality, and skewness data-quality issues by default. This means that feature transformation and generation pipelines that are expected to alleviate these data-quality issues are executed by default. The transformationPolicy parameter in this example activates all data-quality issues except interaction (polynomial). The action produces four CAS output tables. The transformation_out table contains metadata about the feature transformation and generation pipelines. The feature_out table contains metadata about the generated features. The pipeline_out table contains metadata about the best-performing pipelines out of all the explored and executed pipelines. You can control the number of best-performing pipelines by using the topKPipelines parameter. The final output CAS table is the astore_out table, which contains the analytic store scoring object for the generated features.
s.loadactionset(actionset="dataSciencePilot")
s.dataSciencePilot.dsAutoMl(
table = {"name" : "DMAGECR"},
target = "good_bad",
explorationPolicy = {},
screenPolicy = {},
selectionPolicy = {},
transformationPolicy = {"missing":True, "cardinality":True,
"entropy":True, "iqv":True,
"skewness":True, "kurtosis":True, "Outlier":True},
modelTypes = ["decisionTree"],
objective = "AUC",
sampleSize = 20,
topKPipelines = 20,
kFolds = 5,
transformationOut = {"name" : "TRANSFORMATION_OUT", "replace" : True},
featureOut = {"name" : "FEATURE_OUT", "replace" : True},
pipelineOut = {"name" : "PIPELINE_OUT", "replace" : True},
saveState = {"name" : "ASTORE_OUT", "replace" : True}
)
s.fetch(table = {"name": "PIPELINE_OUT"})
s.fetch(table = {"name": "FEATURE_OUT"})
s.fetch(table = {"name": "TRANSFORMATION_OUT"})
The table parameter names the input reference data table. The target parameter names the target variable. The explorationPolicy parameter names the data exploration policy; this example uses the default setting. The screenPolicy parameter names the variable screening policy; this example uses the default setting. The selectionPolicy parameter names the variable selection policy; this example uses the default setting. The transformationPolicy parameter names the feature transformation and generation policy. By selectively switching specific data-quality issues, you can control the amount of feature transformation exploration and generation that is done. The modelTypes parameter specifies the model types to consider in the pipeline sampling. The objective parameter specifies the objective metric to use during the hyperparameter tuning. The sampleSize parameter specifies the number of pipelines to sample. The topKPipelines parameter specifies the number of best-performing pipelines to include in the pipelines CAS output table. The kFolds parameter specifies the number of folds to use for cross validation to assess model fit error as the tuning objective. The transformationOut parameter names the transformation CAS output table. The featureOut parameter names the features CAS output table. The pipelineOut parameter names the pipelines CAS output table. The saveState parameter names the CAS output table to store the analytic store scoring object.
For example output from the action, see the CASL code examples.