Working with Decision Trees
- About Decision Trees
- Data Roles for a Decision Tree
- Options for a Decision Tree
- Derive a Leaf ID Data Item from a Decision Tree
- Display the Overview
- Zoom a Decision Tree
About Decision Trees
A decision tree uses the values of one or more predictor data items to predict the values of a response data item. A decision tree displays a series of nodes as a tree, where the top node is the response data item, and each branch of the tree represents a split in the values of a predictor data item. Decision trees are also known as classification and regression trees.

Each branch of the tree displays the name of the predictor for the branch at the top of the split. The thickness of the branch indicates the number of values that are associated with each node. The predictor values for each node are displayed above the node.
Each node in the tree displays the data for the node either as a histogram (if the response contains continuous data) or as a bar chart (if the response contains discrete data). The histogram or bar chart in each node displays the values of the response data item that are selected by the splits in the tree.
Below the decision tree, an icicle plot of the nodes is displayed. The color of the node in the icicle plot indicates either the predicted level for that node (for a discrete or binned response) or the aggregated value for the node (for an unbinned continuous response). When you select a node in either the decision tree or the icicle plot, the corresponding node is selected in the other location.
Decision trees in SAS Visual Analytics use a modified version of the C4.5 algorithm.
Data Roles for a Decision Tree
For information about setting data roles, see Working with Data Role Assignments in SAS Visual Analytics: Working with Report Data.
The basic data roles for a decision tree are:
- Response
-
specifies the response for the decision tree. You can specify any category or measure. The decision tree attempts to predict the values of the response data item. Each node in the tree displays the values of the response data item.
- Predictors
-
specifies predictors for the decision tree. You can specify one or more categories or measures as predictors. The values of predictor data items are displayed above the nodes in the tree. The order of the data items in the Predictors list does not affect the tree.
Note: If a predictor does not contribute to the predictive accuracy of the tree or the contribution has been pruned, then the predictor is not included in the final tree that is displayed.
Options for a Decision Tree
For information about general options, see Using the Options Pane.
In addition to the general options, you can specify the following object-specific options in the Options pane:
- Missing assignment
-
specifies how missing values are included in the model.
- None
-
observations with missing values are excluded from the model.
- Use in search
-
missing values are considered a unique measurement level and are included in the model.
- As machine smallest
-
missing interval values are set to the smallest possible machine value and missing category values are treated as a unique measurement level.
- Popular
-
observations with missing values are assigned to the child node with the most observations.
- Similar
-
observations with missing values are assigned to the node deemed most similar by a chi-square test for category responses or an F test for measure responses.
- Minimum value
-
specifies the minimum number of observations that are allowed to have missing values before missing values are treated as a distinct category level.
- Growth strategy
-
specifies the parameters that are used to create the decision tree. Select one of the following values:
- Basic
-
specifies a simple tree with a maximum of two branches per split and a maximum of six levels. For details, see Parameter Values for the Basic and Advanced Growth Strategies.
- Advanced
-
specifies a complex tree with a maximum of four branches per split and a maximum of six levels. For details, see Parameter Values for the Basic and Advanced Growth Strategies.
- Custom
-
enables you to select the values for each of the parameters.
If you select Custom as the value for Growth strategy, then the following additional options appear:
- Maximum branches
-
specifies the maximum number of branches for each node split.
- Maximum levels
-
specifies the maximum number of levels in the tree.
- Leaf size
-
specifies the minimum number of values (count) for each node.
- Predictor bins
-
specifies the number of bins that are used for predictor data items.
Note: This option has no effect if the predictor data items contain discrete data. - High-cardinality predictors
-
enables a greater cardinality limit for the categorical predictor values at each node. If the cardinality limit is exceeded for a node, then the predictor is not included in the split candidates for that node.
Note: If this option is not enabled, then the default cardinality limit is 128. - Pruning
-
specifies the level of pruning that is applied to the tree. Pruning removes leaves and branches that contribute the least to the predictive accuracy of the tree. A higher pruning value specifies that more leaves and branches are removed from the tree.
- Reuse predictors
-
specifies that predictors can be used more than once in a branch. That is, a predictor can be used more than once on the path leading to a leaf.
The following parameter values are used for the Basic and Advanced growth strategies:
|
Property |
Basic Value |
Advanced Value |
|---|---|---|
|
Maximum branches |
2 |
4 |
|
Maximum levels |
6 |
6 |
|
Leaf size |
1 |
1 |
|
Bin response variable |
No |
No |
|
Response bins |
10 |
10 |
|
Predictor bins |
2 |
10 |
|
Pruning |
0 |
0 |
|
Reuse predictors |
No |
Yes |
The following option is available under Tree Display:
- Displayed visuals
-
specifies the elements of the object that are displayed.
- Abbreviate numeric values
-
displays the measure values in an abbreviated form. For example, the value 1,142,571 might be abbreviated as 1.1M. The data tip for the measure always displays the full value. You can set the scale and precision of the abbreviated values.
The following options are available under Decision Tree / Icicle Plot:
- Statistic to show
-
specifies whether the nodes of the decision tree display the Count statistic or the Percent statistic. This statistic is also shown in the data tips for the nodes in the icicle plot.
- Legend visibility
-
specifies whether the legend for the decision tree is displayed.
Derive a Leaf ID Data Item from a Decision Tree
You can derive a leaf ID data item to represent the results of a decision tree. The leaf ID data item creates values that correspond to the node IDs in the details table for the decision tree.
You can use the leaf ID data item in a filter to select the values for a decision tree node in other types of objects.
To calculate a leaf ID data item from a decision tree:
- Click
on the object toolbar, and then select Derive a leaf ID variable.
- In the New Leaf ID window, enter a Name for the new derived item.
- Click OK to create the new data item.
Display the Overview
For large trees, the overview enables you to select the portions of the tree that are visible.
To display the overview, click on the object toolbar.
Zoom a Decision Tree
You can zoom a decision tree by scrolling the mouse wheel. The decision tree zooms in and out at the location of the pointer.
If you zoom out on a decision tree with a categorical response or a binned measure, then each leaf node displays a single bar for the greatest value in that node.
If you zoom out on a decision tree with an unbinned measure response, then the color of each leaf node indicates the average response value for the node.
When you have zoomed in on a decision tree, you can reposition the decision tree by clicking the tree and dragging it.