Sampling and Partitioning Action Set
Segment-Stratified Sampling for Each Target When Frequency Variable Is Specified
This section contains PROC CAS code.
Note: Input data must be accessible in your CAS session, either as a CAS table or as a transient-scope table. A CAS table has a two-level name: the first level is your CAS engine libref, and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 2, Shared Concepts. A transient-scope table is called directly from the action and exists in memory for the duration of the action. For more information about accessing data, see SAS Viya: System Programming Guide. For more information about PROC CAS and programming in CASL, see SAS Cloud Analytic Services: CASL Programmer’s Guide and SAS Cloud Analytic Services: CASL Reference.
This example performs segment-stratified partitioning of 5,960 fictitious mortgages for each target variable when a frequency variable is specified, with the groupby parameter bad used as the segment variable. This example essentially performs two stratified partitions simultaneously for every stratum that is defined by the level combinations of each target with the variable that is specified in the groupby parameter. The input data table mycas.hmeq includes information about fictitious mortgages. Each observation represents an applicant for a home equity loan, and all applicants have an existing mortgage.
You can load the sampsio.hmeq data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step:
data mycas.hmeq;
set sampsio.hmeq;
id = _N_;
freqVar = 10;
run;
This DATA step assumes that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref. Note that the frequency variable freqVar is set to 10 for all the observations in the hmeq data set in the CAS engine libref mycas.
The following statements load the sampling action set and then use the stratified action to partition the mycas.hmeq data table for each stratum while considering the frequency of each observation:
proc cas;
loadactionset "sampling";
action stratified result=r/table={name="hmeq",groupby={"bad"}}
samppct=10, samppct2=20, seed=10, target={"job","reason"}, freq="freqVar",
outputTables={names={STRAFreq="straf",SampleFreqMap="sfmap"}},
output={casout={name="out",replace="TRUE"},
copyvars={"id","job","reason","bad","loan","value","delinq","derog"}};
run;
print r.STRAFreq;
print r.SampleFreqMap;
run;
quit;
proc sort data= mycas.out out= myout;
by id;
run;
proc print data=myout(obs=20);
run;
The table parameter names the input data table to be analyzed. The groupby subparameter in the table parameter names the variables to use for segmentation. The samppct parameter requests that 10% of the input data be included in the training partition, and the samppct2 parameter requests that 20% of the input data be included in the testing partition. The seed parameter specifies 10 as the random seed to use in the partitioning process. The by parameter requests that the variable bad be used as the segment variable. The target parameter requests that stratified sampling be performed for each target variable job and reason in each segment level. The freq parameter requests that the variable freqVar be used as the frequency variable. In this example, the action treats each observation as if it appears 10 times. The outputTables parameter outputs the frequency table to the mycas.straf data table and outputs the sample frequency map table to the mycas.sfmap data table. The output parameter requests that the sampled data be stored in a table named mycas.out, and the copyvars parameter lists the variables to be copied from mycas.hmeq to mycas.out. The output data table, mycas.out, also includes one partition indicator variable that shows whether each observation is selected for a partition (1 for training, 2 for testing, or 0 for not selected), as well as two sample frequency variables (one for each target variable in each segment level in the BY variable bad).
Output 22.6.1 shows the frequency information for each level of stratification that is defined by the level combination of the groupby variable bad and each target variable in the mycas.hmeq data table.
Output 22.6.1: Frequency Information Table
| Stratified Sampling Frequency | |||||
|---|---|---|---|---|---|
| Target Name | Target Level | BAD | Number of Obs | Sample Size 1 | Sample Size 2 |
| JOB | 0 | 2560 | 256 | 512 | |
| JOB | Mgr | 0 | 5880 | 588 | 1176 |
| JOB | Office | 0 | 8230 | 823 | 1646 |
| JOB | Other | 0 | 18340 | 1834 | 3668 |
| JOB | ProfExe | 0 | 10640 | 1064 | 2128 |
| JOB | Sales | 0 | 710 | 71 | 142 |
| JOB | Self | 0 | 1350 | 135 | 270 |
| JOB | 1 | 230 | 23 | 46 | |
| JOB | Mgr | 1 | 1790 | 179 | 358 |
| JOB | Office | 1 | 1250 | 125 | 250 |
| JOB | Other | 1 | 5540 | 554 | 1108 |
| JOB | ProfExe | 1 | 2120 | 212 | 424 |
| JOB | Sales | 1 | 380 | 38 | 76 |
| JOB | Self | 1 | 580 | 58 | 116 |
| REASON | 0 | 2040 | 204 | 408 | |
| REASON | DebtCon | 0 | 31830 | 3183 | 6366 |
| REASON | HomeImp | 0 | 13840 | 1384 | 2768 |
| REASON | 1 | 480 | 48 | 96 | |
| REASON | DebtCon | 1 | 7450 | 745 | 1490 |
| REASON | HomeImp | 1 | 3960 | 396 | 792 |
Output 22.6.2 shows the map of each sample frequency to each target.
Output 22.6.2: Sample Frequency Map
| Sample Frequency Map | |
|---|---|
| Sample Frequency Name | Target Name |
| _SampleFreq1_ | JOB |
| _SampleFreq2_ | REASON |
Output 22.6.3 shows the first 20 output sample observations in mycas.out. Each observation is duplicated by three rows in the output, with the _PartInd_ column showing which partition each observation is selected for (1 for training, 2 for testing, or 0 for not selected). The _SampleFreq1_ column shows the sample frequency for each partition for the target variable job. The _SampleFreq2_ column shows the sample frequency for each partition for the target variable reason. Note that the rows for which all the _SampleFreq_ columns have a value of 0 are not included in the output.
Output 22.6.3: Sample Output with Partition Indicator and Sample Frequency for Each Target
| Obs | id | JOB | REASON | BAD | LOAN | VALUE | DELINQ | DEROG | _PartInd_ | _SampleFreq1_ | _SampleFreq2_ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | Other | HomeImp | 1 | 1100 | 39025 | 0 | 0 | 2 | 1 | 1 |
| 2 | 1 | Other | HomeImp | 1 | 1100 | 39025 | 0 | 0 | 0 | 8 | 7 |
| 3 | 1 | Other | HomeImp | 1 | 1100 | 39025 | 0 | 0 | 1 | 1 | 2 |
| 4 | 2 | Other | HomeImp | 1 | 1300 | 68400 | 2 | 0 | 2 | 1 | 0 |
| 5 | 2 | Other | HomeImp | 1 | 1300 | 68400 | 2 | 0 | 0 | 9 | 7 |
| 6 | 2 | Other | HomeImp | 1 | 1300 | 68400 | 2 | 0 | 1 | 0 | 3 |
| 7 | 3 | Other | HomeImp | 1 | 1500 | 16700 | 0 | 0 | 0 | 8 | 4 |
| 8 | 3 | Other | HomeImp | 1 | 1500 | 16700 | 0 | 0 | 1 | 1 | 2 |
| 9 | 3 | Other | HomeImp | 1 | 1500 | 16700 | 0 | 0 | 2 | 1 | 4 |
| 10 | 4 | 1 | 1500 | . | . | . | 2 | 0 | 3 | ||
| 11 | 4 | 1 | 1500 | . | . | . | 0 | 9 | 6 | ||
| 12 | 4 | 1 | 1500 | . | . | . | 1 | 1 | 1 | ||
| 13 | 5 | Office | HomeImp | 0 | 1700 | 112000 | 0 | 0 | 2 | 2 | 3 |
| 14 | 5 | Office | HomeImp | 0 | 1700 | 112000 | 0 | 0 | 1 | 1 | 1 |
| 15 | 5 | Office | HomeImp | 0 | 1700 | 112000 | 0 | 0 | 0 | 7 | 6 |
| 16 | 6 | Other | HomeImp | 1 | 1700 | 40320 | 0 | 0 | 1 | 1 | 1 |
| 17 | 6 | Other | HomeImp | 1 | 1700 | 40320 | 0 | 0 | 2 | 3 | 2 |
| 18 | 6 | Other | HomeImp | 1 | 1700 | 40320 | 0 | 0 | 0 | 6 | 7 |
| 19 | 7 | Other | HomeImp | 1 | 1800 | 57037 | 2 | 3 | 0 | 8 | 8 |
| 20 | 7 | Other | HomeImp | 1 | 1800 | 57037 | 2 | 3 | 1 | 1 | 1 |
Segment-Stratified Sampling for Each Target When Frequency Variable Is Specified
This example is not available for the Lua programming language.
Segment-Stratified Sampling for Each Target When Frequency Variable Is Specified
This example is not available for the Python programming language.
Segment-Stratified Sampling for Each Target When Frequency Variable Is Specified
This example is not available for the R programming language.