RPCA Procedure
Getting Started: RPCA Procedure
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
This example shows how to use the RPCA procedure to decompose the input data set into a low-rank matrix and a sparse matrix. It also demonstrates how to decompose the low-rank matrix to obtain the principal components. In this example, the input data set rpcaData has 50 observations and three variables: Index, X, and Y. Index is simply an observation number. X is a randomly generated number between 0 and 1. Y is X plus a random term between and 0.1. This data set also has three outliers for variable
Y in observations 7, 20, and 33.
The following DATA step creates the rpcaData data set:
data rpcaData;
input index X Y;
datalines;
1 0.522 0.510
2 0.642 0.583
3 0.628 0.543
4 0.826 0.875
5 0.101 0.031
6 0.310 0.311
7 0.447 4.421
8 0.419 0.481
9 0.861 0.874
10 0.418 0.334
11 0.929 1.020
12 0.946 0.946
13 0.548 0.567
14 0.626 0.643
15 0.616 0.581
16 0.684 0.622
17 0.438 0.450
18 0.264 0.174
19 0.705 0.607
20 0.932 3.024
21 0.866 0.836
22 0.145 0.138
23 0.225 0.133
24 0.577 0.515
25 0.815 0.832
26 0.678 0.706
27 0.844 0.892
28 0.996 1.033
29 0.695 0.676
30 0.988 0.930
31 0.684 0.614
32 0.582 0.660
33 0.048 2.004
34 0.230 0.276
35 0.928 0.885
36 0.509 0.539
37 0.187 0.224
38 0.339 0.405
39 0.726 0.782
40 0.459 0.463
41 0.938 1.038
42 0.435 0.451
43 0.520 0.535
44 0.051 0.146
45 0.951 1.017
46 0.173 0.199
47 0.164 0.119
48 0.773 0.834
49 0.428 0.336
50 0.864 0.918
;
run;
You can load rpcaData into your CAS session by using your CAS engine libref in the first statement of the following DATA step:
data mylib.rpcaData;
set rpcaData;
run;
The following code runs PROC RPCA and outputs the decomposition results:
proc rpca data=mylib.rpcaData
decomp=svd
outlowrank=mylib.lowrankmat
outsparse=mylib.sparsemat;
id index;
outdecomp svdleft=mylib.svdleft
svddiag=mylib.svddiag
svdright=mylib.svdright;
run;
These statements produce the mylib.lowrankmat and mylib.sparsemat tables, which are the decompositions of mylib.rpcaData. They also produce the mylib.svdleft, mylib.svddiag, and mylib.svdright tables, which are the decompositions of the mylib.lowrankmat table. The DECOMP= option produces the SVD decomposition of the low-rank matrix.
Figure 1: ODS Tables
| Model Information | |
|---|---|
| Data Source | RPCADATA |
| RPCA Method | Augmented Lagrange Multiplier |
| SVD Method | Eigenvalue Decomposition |
| Lambda | 0.1414214 |
| Lambda Weight | 1 |
| Cumulative Eigenvalue Percent Tolerance Specified | 1 |
| Cumulative Eigenvalue Percent Tolerance Applied | 1 |
| Number of Observations | 50 |
|---|---|
| Number of Variables | 2 |
| Number of Observations Used | 50 |
| Number of Observations with Missing Values | 0 |
| Results Summary | |
|---|---|
| Sparsity of Sparse Matrix | 0.1600 |
| Rank of Low-Rank Matrix | 1 |
| Number of Iterations | 29 |
| Solution Status | Optimal |
| Run Time (Seconds) | 0.18 |
| Total SVD Time (Seconds) | 0.11 |
Figure 1 displays the "Model Information," "Dimensions," and "Results Summary" tables. The "Model Information" table shows the default parameters for the RPCA method, the SVD method, Lambda, and LambdaWeight. The "Dimensions" table shows the number of observations and variables in the input table and the number of observations that have missing values. (PROC RPCA ignores the observations that have missing values.) The "Results Summary" table shows that the solution status is optimal; that is, the RPCA algorithm converged based on the tolerance value (this example uses the default value of ) within the maximum number of iterations (this example uses the default value of 1,000). In the "Results Summary" table, you can also see that the algorithm converged after 29 iterations. Furthermore, you can observe in this table that the rank of the low-rank matrix is 1, as expected from the fact that variable
Y is highly correlated with variable X. Note that the sparsity value is 0.16, which indicates that the sparse matrix contains many nonzero values. As you increase the value of the LambdaWeight parameter, the sparsity of the sparse matrix increases.
The default value of the LambdaWeight parameter is 1. You can use the following PROC RPCA statement to change this parameter to 3.5. Figure 2 shows that the sparsity of the sparse matrix increases to 0.97 and the rank of the low-rank matrix increases to 2.
proc rpca data=mylib.rpcaData
lambdaweight = 3.5
outsparse=mylib.sparsemat2;
id index;
run;
proc print data=mylib.sparsemat2;
run;
Figure 2: Results Summary When LambdaWeight Is 3.5
| Results Summary | |
|---|---|
| Sparsity of Sparse Matrix | 0.9700 |
| Rank of Low-Rank Matrix | 2 |
| Number of Iterations | 17 |
| Solution Status | Optimal |
| Run Time (Seconds) | 0.10 |
| Total SVD Time (Seconds) | 0.06 |
If you look at mylib.sparsemat2, you can see that the only nonzero values in the sparse matrix are related to the outlier values that were introduced for variable Y (in observations 7, 20, and 33).
Alternatively, you can use the ANOMALYDETECTION statement in PROC RPCA and then use the ASTORE procedure to detect the outliers:
proc rpca data=mylib.rpcaData
scale center;
id index;
anomalydetection;
savestate rstore=mylib.store;
run;
proc astore;
setoption rpca_projection_type 2;
score rstore=mylib.store data=mylib.rpcaData out=mylib.scored;
run;
proc print data=mylib.scored;
run;
If you look at mylib.scored, you can see that the same outliers (observations 7, 20, and 33) are detected. The two methods that are described in this section are not exactly the same, but with proper settings they both detect outliers.