RPCA Procedure

Getting Started: RPCA Procedure

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.

This example shows how to use the RPCA procedure to decompose the input data set into a low-rank matrix and a sparse matrix. It also demonstrates how to decompose the low-rank matrix to obtain the principal components. In this example, the input data set rpcaData has 50 observations and three variables: Index, X, and Y. Index is simply an observation number. X is a randomly generated number between 0 and 1. Y is X plus a random term between negative 0.1 and 0.1. This data set also has three outliers for variable Y in observations 7, 20, and 33.

The following DATA step creates the rpcaData data set:

    data rpcaData;
    input index X Y;
    datalines;
    1    0.522    0.510
    2    0.642    0.583
    3    0.628    0.543
    4    0.826    0.875
    5    0.101    0.031
    6    0.310    0.311
    7    0.447    4.421
    8    0.419    0.481
    9    0.861    0.874
    10    0.418    0.334
    11    0.929    1.020
    12    0.946    0.946
    13    0.548    0.567
    14    0.626    0.643
    15    0.616    0.581
    16    0.684    0.622
    17    0.438    0.450
    18    0.264    0.174
    19    0.705    0.607
    20    0.932    3.024
    21    0.866    0.836
    22    0.145    0.138
    23    0.225    0.133
    24    0.577    0.515
    25    0.815    0.832
    26    0.678    0.706
    27    0.844    0.892
    28    0.996    1.033
    29    0.695    0.676
    30    0.988    0.930
    31    0.684    0.614
    32    0.582    0.660
    33    0.048    2.004
    34    0.230    0.276
    35    0.928    0.885
    36    0.509    0.539
    37    0.187    0.224
    38    0.339    0.405
    39    0.726    0.782
    40    0.459    0.463
    41    0.938    1.038
    42    0.435    0.451
    43    0.520    0.535
    44    0.051    0.146
    45    0.951    1.017
    46    0.173    0.199
    47    0.164    0.119
    48    0.773    0.834
    49    0.428    0.336
    50    0.864    0.918
    ;
    run;

You can load rpcaData into your CAS session by using your CAS engine libref in the first statement of the following DATA step:

    data mylib.rpcaData;
    set rpcaData;
    run;

The following code runs PROC RPCA and outputs the decomposition results:

 proc rpca data=mylib.rpcaData
           decomp=svd
           outlowrank=mylib.lowrankmat
           outsparse=mylib.sparsemat;
 id index;
 outdecomp svdleft=mylib.svdleft
           svddiag=mylib.svddiag
           svdright=mylib.svdright;
 run;

These statements produce the mylib.lowrankmat and mylib.sparsemat tables, which are the decompositions of mylib.rpcaData. They also produce the mylib.svdleft, mylib.svddiag, and mylib.svdright tables, which are the decompositions of the mylib.lowrankmat table. The DECOMP= option produces the SVD decomposition of the low-rank matrix.

Figure 1: ODS Tables

Model Information
Data SourceRPCADATA
RPCA MethodAugmented Lagrange Multiplier
SVD MethodEigenvalue Decomposition
Lambda0.1414214
Lambda Weight1
Cumulative Eigenvalue Percent Tolerance Specified1
Cumulative Eigenvalue Percent Tolerance Applied1

Number of Observations50
Number of Variables2
Number of Observations Used50
Number of Observations with Missing Values0

Results Summary
Sparsity of Sparse Matrix0.1600
Rank of Low-Rank Matrix1
Number of Iterations29
Solution StatusOptimal
Run Time (Seconds)0.18
Total SVD Time (Seconds)0.11


Figure 1 displays the "Model Information," "Dimensions," and "Results Summary" tables. The "Model Information" table shows the default parameters for the RPCA method, the SVD method, Lambda, and LambdaWeight. The "Dimensions" table shows the number of observations and variables in the input table and the number of observations that have missing values. (PROC RPCA ignores the observations that have missing values.) The "Results Summary" table shows that the solution status is optimal; that is, the RPCA algorithm converged based on the tolerance value (this example uses the default value of 10 Superscript negative 7) within the maximum number of iterations (this example uses the default value of 1,000). In the "Results Summary" table, you can also see that the algorithm converged after 29 iterations. Furthermore, you can observe in this table that the rank of the low-rank matrix is 1, as expected from the fact that variable Y is highly correlated with variable X. Note that the sparsity value is 0.16, which indicates that the sparse matrix contains many nonzero values. As you increase the value of the LambdaWeight parameter, the sparsity of the sparse matrix increases.

The default value of the LambdaWeight parameter is 1. You can use the following PROC RPCA statement to change this parameter to 3.5. Figure 2 shows that the sparsity of the sparse matrix increases to 0.97 and the rank of the low-rank matrix increases to 2.

 proc rpca data=mylib.rpcaData
           lambdaweight = 3.5
           outsparse=mylib.sparsemat2;
 id index;
 run;

 proc print data=mylib.sparsemat2;
 run;

Figure 2: Results Summary When LambdaWeight Is 3.5

Results Summary
Sparsity of Sparse Matrix0.9700
Rank of Low-Rank Matrix2
Number of Iterations17
Solution StatusOptimal
Run Time (Seconds)0.10
Total SVD Time (Seconds)0.06


If you look at mylib.sparsemat2, you can see that the only nonzero values in the sparse matrix are related to the outlier values that were introduced for variable Y (in observations 7, 20, and 33).

Alternatively, you can use the ANOMALYDETECTION statement in PROC RPCA and then use the ASTORE procedure to detect the outliers:

 proc rpca data=mylib.rpcaData
           scale center;
 id index;
 anomalydetection;
 savestate rstore=mylib.store;
 run;

 proc astore;
 setoption rpca_projection_type 2;
 score rstore=mylib.store data=mylib.rpcaData out=mylib.scored;
 run;

 proc print data=mylib.scored;
 run;

If you look at mylib.scored, you can see that the same outliers (observations 7, 20, and 33) are detected. The two methods that are described in this section are not exactly the same, but with proper settings they both detect outliers.

Last updated: September 04, 2026