RECENGINE Procedure
Getting Started: RECENGINE Procedure
Note: Input data must be in a table that is accessible in your session. You can refer to this table by using a two-level name. The first level is a libref, and the second level is the table name. For more information, see the section Using SAS Viya Workbench in Chapter 2, Shared Concepts.
This example shows how to use the RECENGINE procedure to train a recommender system model on observations in a data table. The usersCommunity data set includes a subset of the click events over a period of time on a peer-to-peer-support community website. Each web page (item) on the site belongs to a message board that includes a number of topics of specific interest. The same user might have viewed the same web page multiple times. Also, in this data set, it is indicated whether or not the questions within the topic have been resolved. For each click event, the user ID and item ID are recorded. This data set includes 40,000 click events (observations).
The following code includes the DATA step that generates the data table mylib.usersCommunity. You can find the complete DATA set here:
https://support.sas.com/documentation/onlinedoc/viya/examples.htm
You can load the usersCommunity data table into your SAS library by using your libref as in the following statements:
data mylib.usersCommunity;
input itemid boardid userid isSolvedTopic;
datalines;
38 4817 6313 0
15 7589 1355 1
46 9560 2926 1
16 11635 3985 1
38 13023 4163 0
... more lines ...
The following statements train a recommender system model on the usersCommunity data by using PROC RECENGINE:
proc recengine data=mylib.usersCommunity method=bpr nthreads=1 outmodel=mylib.factors;
input userid itemid/level=nominal;
userid userid;
itemid itemid;
bpr maxIter=20 nFactors=10 learnStep=0.01 regl2=0.01 seed=1;
savestate rstore=mylib.state;
run;
The following statements print the first 10 observations in the mylib.factors data table, which includes some of the rows that are related to the latent factors for users. The output is shown in Figure 1.
proc print data=mylib.factors(obs=10);
run;
Figure 1: User Factors
| Obs | Variable | Level | Bias | Factor1 | Factor2 | Factor3 | Factor4 | Factor5 | Factor6 | Factor7 | Factor8 | Factor9 | Factor10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | _GLOBAL_ | 0 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | 0.00000 | |
| 2 | itemid | 8 | 0 | -1.12485 | 4.45192 | 2.19208 | -0.60274 | -1.77642 | 2.14684 | -0.59482 | -0.39089 | 1.39398 | 2.05710 |
| 3 | itemid | 9 | 0 | 0.03624 | 3.21847 | -3.67782 | 4.29316 | -1.04534 | 3.83772 | 4.33494 | -1.09129 | 0.34840 | -2.91478 |
| 4 | itemid | 10 | 0 | -2.78328 | 1.86587 | -2.04741 | -2.00398 | -1.13518 | 0.82657 | -1.10613 | -2.48457 | -1.70573 | 0.79410 |
| 5 | itemid | 11 | 0 | 0.12972 | 0.73663 | 1.27758 | 2.48206 | -1.89837 | -0.45707 | 2.19836 | -3.74934 | -2.91389 | -3.74224 |
| 6 | itemid | 15 | 0 | 6.11635 | -2.34953 | -2.74044 | -4.51393 | -0.15203 | 0.50069 | -2.16138 | 1.98882 | -3.95932 | 1.74290 |
| 7 | itemid | 16 | 0 | 1.31643 | 1.50721 | 1.40146 | -1.93326 | -2.93120 | 7.88654 | 0.02650 | -1.51039 | -2.15252 | -3.79505 |
| 8 | itemid | 17 | 0 | -1.97416 | 3.98234 | -0.82852 | -1.04245 | -1.02738 | 0.37425 | -0.77400 | -1.92542 | -4.39585 | -0.70957 |
| 9 | itemid | 18 | 0 | 4.59144 | 0.40213 | -1.29228 | -2.76491 | -2.50467 | -1.08777 | 1.55707 | 0.42976 | -2.94995 | 1.10726 |
| 10 | itemid | 22 | 0 | 1.71007 | -2.33381 | -2.07336 | -4.04132 | -2.21438 | 2.65438 | -0.57034 | -1.43373 | -2.61990 | -2.00019 |
The following statements print the last 10 observations in the mylib.factors data table, which includes some of the rows that are related to the latent factors for items. The output is shown in Figure 2.
data _null_;
if 0 then set mylib.factors nobs=n;
call symputx("n",left(put(n,best.)),'l');
run;
proc print data=mylib.factors(firstobs=%eval(&n-10));
run;
Figure 2: Item Factors
| Obs | Variable | Level | Bias | Factor1 | Factor2 | Factor3 | Factor4 | Factor5 | Factor6 | Factor7 | Factor8 | Factor9 | Factor10 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 5530 | userid | 6493 | 0 | -2.05324 | -0.74745 | -0.43335 | -0.14207 | 1.27943 | -1.48819 | -0.50519 | 3.75641 | 1.75114 | 0.92086 |
| 5531 | userid | 6494 | 0 | -0.99664 | -1.52075 | 0.15280 | 0.74807 | 2.66734 | -0.66143 | -1.11315 | 0.71371 | 1.50215 | 2.22022 |
| 5532 | userid | 6495 | 0 | -1.40196 | 0.28836 | -1.22295 | 1.40126 | 1.68950 | 0.66801 | -2.45429 | 1.66835 | 0.68817 | -0.89730 |
| 5533 | userid | 6496 | 0 | -1.74555 | -0.90889 | 0.53717 | 2.17322 | -1.58115 | -2.26984 | -1.04863 | 1.60068 | -0.18941 | -0.89444 |
| 5534 | userid | 6497 | 0 | -1.33361 | 0.22481 | -1.37731 | 0.35985 | -1.27575 | -1.15299 | -1.01390 | -0.48978 | 2.41042 | 1.13173 |
| 5535 | userid | 6498 | 0 | -1.22752 | -0.13601 | -0.94915 | -0.38083 | 2.77113 | -2.06585 | -1.59524 | -0.20439 | 2.10712 | 1.84687 |
| 5536 | userid | 6499 | 0 | -2.44133 | -0.20218 | -1.28871 | 1.65504 | 1.48795 | -1.32646 | -1.47451 | 1.66530 | 1.57993 | 1.82965 |
| 5537 | userid | 6500 | 0 | -0.81307 | 1.30038 | -0.71154 | -0.00072 | 1.57934 | -1.71881 | -0.65367 | 2.54162 | 0.44767 | 0.77166 |
| 5538 | userid | 6501 | 0 | -1.50311 | -0.84557 | -0.60752 | 0.90024 | 2.62379 | -1.63011 | -0.52599 | 2.06889 | 1.77451 | 2.02186 |
| 5539 | userid | 6502 | 0 | -0.27437 | -1.01527 | -0.03764 | 1.40936 | 0.53061 | 1.07937 | 0.90707 | 0.13427 | -1.68125 | -0.64236 |
| 5540 | userid | 6503 | 0 | -1.75542 | -0.40282 | -1.86557 | 1.61692 | 0.57108 | -2.03655 | 0.70714 | 1.12797 | -1.05233 | 1.49678 |
You can use the mylib.factors output table to rank items for each user. The following code shows how to find the top two items for a sample user (that is, the one whose user ID is 6313).
First, you create the scoring data table that includes the ID of the sample user of interest:
data mylib.user6313;
input userid;
datalines;
6313
;
run;
Next, you use the ASTORE procedure to score this user:
proc astore;
setoption REC_TOP_N 2;
score data=mylib.user6313 rstore=mylib.state out=mylib.rankedItems;
run;
quit;
The following statements print the top two recommended items for user 6313. The output is shown in Figure 3.
proc print data=mylib.rankedItems;
run;
Figure 3: Top Two Items for User 6313
| Obs | userid | _RECOMMENDED_itemid | _RANK | _SCORE |
|---|---|---|---|---|
| 1 | 6313 | 16 | 1 | 12.1376 |
| 2 | 6313 | 15 | 2 | 10.2563 |