RECENGINE Procedure

Getting Started: RECENGINE Procedure

Note: Input data must be in a table that is accessible in your session. You can refer to this table by using a two-level name. The first level is a libref, and the second level is the table name. For more information, see the section Using SAS Viya Workbench in Chapter 2, Shared Concepts.

This example shows how to use the RECENGINE procedure to train a recommender system model on observations in a data table. The usersCommunity data set includes a subset of the click events over a period of time on a peer-to-peer-support community website. Each web page (item) on the site belongs to a message board that includes a number of topics of specific interest. The same user might have viewed the same web page multiple times. Also, in this data set, it is indicated whether or not the questions within the topic have been resolved. For each click event, the user ID and item ID are recorded. This data set includes 40,000 click events (observations).

The following code includes the DATA step that generates the data table mylib.usersCommunity. You can find the complete DATA set here:

https://support.sas.com/documentation/onlinedoc/viya/examples.htm

You can load the usersCommunity data table into your SAS library by using your libref as in the following statements:

   data mylib.usersCommunity;
      input itemid boardid userid isSolvedTopic;
      datalines;
      38 4817 6313 0
      15 7589 1355 1
      46 9560 2926 1
      16 11635 3985 1
      38 13023 4163 0

   ... more lines ...   

The following statements train a recommender system model on the usersCommunity data by using PROC RECENGINE:

proc recengine data=mylib.usersCommunity method=bpr nthreads=1 outmodel=mylib.factors;
   input userid itemid/level=nominal;
   userid userid;
   itemid itemid;
   bpr maxIter=20 nFactors=10 learnStep=0.01 regl2=0.01 seed=1;
   savestate rstore=mylib.state;
run;

The following statements print the first 10 observations in the mylib.factors data table, which includes some of the rows that are related to the latent factors for users. The output is shown in Figure 1.

proc print data=mylib.factors(obs=10);
run;

Figure 1: User Factors

ObsVariableLevelBiasFactor1Factor2Factor3Factor4Factor5Factor6Factor7Factor8Factor9Factor10
1_GLOBAL_ 00.000000.000000.000000.000000.000000.000000.000000.000000.000000.00000
2itemid80-1.124854.451922.19208-0.60274-1.776422.14684-0.59482-0.390891.393982.05710
3itemid900.036243.21847-3.677824.29316-1.045343.837724.33494-1.091290.34840-2.91478
4itemid100-2.783281.86587-2.04741-2.00398-1.135180.82657-1.10613-2.48457-1.705730.79410
5itemid1100.129720.736631.277582.48206-1.89837-0.457072.19836-3.74934-2.91389-3.74224
6itemid1506.11635-2.34953-2.74044-4.51393-0.152030.50069-2.161381.98882-3.959321.74290
7itemid1601.316431.507211.40146-1.93326-2.931207.886540.02650-1.51039-2.15252-3.79505
8itemid170-1.974163.98234-0.82852-1.04245-1.027380.37425-0.77400-1.92542-4.39585-0.70957
9itemid1804.591440.40213-1.29228-2.76491-2.50467-1.087771.557070.42976-2.949951.10726
10itemid2201.71007-2.33381-2.07336-4.04132-2.214382.65438-0.57034-1.43373-2.61990-2.00019


The following statements print the last 10 observations in the mylib.factors data table, which includes some of the rows that are related to the latent factors for items. The output is shown in Figure 2.


data _null_;
   if 0 then set mylib.factors nobs=n;
   call symputx("n",left(put(n,best.)),'l');
run;

proc print data=mylib.factors(firstobs=%eval(&n-10));
run;

Figure 2: Item Factors

ObsVariableLevelBiasFactor1Factor2Factor3Factor4Factor5Factor6Factor7Factor8Factor9Factor10
5530userid64930-2.05324-0.74745-0.43335-0.142071.27943-1.48819-0.505193.756411.751140.92086
5531userid64940-0.99664-1.520750.152800.748072.66734-0.66143-1.113150.713711.502152.22022
5532userid64950-1.401960.28836-1.222951.401261.689500.66801-2.454291.668350.68817-0.89730
5533userid64960-1.74555-0.908890.537172.17322-1.58115-2.26984-1.048631.60068-0.18941-0.89444
5534userid64970-1.333610.22481-1.377310.35985-1.27575-1.15299-1.01390-0.489782.410421.13173
5535userid64980-1.22752-0.13601-0.94915-0.380832.77113-2.06585-1.59524-0.204392.107121.84687
5536userid64990-2.44133-0.20218-1.288711.655041.48795-1.32646-1.474511.665301.579931.82965
5537userid65000-0.813071.30038-0.71154-0.000721.57934-1.71881-0.653672.541620.447670.77166
5538userid65010-1.50311-0.84557-0.607520.900242.62379-1.63011-0.525992.068891.774512.02186
5539userid65020-0.27437-1.01527-0.037641.409360.530611.079370.907070.13427-1.68125-0.64236
5540userid65030-1.75542-0.40282-1.865571.616920.57108-2.036550.707141.12797-1.052331.49678


You can use the mylib.factors output table to rank items for each user. The following code shows how to find the top two items for a sample user (that is, the one whose user ID is 6313).

First, you create the scoring data table that includes the ID of the sample user of interest:

   data mylib.user6313;
   input   userid;
   datalines;
   6313
   ;
   run;

Next, you use the ASTORE procedure to score this user:


proc astore;
   setoption REC_TOP_N 2;
   score data=mylib.user6313 rstore=mylib.state out=mylib.rankedItems;
   run;
quit;

The following statements print the top two recommended items for user 6313. The output is shown in Figure 3.

proc print data=mylib.rankedItems;
run;

Figure 3: Top Two Items for User 6313

Obsuserid_RECOMMENDED_itemid_RANK_SCORE
1631316112.1376
2631315210.2563


Last updated: September 23, 2026