NETWORK Procedure

Example 2.12 Connected Components for US Patent Citations

This example looks at the structural relationship of US patent citations by using a large data set that is maintained by the Stanford Network Analysis Project (SNAP) (Leskovec 2014). The citation graph includes over 16 million citations made to patents between 1975 and 1999.

The following statements construct the links data table mylib.Patents from a local copy of the raw patent citation data:

filename in 'cit-Patents.txt';
data mylib.Patents;
   infile in firstobs=5 dlm='09'X;
   input from to;
run;

The following statements find the connected components of the citation graph by using a distributed union-find algorithm. This algorithm takes advantage of all the machines in your configured session.

proc network
   links        = mylib.Patents
   outNodes     = mylib.NodeSetOut
   distributed  = true;
   connectedComponents
      out       = mylib.ConCompOut;
run;
%put &_NETWORK_;

The progress of the procedure is shown in Output 2.12.1.

Output 2.12.1: PROC NETWORK Log: Connected Components for US Patent Citations

NOTE: ------------------------------------------------------------------------------------------
NOTE: Running NETWORK.                                                                          
NOTE: ------------------------------------------------------------------------------------------
NOTE: The number of nodes in the input graph is 3774768.                                        
NOTE: The number of links in the input graph is 16518948.                                       
NOTE: Processing connected components using 256 threads across 16 machines.                     
NOTE: The graph has 3627 connected components.                                                  
NOTE: Processing connected components used 1.50 (cpu: 7.16) seconds.                            
NOTE: The Cloud Analytic Services server processed the request in 3.615459 seconds.             
NOTE: The data set MYLIB.NODESETOUT has 3774768 observations and 2 variables.                   
NOTE: The data set MYLIB.CONCOMPOUT has 3627 observations and 2 variables.                      
STATUS=OK  PROBLEM_TYPE=CONNECTEDCOMPONENTS  SOLUTION_STATUS=OK  NUM_COMPONENTS=3627            
CPU_TIME=37.93  REAL_TIME=3.62                                                                  


The 10 biggest components are shown in Output 2.12.2. It is interesting to note that the vast majority of patents (over 99%) are all contained in the same component. This is not too surprising, because many of the seminal patent claims are required in order to understand subsequent inventions.

Output 2.12.2: Ten Largest Components for US Patent Citations

concompnodes
13764117
6919
29516
30115
4314
16314
268014
322214
19413
128313


Last updated: August 07, 2026