SAS Cloud Analytic Services: Troubleshooting
- Insufficient Disk Space (CAS_DISK_CACHE)
- Cannot Access Permstore
- CAS Pods in Pending Status
- Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error
- CAS Pods Restart in a Kubernetes 1.28 Cluster That Is Configured with Linux cgroup v2
- CAS Resources Are Defined Even Though CAS Servers Are Removed from the SAS Viya Programming-Only Deployment
Insufficient Disk Space (CAS_DISK_CACHE)
Issue: An
upload fails with a message such as Failed to
open temporary file for upload (80BFE801): /tmp/cascache1/_f_43d6c87c_7f5d854996e8.sas7bdat
Explanation: Insufficient disk space in CAS_DISK_CACHE on the CAS controller.
Resolution: Add more disk space.
Cannot Access Permstore
Issue: The
following message is displayed. A CAS server
is currently running and has exclusive access to the access control
storage location.
Explanation: This error message is displayed when a CAS server fails to access the permstore and terminates because the permstore is already being accessed by another CAS server.
CAS Pods in Pending Status
Issue: One
or more CAS pods are stuck in Pending status.
Explanation: The CAS pods that are in a pending state cannot currently be scheduled in the Kubernetes cluster. This state can occur if the available nodes have insufficient resources (memory or CPU), are missing the appropriate taints or labels, or the pod or node affinity rules are not met.
Resolution: If you describe the pod, the scheduler provides a reason for the pending state.
kubectl describe pods podName
Inadequate resources can be fixed by adjusting resource requests, deleting pods, or adding new nodes to your cluster. Ensure that nodes have the appropriate taints and labels and that all the affinity rules are met.
Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error
Issue: The following message
is displayed in the server log. Child terminated by signal: PID nnn,
signal 9, status 0x00000009. The syslog for the Kubernetes node(s)
where CAS pods are scheduled, also records out-of-memory (OOM) messages that varies
by the
provider.
Explanation: This error message indicates that a session was killed. One possible cause is that the OOM killer terminated a session because of high memory use.
Tips for reducing OOM error issues follow:
TIP 1:
Reducing the percentage of node memory that can be used for writeback (lazy writes) can reduce chances of cgroup OOM killer events for workloads that have heavy disk write activity to CAS node local storage like cascache.
You can modify the following two OS configurations:
- Reduce
vm.dirty_ratiofrom 20 to 10. - Reduce
vm.dirty_background_ratiofrom 10 to 5.
Refer the documentation of your respective
environment on how to modify these settings on CAS nodes. If you have access to the
nodes
where the CAS pods are running, then you can run ssh to each CAS
node, sudo to root, and run sysctl to modify
the suggested values.
TIP 2:
Reducing container memory request = limit to a
percentage (for example, 90%) of the Remaining Allocatable memory on the node can
help
mitigate issues that are due to a host OOM error. You can see the values of Allocatable
and
Allocated Memory by running
kubectl -n name-of-namespace describe node.
Note that by reducing this value, you might not have enough space for your in-memory
CAS
tables.
TIP 3:
To prevent OOM failures, you can also set a backing store to support memory allocations. For more information, see Backing Store for CAS Memory Allocations.
- All the tips provided also apply to the compute node if any of the SAS processes are being OOM killed.
- Kubernetes 1.32 has introduced a
configuration flag - singleprocessOOMKill to disable the OOM
group kill for
cgroup v2. For more information, see Kubelet Configuration. Contact your Kubernetes vendor for instructions on modifying kublet configuration flags.
CAS Pods Restart in a Kubernetes 1.28 Cluster That Is Configured with Linux cgroup v2
Explanation: Kubernetes
1.28 implemented a change in behavior that impacts CAS if your cluster is configured
with
Linux cgroup v2 and encounters a Linux out-of-memory (OOM) kill on
CAS sessions.
To determine whether your site is experiencing Linux OOM kill situations, see Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error.
Resolution:
- For information about how to tune CAS to avoid Linux OOM failures, see Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error.
- To prevent OOM failures, you can also set a backing store to support memory allocations. For more information, see Backing Store for CAS Memory Allocations.
- Kubernetes 1.32 has introduced a configuration
flag - singleprocessOOMKill to disable the OOM group kill for
cgroup v2. For more information, see Kubelet Configuration. Contact your Kubernetes vendor for instructions about modifying kubelet configuration flags.
CAS Resources Are Defined Even Though CAS Servers Are Removed from the SAS Viya Programming-Only Deployment
Explanation: If CAS Servers are not defined in your SAS Viya Programming-Only deployment, the following three CAS resources are still deployed:
sas-cas-controlpodsas-cas-operatorpodcasdeploymentCustom Resource Definition (CRD)
Resolution: These resources are used by other services in your SAS Viya Programming-Only deployment and must remain deployed. The pods require minimal resources and do not require a dedicated CAS NodePool.