Pod Affinity
Pod affinity allows you to influence where a new Pod is scheduled by specifying rules based on the labels of existing Pods. The goal is to co-locate Pods on the same node.
Example: Co-locating Pods
First, deploy a Pod with specific labels:
apiVersion: v1
kind: Pod
metadata:
name: nginx-pod
labels:
app: webserver
tier: frontend
spec:
containers:
- name: nginx
image: nginx:latest
Then, deploy another Pod that should be scheduled on the same node as the webserver Pod:
apiVersion: v1
kind: Pod
metadata:
name: redis-pod
labels:
app: database
tier: backend
spec:
affinity:
podAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app
operator: In
values:
- webserver
topologyKey: "kubernetes.io/hostname"
containers:
- name: redis
image: redis:7.2.5
When you apply these configurations, Kubernetes' scheduler will attempt to place redis-pod on the same node as nginx-pod, provided the webserver label matches.
Pod Anti-Affinity
Pod anti-affinity achieves the opposite of affinity; it aims to prevent new Pods from being scheduled on the same node as specific existing Pods.
Example: Avoiding Co-location
Deploy a Pod that will serve as a target for anti-affinity:
apiVersion: v1
kind: Pod
metadata:
name: worker-pod-1
labels:
app: batch-job
worker: worker-1
spec:
containers:
- name: worker
image: ubuntu:latest
Now, create a new Pod that should not run on the same node as worker-pod-1:
apiVersion: v1
kind: Pod
metadata:
name: worker-pod-2
labels:
app: batch-job
worker: worker-2
spec:
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: worker
operator: In
values:
- worker-1
topologyKey: "kubernetes.io/hostname"
containers:
- name: worker
image: ubuntu:latest
With this configuration, the scheduler will ensure that worker-pod-2 is placed on a different node than worker-pod-1.
Taints and Tolerations
Taints and tolerations work together to control Pod scheduling by making nodes repel Pods or by allowing Pods to tolerate certain node attributes.
- Taints are applied to nodes, marking them as undesirable for certain Pods.
- Tolerations are applied to Pods, allowing them to be scheduled on nodes with matching taints.
Taint Effects:
- NoSchedule: Pods without a matching toleration will not be scheduled on the node. Existing Pods are unaffected.
- PreferNoSchedule: The scheduler will try to avoid scheduling Pods without a matching toleration, but may do so if no other suitable nodes are available.
- NoExecute: Pods without a matching toleration will not be scheduled, and any existing Pods on the node that do not tolerate the taint will be evicted.
Applying a Taint:
To apply a taint to a node:
kubectl taint node <node-name> <key>=<value>:<effect>
Example:
kubectl taint node node-1 env=prod:NoSchedule
Removing a Taint:
kubectl taint node <node-name> <key>:<effect>-
Pod Tolerations:
Tolerations are defined within the Pod's specification:
apiVersion: v1
kind: Pod
metadata:
name: tolerated-pod
spec:
containers:
- name: app
image: my-app
tolerations:
- key: "env"
operator: "Equal"
value: "prod"
effect: "NoSchedule"
This Pod can now be scheduled on nodes with the env=prod:NoSchedule taint.
Other forms of tolerations include:
operator: Exists: Matches any taint with the specified key, regardless of value.- Omitting
key: Matches any taint key. - Omitting
effect: Matches any effect for the specified key. operator: Existswithoutkeyoreffect: Tolerates all taints.
The tolerationSeconds field can be used with the NoExecute effect to specify how long a Pod should remain on a node before being evicted if the taint is added after the Pod has started.
API Resources: PersistentVolume (PV) and PersistentVolumeClaim (PVC)
Persistent Volumes (PVs) and Persistent Volume Claims (PVCs) are fundamental Kubernetes objects for managing persistent storage. They abstract the underlying storage implementation.
- PersistentVolume (PV): A storage resource in the cluster, such as an NFS share or a cloud provider's block storage. PVs are cluster resources, independent of any single Pod.
- PersistentVolumeClaim (PVC): A request for storage by a user. It specifies the desired size, access mode, and storage class. PVCs consume PV resources.
- StorageClass (SC): Defines different "classes" of storage. It allows for dynamic provisioning of PVs. When a PVC requests a StorageClass, Kubernetes can automatically provision a PV that matches the PVC's requirements.
Static Provisioning (PV and PVC)
In static provisioning, administrators manually create PVs, and users create PVCs that bind to these pre-existing PVs.
PV Example:
apiVersion: v1
kind: PersistentVolume
metadata:
name: manual-pv
spec:
storageClassName: manual-storage
capacity:
storage: 1Gi
accessModes:
- ReadWriteOnce # RWO: Can be mounted as read-write by a single node.
- ReadOnlyMany # ROX: Can be mounted read-only by many nodes.
- ReadWriteMany # RWX: Can be mounted read-write by many nodes.
hostPath: # Example using a local directory on a node.
path: /mnt/data/pv001
PVC Example:
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: manual-pvc
spec:
storageClassName: manual-storage
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 500Mi
Kubernetes matches PVCs to PVs based on:
- StorageClass: Must match.
- Access Modes: Must be compatible.
- Capacity: The PV's capacity must be greater than or equal to the PVC's requested capacity.
- Best Fit: If multiple PVs match, the one with the closest capacity to the request is chosen.
Dynamic Provisioning (StorageClass)
Dynamic provisioning uses StorageClasses to automatically create PVs when a PVC is created.
StorageClass Example:
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: dynamic-storage
provisioner: kubernetes.io/no-provisioner # Replace with actual provisioner (e.g., nfs.csi.k8s.io)
parameters:
# Parameters specific to the provisioner
server: nfs-server.example.com
path: /exports
A PVC using this StorageClass will trigger the provisioner to create a new PV.
Local Persistent Storage
Local persistent storage utilizes the local disk of a node for storage. This offers high performence and low latency but comes with node-specific limitations.
Key Considerations:
- Node Affinity: Pods using local PVs are typically tied to the node where the PV resides.
nodeAffinityis crucial to ensure the Pod is scheduled on the correct node. - Data Durability: If the node fails, the data on its local storage might become inaccessible. Implement robust backup strategies.
- Scheduling Constraints: Ensure the node has sufficient resources for the Pod.
- Reclaim Policy: Recommended policies are
Retain(keeps data after PVC deletion) orDelete(removes data).Recycleis not suitable for local PVs.
Example: Local PV with Node Affinity
First, ensure a directory exists on the target node, e.g., /mnt/local-storage.
Local PV Definition:
apiVersion: v1
kind: PersistentVolume
metadata:
name: local-pv-node1
spec:
storageClassName: local-storage
capacity:
storage: 5Gi
accessModes:
- ReadWriteOnce
persistentVolumeReclaimPolicy: Retain
local:
path: /mnt/local-storage
nodeAffinity:
required:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/hostname
operator: In
values:
- node-1 # Ensures this PV is only available on node-1
PVC Definition:
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: local-pvc-node1
spec:
storageClassName: local-storage
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 5Gi
Pod Definition using Local PVC:
apiVersion: v1
kind: Pod
metadata:
name: local-app-pod
spec:
containers:
- name: app-container
image: nginx:latest
volumeMounts:
- name: local-storage-volume
mountPath: /data
volumes:
- name: local-storage-volume
persistentVolumeClaim:
claimName: local-pvc-node1
With these configurations, the local-app-pod will be scheduled on node-1 and will mount the local storage at /data. Data written to /data will persist on node-1's /mnt/local-storage directory.