Kubernetes Scheduling and Controllers
Kubernetes is often introduced as a platform that “runs containers,” but that description misses the control system that makes the platform useful. A workload object declares desired state, controllers continually compare desired state with observed state, and the scheduler decides where newly created Pods can run. Those responsibilities are related, but they are not the same. Understanding their separation is central to the current KCNA Kubernetes Fundamentals domain.
This topic is intentionally narrower than Kubernetes architecture. The KCNA includes scheduling inside a Fundamentals domain that accounts for a large share of the blueprint, while broader architecture and container concepts are covered elsewhere in the production plan. The useful skill here is following a workload from declaration to placement and then understanding which control loop reacts when reality changes.
That perspective explains many everyday Kubernetes behaviors. A Deployment can desire three replicas while no Pod is currently running. The controller responds by creating what is missing. The scheduler then selects eligible nodes for pending Pods. Kubelets on those nodes attempt to run the assigned Pods. If a Pod disappears later, the controller may create a replacement and the scheduler makes another placement decision.
A controller watches part of the Kubernetes API and works to move the cluster toward the state declared by users or other controllers. That pattern is reconciliation: observe, compare, act, and repeat. It is different from a one-time script that assumes its command succeeded forever.
A Deployment controller, for example, manages replica behavior through ReplicaSets. If the desired number of replicas is three and only two eligible Pods exist, the controller creates work that moves the system back toward three. If the workload specification changes, the controller manages the transition according to the rollout strategy.
This is why declarative systems are resilient to individual object loss. The controller is not preserving one particular Pod; it is preserving the declared workload state. A replacement Pod can have a different identity and run on a different node while still satisfying the same desired state.
The scheduler focuses on pending Pods that do not yet have a node assignment. It evaluates available nodes against workload requirements and cluster constraints, filters out ineligible choices, and selects an appropriate placement according to its scheduling logic.
This separation is easy to confuse. If a Deployment is missing replicas, the scheduler does not decide to create more Pods. The controller creates the missing workload objects. The scheduler becomes involved only when those Pods need placement.
Troubleshooting improves immediately when this boundary is understood. “No new Pod exists” points toward the controller or workload definition. “Pod exists but remains Pending” points toward scheduling, resource, storage, affinity, taint, or other placement constraints.
CPU and memory requests tell Kubernetes what amount of resource a Pod needs for scheduling purposes. The scheduler uses those requests when deciding whether a node has enough allocatable capacity. Limits influence runtime resource behavior but should not be confused with the placement calculation itself.
A cluster can therefore have plenty of unused physical memory and still leave a Pod Pending if the scheduler’s view of requested capacity says no eligible node can satisfy the new request. Conversely, workloads with unrealistically low requests can be packed onto nodes in a way that creates runtime contention.
Requests are both technical and economic signals. They influence density, fragmentation, and the ability to place future work. KCNA candidates should understand why accurate requests improve cluster behavior even without becoming experts in capacity modeling.
Some workloads need more than raw capacity. Pod affinity can encourage or require workloads to run near other Pods. Anti-affinity can separate replicas so one failure does not affect all of them. Node affinity can steer workloads toward nodes with specific labels or characteristics.
These rules make placement align with application needs, but strong constraints can reduce scheduling flexibility. A required rule may leave a Pod Pending even when other nodes have free resources. A preference gives the scheduler more freedom but may not deliver the desired topology under pressure.
Good design uses hard requirements only when the workload truly cannot operate without them. Everything else should be treated as a preference where practical. That keeps the cluster resilient when capacity or topology changes.
Taints let nodes repel Pods that do not have matching tolerations. This is useful when nodes are dedicated to special hardware, sensitive workloads, infrastructure services, or another category that should not accept arbitrary workloads.
A toleration does not guarantee placement on the tainted node; it merely permits the Pod to be considered despite the taint. Other scheduling rules and resource requirements still apply. Candidates often confuse toleration with attraction, but it is better understood as permission to cross a placement barrier.
When a Pod remains Pending, taints should be checked alongside affinity, resource requests, node readiness, storage requirements, and other scheduling constraints. Placement is the combined result of these conditions.
A Deployment is designed around interchangeable replicas and rolling application updates. A StatefulSet gives Pods stable identities and ordered behavior useful for stateful applications. A DaemonSet tries to run a Pod on each applicable node, which suits node-level agents. Jobs and CronJobs represent finite or scheduled work rather than continuously available services.
These controllers are not merely different YAML shapes. They communicate different desired-state promises to Kubernetes. The controller’s reconciliation logic reflects those promises, which in turn changes what happens after failure, scale changes, or node additions.
Choosing the wrong controller can create operational friction even if the container itself runs. Candidates should ask whether replicas are interchangeable, whether stable identity matters, whether work must exist on every node, and whether the task should terminate.
Suppose a node fails and a Deployment Pod disappears. The controller observes that the desired replica count is no longer satisfied and creates a replacement Pod. That replacement begins Pending because it has no node. The scheduler evaluates eligible nodes and assigns one. The kubelet on that node attempts to start the containers.
This sequence illustrates why Kubernetes recovery is distributed across control components. The controller restores the desired count. The scheduler chooses placement. The node agent realizes the assigned Pod. If the cluster cannot satisfy the constraints, the process can stall at a specific stage that tells the administrator where to investigate.
This is more useful than thinking “Kubernetes restarts the app.” The platform continuously reconciles several kinds of state, and each component has a narrower responsibility.
A Pending Pod gives the operator a strong clue: the workload object exists, but placement or startup prerequisites are not satisfied. Events and Pod status can explain whether the scheduler found insufficient resources, an affinity conflict, a taint without toleration, a storage binding problem, or another eligibility issue.
The correct response is not automatically to add nodes. If the request is wrong, adding capacity treats the symptom. If a required affinity rule is impossible, more generic nodes may not help. If a taint is intentional, removing it may violate the cluster design.
Operational reasoning means reading the scheduler’s evidence, identifying the constraint, and deciding whether the workload or the cluster should change.
Controller health is visible through drift from desired state. Controllers are successful when observed state converges on declared state. Repeated failure to converge is a useful signal. A Deployment that continually creates Pods that immediately fail is not “self-healing” in a healthy sense; it is a control loop repeatedly encountering an unresolved problem.
Administrators should therefore monitor more than object counts. Rollout status, restart behavior, readiness, events, and application health explain whether reconciliation is producing a usable service. A controller can maintain the requested replica count while every replica is technically running but functionally unhealthy.
The KCNA certification level does not require deep scheduler implementation knowledge, but it does require understanding this control model well enough to interpret cluster behavior.
Storage can also participate in scheduling. A Pod that needs a persistent volume may be constrained by where that storage can attach or by how the StorageClass provisions it. This is why a Pending Pod can have sufficient CPU and memory available yet still be unschedulable. Placement is the intersection of compute resources, topology, policy, and workload dependencies.
Priority and preemption add another operational dimension. Higher-priority workloads can influence which Pods obtain scarce capacity, but priority should reflect service importance rather than become a routine substitute for capacity planning. A cluster in which every workload claims the highest priority has lost the information the scheduler needs to make meaningful trade-offs.
A useful study exercise starts with a Deployment specifying three replicas and resource requests. Follow what happens when it is submitted, when the ReplicaSet creates Pods, when the scheduler assigns nodes, and when kubelets start containers. Then change one condition at a time.
Add a taint, require an unavailable node label, increase a resource request, remove a node, or scale the Deployment. For each change, predict which component notices it first, whether a new Pod is created, whether that Pod can be scheduled, and what evidence an operator would see.
This exercise keeps scheduling and controller behavior connected without collapsing them into one vague “orchestration” concept. Linux Foundation build on this foundation: cloud-native operations become easier to understand once the candidate can trace how declared workload intent becomes concrete placement and continuous reconciliation.
