GET _cluster/health
{
"status" : "yellow",
"unassigned_shards" : 12,
...
}
The cluster has been yellow for a week, or it just went red and searches are failing. Both colors mean the same thing: some shards are not assigned to any node. The difference is which shards, and Elasticsearch will tell you exactly why.
What the colors mean
- Green: every primary and replica shard is assigned.
- Yellow: every primary is assigned, so all data is searchable, but at least one replica isn't. You've lost redundancy, not data. Another node failure could turn it red.
- Red: at least one primary is unassigned. Searches on that index return partial results or fail, and writes to those shards fail.
Step 1: find the unassigned shards
GET _cat/shards?v&h=index,shard,prirep,state,unassigned.reason&s=state
prirep is p for primary and r for replica. unassigned.reason gives a first hint: NODE_LEFT, ALLOCATION_FAILED, INDEX_CREATED, CLUSTER_RECOVERED, and so on.
Step 2: ask the cluster why
GET _cluster/allocation/explain
{
"index": "logs-2026.09.26",
"shard": 0,
"primary": false
}
With no body, it explains the first unassigned shard it finds. The response lists every node and the decider that said no, for example:
"deciders": [{
"decider": "disk_threshold",
"decision": "NO",
"explanation": "the node is above the high watermark cluster setting [cluster.routing.allocation.disk.watermark.high=90%] ..."
}]
Always start here. It's much faster than guessing.
The common causes
1. A single-node cluster with replicas. Elasticsearch never puts a replica on the same node as its primary. With one node and number_of_replicas: 1 (the default for many indices), the cluster is yellow forever. Either add a node, or accept no replicas on a dev box:
PUT logs-*/_settings
{ "index": { "number_of_replicas": 0 } }
More generally, replicas + 1 must be ≤ the number of eligible data nodes.
2. Disk watermarks. By default Elasticsearch stops allocating shards to a node above 85% disk use (low watermark), moves shards away above 90% (high), and marks indices read-only above 95% (flood stage). Free space by deleting old indices (ILM should do this for you), adding disk, or adding nodes. The flood-stage read-only block is released automatically once usage falls below the high watermark.
3. A node left. After a restart or crash, replicas wait index.unassigned.node_left.delayed_timeout (1 minute by default) for the node to come back before being rebuilt elsewhere. For planned restarts, pause replica allocation first so shards don't shuffle needlessly:
PUT _cluster/settings
{ "persistent": { "cluster.routing.allocation.enable": "primaries" } }
Then set it back to null (all) afterwards. Forgetting this is itself a common cause of permanent yellow.
4. Allocation filtering or tiers. Settings like index.routing.allocation.require._name, include._tier_preference, or node attributes can exclude every eligible node. The explain API names the filter.
5. Allocation failed too many times. After 5 failed attempts (for example a corrupted file or a transient out-of-memory), Elasticsearch stops retrying. Fix the cause, then:
POST _cluster/reroute?retry_failed=true
When it's red
A red status means a primary has no copy anywhere Elasticsearch trusts.
-
If the node holding it is coming back, bring it back. The primary recovers on its own.
-
If the data is gone, restore the index from a snapshot:
POST _snapshot/my_repo/nightly-2026.09.26/_restore { "indices": "orders-v3" } -
As a last resort, the reroute API's
allocate_stale_primaryorallocate_empty_primarycommands force a copy to become primary. Both accept data loss (stale data or an empty shard). Only use them when you understand what will be lost and have no snapshot.
Prevention
- Run at least 3 master-eligible nodes and enough data nodes for your replica count.
- Use ILM to roll over and delete indices before disks fill.
- Alert on disk use at 75–80%, below the low watermark.
- Take regular snapshots: they're the only real fix for a lost primary.
- Keep shard counts reasonable. Thousands of tiny shards slow recovery and allocation.
Checklist
- Yellow means missing replicas; red means missing primaries.
- Run
_cat/shardsthen_cluster/allocation/explain, and read the decider. - The usual fixes: replicas vs. node count, disk watermarks, allocation settings,
retry_failed. - For red: bring the node back or restore from a snapshot. Forced allocation loses data.
Get the weekly commit
New database deep dives every week.
