Customize alerting

View as Markdown

Once you have configured a receiver, you can change which alerts reach it, send some alerts to other receivers, tune the bundled rules, and silence alerts during maintenance.

The examples on this page set the monitoring module’s Terraform inputs, and require v15.0.0 or later of the Materialize Terraform Modules. With Helm, the same settings are chart values under alerting and rules, usually the same name in camel case, such as alerting.timeIntervals for time_intervals. Configuring Alerting through Terraform ⧉ lists the chart value for each.

Commands on this page use <alertmanager-namespace> for the namespace Alertmanager runs in: the release namespace, which the monitoring module sets from its namespace (monitoring in the examples), or alertmanager under the chart’s split-namespace profile.

How routing works

Every bundled rule sets a severity label: critical, warning, or notice. The selected preset maps each severity to a class, and an alert reaches every receiver whose class includes it. The shipped presets express how much the deployment depends on Materialize:

severity critical-infrastructure important (default) evaluation
critical page high normal
warning high normal normal
notice normal low suppressed

suppressed notifies nobody. A suppressed alert still fires and still shows in Alertmanager and in Grafana.

Every bundled rule also carries an audience label, which you can route on: platform for the Materialize deployment, its system clusters, and the Kubernetes platform under it, or workload for what runs on it, such as a user cluster falling behind or running out of memory.

Page on critical alerts

The critical-infrastructure preset maps critical to page. This example pages through PagerDuty and sends everything else to Slack:

module "monitoring" {
  # ...

  alerting = {
    preset = "critical-infrastructure"

    receivers = {
      oncall = {
        class = "page"
        route = { group_wait = "10s", repeat_interval = "1h" }
        config = {
          pagerduty_configs = [{
            routing_key_file = "/etc/alertmanager/secrets/alertmanager-receivers/pagerduty-key"
            send_resolved    = true
          }]
        }
      }
      platform = {
        class = ["high", "normal"]
        config = {
          slack_configs = [{
            channel       = "#platform-alerts"
            api_url_file  = "/etc/alertmanager/secrets/alertmanager-receivers/slack-url"
            send_resolved = true
          }]
        }
      }
    }
  }

  alerting_receiver_secrets = {
    "pagerduty-key" = var.pagerduty_routing_key
    "slack-url"     = var.platform_slack_webhook
  }
}

Under this preset warning maps to high and notice to normal, so the platform receiver gets both. The route options on oncall apply only to alerts routed to it, so pages are grouped and repeated faster than anything else.

Define your own preset

To map severities to classes of your own, add an entry under presets and select it with preset:

module "monitoring" {
  # ...

  alerting = {
    preset = "oncall-lite"
    presets = {
      oncall-lite = {
        critical = "page"
        warning  = "ticket"
        notice   = "suppressed"
      }
    }

    receivers = {
      oncall = {
        class = "page"
        config = {
          opsgenie_configs = [{
            api_key_file = "/etc/alertmanager/secrets/alertmanager-receivers/opsgenie-key"
          }]
        }
      }
      tickets = {
        class = "ticket"
        config = {
          webhook_configs = [{
            url = "https://tickets.example.com/hooks/alertmanager"
            http_config = {
              authorization = {
                credentials_file = "/etc/alertmanager/secrets/alertmanager-receivers/tickets-token"
              }
            }
          }]
        }
      }
    }
  }

  alerting_receiver_secrets = {
    "opsgenie-key"  = var.opsgenie_api_key
    "tickets-token" = var.tickets_token
  }
}

A cell set under presets for a shipped preset name, such as important, overrides only that cell and keeps the rest. An alert whose severity label is missing or unknown is routed as unknown_severity, which defaults to warning.

Route alerts to the team that owns them

routes.extra takes routes in Alertmanager’s own format, and places them ahead of the preset’s severity routes, so a specific match wins and the preset remains the fallback. This example sends every workload alert to the team that owns the clusters:

module "monitoring" {
  # ...

  alerting = {
    receivers = {
      chat = {
        class = ["high", "normal", "low"]
        config = {
          slack_configs = [{
            channel      = "#materialize-alerts"
            api_url_file = "/etc/alertmanager/secrets/alertmanager-receivers/slack-url"
          }]
        }
      }
      data-team = {
        config = {
          slack_configs = [{
            channel      = "#data-platform"
            api_url_file = "/etc/alertmanager/secrets/alertmanager-receivers/data-team-slack-url"
          }]
        }
      }
    }

    routes = {
      extra = [{
        matchers = ["audience=workload"]
        receiver = "data-team"
        continue = true
      }]
    }
  }

  alerting_receiver_secrets = {
    "slack-url"           = var.slack_webhook_url
    "data-team-slack-url" = var.data_team_slack_webhook
  }
}

data-team has no class, so only this route reaches it. Because the route sets continue = true, a workload alert also continues through the preset to chat. Without it, the alert would go to data-team only. Each receiver an extra route names must be defined under receivers, or the apply fails.

Tune the bundled rules

alert_rules sets which bundled rules install, and adjusts them without changing their expressions:

module "monitoring" {
  # ...

  alert_rules = {
    # User-cluster freshness is opt-in, because some clusters are behind by design.
    selected = ["cluster-falling-behind", "cluster-stale"]
    disabled = ["pods-stuck-in-waiting"]

    overrides = {
      # Large clusters here take hours to hydrate with nothing wrong.
      cluster-hydration-stuck   = { for_duration = "6h" }
      cluster-replica-oomkilled = { labels = { severity = "notice" } }
    }

    excluded_namespaces = ["materialize-scratch"]
  }
}
Attribute Purpose
selected Rules to install beyond the default set: alert names, rule-group names, or "*" for every rule that applies. Evaluate a rule outside the default set against your deployment before relying on it.
disabled Alert names never to install.
overrides Per alert, for_duration (the rule’s for) and labels. An override’s severity must stay critical, warning, or notice, and its audience platform or workload, so the routes still match it.
excluded_namespaces Namespaces no rule alerts on.

For the remaining attributes, including the namespaces your Materialize environments run in and the infrastructure workload tiers, see Tuning the bundled rules ⧉. For every bundled rule and whether it is in the default set, see Common Alerts ⧉.

NOTE: The rule evaluator imports every PrometheusRule in the cluster, not only the bundled ones. If another chart in the cluster ships its own PrometheusRule resources, those alerts are also evaluated and notified through this Alertmanager.

Silence alerts

A Materialize upgrade restarts pods and rehydrates clusters, which can fire alerts that resolve on their own. A silence stops the notifications for alerts matching a set of labels until it expires. The alerts still fire and still show in Alertmanager and Grafana.

In Grafana, go to Alerting > Silences, select the Alertmanager data source, and create a silence.
kubectl --namespace <alertmanager-namespace> exec alertmanager-0 -c alertmanager -- \
  amtool silence add namespace=materialize-environment \
    --duration=2h --author="$USER" --comment="Planned Materialize upgrade"

The command has two namespaces in it. --namespace is where Alertmanager runs, and namespace=materialize-environment matches the alerts to silence, here those about the Materialize environment’s namespace.

amtool silence query lists active silences, and amtool silence expire <id> ends one early.

Silences are shared between the Alertmanager replicas and survive the loss of either one. Give each a duration that covers the work and no more, so it does not hide the next incident on the same labels.

For a recurring schedule, define it under time_intervals and reference it from a receiver’s route.mute_time_intervals:

module "monitoring" {
  # ...

  alerting = {
    time_intervals = [{
      name = "change-window"
      time_intervals = [{
        weekdays = ["saturday"]
        times    = [{ start_time = "02:00", end_time = "06:00" }]
        location = "America/New_York"
      }]
    }]

    receivers = {
      chat = {
        class = ["high", "normal", "low"]
        route = { mute_time_intervals = ["change-window"] }
        config = {
          slack_configs = [{
            channel      = "#materialize-alerts"
            api_url_file = "/etc/alertmanager/secrets/alertmanager-receivers/slack-url"
          }]
        }
      }
    }
  }

  alerting_receiver_secrets = {
    "slack-url" = var.slack_webhook_url
  }
}

A mute window mutes everything routed to that receiver during the window, including an unrelated incident. For inhibition rules and the other options, see Maintenance Windows ⧉.

Manage receiver credentials

When alerting_receiver_secrets is set, the module creates the alertmanager-receivers Secret from it:

  • The plan fails when a receiver reads a key the map does not set, and the error names the key.

  • The values stay out of plan output, because the input is sensitive, but they are stored in Terraform state. Restrict who can read your state accordingly.

  • Rotating a credential needs no restart. Alertmanager reads the file on every send, and the kubelet refreshes the mounted Secret within about a minute.

Amazon SNS is the only integration that authenticates with the pod’s own cloud identity rather than a credential. Every other integration reads a credential from the Secret, including email sent through Amazon SES or Azure Communication Services. For the setup of each, see Cloud provider services ⧉.

Manage the Secret outside Terraform

To keep receiver credentials out of Terraform state, leave alerting_receiver_secrets empty. The module then creates no Secret, so External Secrets Operator, Vault Agent, or a cloud secret store’s CSI driver can own alertmanager-receivers instead. Create it in the namespace Alertmanager runs in. Alertmanager mounts it as optional, so its pods start before the Secret exists.

WARNING! With no alerting_receiver_secrets, the plan no longer checks that the Secret holds every key your receivers read. A missing key fails only when Alertmanager next tries to send through that receiver.

Run Alertmanager in its own namespace

The chart’s split-namespace profile runs Alertmanager in a namespace of its own, alertmanager. When you apply that profile through additional_values, also set alertmanager_namespace = "alertmanager" so the Secret is created where Alertmanager can mount it. The module cannot infer this from the profile. Commands that reach Alertmanager, such as amtool, then run in the alertmanager namespace.

Troubleshooting

Most configuration mistakes fail before anything is installed. The rest fail the Helm render during the apply, or are rejected by Alertmanager itself:

Mistake Fails at
A receiver reads a key that alerting_receiver_secrets does not set Plan, naming the key. Only when alerting_receiver_secrets is set
An override’s for_duration is not a duration Plan
preset is not a shipped preset or a key of presets Plan
A template name does not end in .tmpl Plan
An unknown capability, alert, or rule-group name Apply
An inline credential in a receiver or in global Apply
Once any receiver is configured, the preset maps a severity to a class no receiver serves Apply
An extra route names an undefined receiver or time interval Apply
An invalid field inside a receiver’s config Alertmanager reload. Alertmanager keeps the previous configuration, and alertmanager_config_last_reload_successful drops to 0

terraform plan shows the composed Helm values in the monitoring module’s helm_release resource, with alerting_receiver_secrets shown as (sensitive value). To check where an alert is routed once applied, see Confirm where alerts go.

See also

Back to top ↑