Skip to content
This is the documentation for main, which is not released yet. Read the latest.

Monitoring

Rungar tells you what it is doing in three ways: Prometheus metrics, for dashboards and alerts; rungar status and the ls commands, for a look at the fleet now; and the daemon’s log, for what happened and why.

Metrics

The daemon serves Prometheus metrics once enabled in its configuration:

/etc/rungar/config.yaml
metrics:
  enable: true
  listen: 127.0.0.1:9102

A running daemon takes it up when restarted:

$ sudo systemctl restart rungar
$ curl -s 127.0.0.1:9102/metrics | grep '^rungar_runners_current'
rungar_runners_current{scale_set="rungar-c2-m4",state="busy"} 3
rungar_runners_current{scale_set="rungar-c2-m4",state="idle"} 1
rungar_runners_current{scale_set="rungar-c2-m4",state="starting"} 0

The endpoint, /metrics, has no authentication. It says which scale sets there are, how busy they are and which providers Rungar uses, but nothing about the jobs or the credentials. It listens on loopback unless told otherwise: serve it on an address only your Prometheus can reach, and scrape it:

prometheus.yml
scrape_configs:
  - job_name: rungar
    static_configs:
      - targets: ["rungar1.example.com:9102"]

Every metric is listed in the metrics reference.

What to watch

How long jobs wait

rungar_job_runner_wait_seconds is how long each job waited for a runner, from GitHub giving it to the scale set to a runner taking it: the part of the wait Rungar answers for, and the number to set a service level on. The share of jobs that got a runner within two minutes, over a day:

sum by (scale_set) (rate(rungar_job_runner_wait_seconds_bucket{le="120"}[1d]))
/
sum by (scale_set) (rate(rungar_job_runner_wait_seconds_count[1d]))

and the wait 95 jobs in 100 were within, over the last hour:

histogram_quantile(0.95, sum by (scale_set, le) (rate(rungar_job_runner_wait_seconds_bucket[1h])))

A service level is set on a bucket’s bound – 5, 10, 15, 20, 30, 45, 60, 90, 120, 180, 300 seconds and on – so that it is counted exactly. rungar_job_wait_seconds is the whole wait, since the job was queued, which adds GitHub’s own routing of it; both are measured by GitHub’s clock.

A wait is a runner being created and booting, which rungar_runner_create_duration_seconds and rungar_runner_boot_duration_seconds split by provider: the first is the provider’s, the second the image’s and the network’s. A job that finds a runner idle waits for neither. See Troubleshooting for a runner whose wait was long.

The fleet and the daemon

Whether the scale sets keep up. rungar_runners_desired is how many runners a scale set should have, and rungar_runners_current how many it has. Behind for a moment is normal – a runner takes a while to boot – but behind for long means its jobs are waiting:

rungar_runners_desired - on (scale_set) sum by (scale_set) (rungar_runners_current)

A paused scale set should have none, so its jobs waiting do not show here: rungar_scale_set_paused is 1 while it is.

rungar_scale_up_failures_total says why: no_capacity when no provider took the runner, and error when something else stopped it – GitHub refused to register it or timed out, or a provider’s refusal left a machine behind that could not be deleted – which is something to fix. no_capacity is the providers being full, which is the fleet being too small or a size too large for any provider, or the providers failing: every one that was tried refused, for whatever reason. Which it was is in rungar_provider_calls_total, below.

Whether the runners start. A runner stays starting until it connects to GitHub, and one that never does is removed after start_timeout and replaced, with a warning in the log saying so, and counted in rungar_runners_removed_total with reason="never_connected":

sum by (scale_set, provider) (increase(rungar_runners_removed_total{reason="never_connected"}[30m]))

Runner after runner of a scale set removed that way is an image without the runner in it, or a network that cannot reach GitHub; see Troubleshooting.

Whether each scale set is listening. A scale set hears of its jobs by polling GitHub, which answers at least every minute or so, jobs or none. rungar_scale_set_last_poll_timestamp_seconds is when it last did; one that falls behind is a scale set deaf to its jobs, which the daemon-wide GitHub metrics do not show while the others are fine:

time() - rungar_scale_set_last_poll_timestamp_seconds > 180

rungar_scale_set_last_reconcile_timestamp_seconds is the same for reconciliation, which should never be more than a reconcile_interval or two behind, and rungar_reconcile_duration_seconds how long it takes.

The providers. rungar_provider_reachable is 0 for one that did not answer; its runners are left alone for five minutes, and then replaced elsewhere. A provider that refuses a runner, full or failing, is skipped by the scale set that found it so, which tries its next provider and says so in the events. Every call to a provider is counted by result in rungar_provider_calls_total, refusals included: call="create" with result="no_capacity" is a provider full, and with result="error" one failing, which is something to fix. Calls are also timed in rungar_provider_call_duration_seconds: a provider slow to list slows every reconciliation, and one slow to create slows every runner created on it.

GitHub. rungar_github_requests_total counts every request to GitHub by the status it was answered with, or error when none came. A credential that expired or was revoked is a rate of 401s; one missing a permission, 403s or 404s; an outage, 5xxs and errors. Each scale set’s message session polls GitHub all the time, so a working daemon always has a rate of 2xxs, and one with none is not hearing from GitHub.

The daemon. rungar_build_info carries the version and commit running, for a dashboard to show or to join on, and rungar_start_time_seconds when it started: a value that keeps changing is a daemon restarting.

The jobs. rungar_jobs_completed_total by result is how the jobs on the fleet end. A rise in failed that is not the code’s own is often the runners’: out of disk, out of memory, or an image missing a tool.

Alerts worth having

rungar-alerts.yml
groups:
  - name: rungar
    rules:
      - alert: RungarDown
        expr: up{job="rungar"} == 0
        for: 5m

      - alert: RungarRestarting
        expr: changes(rungar_start_time_seconds[1h]) > 3
        annotations:
          summary: "Rungar has restarted more than 3 times in an hour: see the log"

      - alert: RungarGitHubRefusing
        expr: sum(rate(rungar_github_requests_total{code=~"401|403"}[5m])) > 0
        for: 5m
        annotations:
          summary: "GitHub is refusing Rungar's credential"

      - alert: RungarGitHubUnreachable
        expr: sum(rate(rungar_github_requests_total{code=~"2.."}[5m])) == 0
        for: 5m
        annotations:
          summary: "Rungar has had no successful answer from GitHub for 5 minutes"

      - alert: RungarProviderUnreachable
        expr: rungar_provider_reachable == 0
        for: 5m
        annotations:
          summary: "Provider {{ $labels.provider }} has been unreachable for 5 minutes"

      - alert: RungarScaleSetBehind
        expr: >
          rungar_runners_desired
          - on (scale_set) sum by (scale_set) (rungar_runners_current) > 0
        for: 15m
        annotations:
          summary: "{{ $labels.scale_set }} has had fewer runners than its jobs need for 15 minutes"

      - alert: RungarCannotScaleUp
        expr: increase(rungar_scale_up_failures_total{reason="error"}[15m]) > 0
        annotations:
          summary: "{{ $labels.scale_set }} could not create runners: see the log"

      - alert: RungarProviderCannotCreate
        expr: increase(rungar_provider_calls_total{call="create",result="error"}[15m]) > 0
        annotations:
          summary: "Provider {{ $labels.provider }} failed to create runners: see the events"

      - alert: RungarRunnersNeverConnect
        expr: >
          sum by (scale_set, provider)
          (increase(rungar_runners_removed_total{reason="never_connected"}[30m])) >= 3
        annotations:
          summary: "{{ $labels.scale_set }}'s runners on {{ $labels.provider }} keep failing to connect to GitHub"

      - alert: RungarJobsWaitTooLong
        expr: >
          histogram_quantile(0.95,
            sum by (scale_set, le) (rate(rungar_job_runner_wait_seconds_bucket[30m]))) > 300
        for: 15m
        annotations:
          summary: "{{ $labels.scale_set }}'s jobs wait more than 5 minutes for a runner, 1 in 20 of them"

      - alert: RungarScaleSetNotPolling
        expr: time() - rungar_scale_set_last_poll_timestamp_seconds > 300
        annotations:
          summary: "{{ $labels.scale_set }} has not heard from GitHub for 5 minutes: it is not getting its jobs"

      - alert: RungarReconcileStalled
        expr: time() - rungar_scale_set_last_reconcile_timestamp_seconds > 600
        annotations:
          summary: "{{ $labels.scale_set }} has not reconciled with the fleet for 10 minutes: see the log at debug"

      - alert: RungarProviderFailing
        expr: >
          sum by (provider) (rate(rungar_provider_calls_total{result="error"}[15m]))
          / sum by (provider) (rate(rungar_provider_calls_total[15m])) > 0.5
        for: 15m
        annotations:
          summary: "More than half the calls to provider {{ $labels.provider }} fail"

      - alert: RungarScaleSetPaused
        expr: rungar_scale_set_paused == 1
        for: 4h
        annotations:
          summary: "{{ $labels.scale_set }} has been paused for 4 hours: its jobs are waiting"

RungarJobsWaitTooLong and RungarScaleSetBehind are the ones that matter most: jobs waiting, whatever the cause. The others say which cause. Set the first’s five minutes to the service level you promise, on a bucket’s bound. RungarProviderCannotCreate catches a provider failing even while the scale sets place their runners on the others and no job waits. RungarScaleSetPaused is for the cause the first two cannot see, a pause forgotten; drop it for a scale set the configuration pauses on purpose.

Metrics count runners, but do not name them. To see which runners a scale set has, and what each is doing:

$ sudo -u rungar rungar runners ls --scale-set rungar-c2-m4

Beyond the metrics

For a look at the fleet now, rungar status and the ls and inspect commands name every scale set, provider and runner; see Draining, pausing and removing. For what happened and why, the daemon’s log and its events; see Running the daemon. A runner’s events are its timeline – created, connected, its job started, removed – with how long each step took:

$ sudo -u rungar rungar events --runner rungar-c2-m4-1ff1015a

Next steps