Skip to content
This is the documentation for main, which is not released yet. Read the latest.

Metrics

The Prometheus metrics the daemon serves at /metrics, once enabled with metrics.enable. Every metric of Rungar’s own is named rungar_. See Monitoring for scraping and alerts.

The endpoint also serves the Go runtime’s go_* metrics and the daemon process’s process_* metrics.

Daemon

MetricTypeLabelsDescription
rungar_build_infogaugeversion, commit, go_versionBuild information for the running daemon. Always 1. The build is in its labels.
rungar_start_time_secondsgaugeWhen the daemon started, in seconds since the Unix epoch. Uptime is time() - rungar_start_time_seconds.

GitHub

MetricTypeLabelsDescription
rungar_github_requests_totalcountercodeRequests made to GitHub, by the HTTP status it answered with. code is the status, such as 200, 401 or 503, or error when no answer came: a refused connection, a timeout, a TLS failure. Every attempt is counted, retries too. Each scale set’s message session polls GitHub, so a working daemon always has a 2xx rate above zero.

An expired or revoked credential shows as a rate of 401s, one without a permission as 403s or 404s, and an outage as 5xxs and errors – before, or instead of, the daemon restarting. See Monitoring for alerts on them.

Scale sets

MetricTypeLabelsDescription
rungar_runners_mingaugescale_setRunners the scale set keeps idle and ready: the min_runners in force. Its min_runners, or that of its schedule’s window in force, as of its latest scaling decision.
rungar_runners_desiredgaugescale_setRunners the scale set should have, as of GitHub’s latest statistics. Its assigned jobs plus the min_runners in force, at most max_runners; 0 while it is paused.
rungar_scale_set_pausedgaugescale_setWhether the scale set is paused: 1 if it is, 0 if it is taking jobs. Set by paused in the configuration, or by rungar scale-sets pause and resume until the daemon restarts. A paused scale set’s jobs wait on GitHub.
rungar_runners_currentgaugescale_set, stateRunners currently alive, by state. state is starting, idle or busy. Every state is present, at 0 if none.
rungar_runners_created_totalcounterscale_set, providerRunners created.
rungar_runners_removed_totalcounterscale_set, provider, reasonRunners removed, by scale set, provider and reason. reason is job_completed, when its job completed; scaled_down, when the scale set no longer needed it; never_connected, when it did not connect to GitHub in time; unregistered, when GitHub no longer had its registration; disconnected, when it was disconnected from GitHub for too long; stuck, when it was running a job while disconnected for too long; outdated, when it was created from another revision of its runner block; expired, when it was older than max_idle_age or max_age; requested, when it was removed with rungar runners rm or with its scale set; or orphaned, when GitHub had it disconnected with no machine for longer than start_timeout, and only its registration was removed, with provider empty. Runners created and removed as never_connected in a loop point to a broken image.
rungar_runners_lost_totalcounterscale_set, provider, reasonRunners lost, by scale set, provider and reason. A runner is lost when it ends without the daemon removing it or its job completing. reason is ended, when its machine ended without its job completing – it crashed, never connected, or was taken back, as a Spot machine can be – and its machine was deleted; or unreachable, when its provider had been unreachable for five minutes, and the scale set replaced it elsewhere. A rise in ended points to runners crashing or Spot machines being taken back.
rungar_scale_up_failures_totalcounterscale_set, reasonTimes a runner was wanted and could not be created, by reason. reason is no_capacity, when no provider took the runner: each was full, refused it or could not be tried. Or it is error, when GitHub refused to register the runner, creating it timed out or was cancelled, or a provider that refused it left a machine that could not be deleted. A provider that keeps refusing runners counts as no_capacity here; rungar_provider_calls_total{call="create",result="error"} shows which.
rungar_jobs_started_totalcounterscale_setWorkflow jobs started on this fleet’s runners.
rungar_jobs_completed_totalcounterscale_set, resultWorkflow jobs completed, by result. result is as GitHub gives it: succeeded, failed, canceled.
rungar_jobs_assignedgaugescale_setJobs GitHub has assigned the scale set and not completed, running or waiting. As of GitHub’s latest statistics. Those waiting for a runner are this less the busy runners, rungar_runners_current{state="busy"}.
rungar_job_wait_secondshistogramscale_setHow long each job waited, from being queued on GitHub to a runner taking it. What the job’s author waits. It includes GitHub’s routing of the job to the scale set, which rungar_job_runner_wait_seconds leaves out. Measured by GitHub’s clock. Buckets, in seconds: 5, 10, 15, 20, 30, 45, 60, 90, 120, 180, 300, 600, 900, 1800.
rungar_job_runner_wait_secondshistogramscale_setHow long each job waited for a runner, from GitHub assigning it to the scale set to a runner taking it. The part of the wait Rungar answers for: creating a runner and its booting, or none with one idle. The number a service level is set on. Measured by GitHub’s clock. Buckets, in seconds: 5, 10, 15, 20, 30, 45, 60, 90, 120, 180, 300, 600, 900, 1800.
rungar_runner_create_duration_secondshistogramscale_set, providerHow long creating each runner took, from registering it with GitHub to its provider having created its machine. provider is the one that created it, and the time includes the providers tried before it. Runners that could not be created are not counted. Buckets, in seconds: 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, 60, 120, 300.
rungar_runner_boot_duration_secondshistogramscale_set, providerHow long each runner took to connect to GitHub once its machine was created. The image’s and the network’s part of a job’s wait. A runner that takes a waiting job is seen connecting to within a second or two; one that does not, to within 20 seconds. Adopted runners, and runners that never connect, are not counted. Buckets, in seconds: 5, 10, 15, 20, 30, 45, 60, 90, 120, 180, 300, 600, 900, 1800.
rungar_scale_set_last_poll_timestamp_secondsgaugescale_setWhen GitHub last answered the scale set’s poll for jobs, in seconds since the Unix epoch. GitHub answers at least every minute or so, with jobs or without. One falling behind is a scale set not hearing of its jobs, while others may be.
rungar_scale_set_last_reconcile_timestamp_secondsgaugescale_setWhen the scale set last reconciled with the fleet, in seconds since the Unix epoch. Set when a reconciliation lists the providers, or those that are reachable. One falling behind reconcile_interval is reconciliation stuck, or failing: the log says why at debug.
rungar_reconcile_duration_secondshistogramscale_setHow long each reconciliation with the fleet took. Most of it is listing the providers and asking GitHub about the runners. Buckets, in seconds: 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, 60, 120, 300.

The runner gauges are updated on every reconciliation – every reconcile_interval, 30 seconds unless set – and the rest as things happen. The runners counted are those of the scale set Rungar knows of: its own, and those it has adopted.

Each histogram is served as its _bucket, _count and _sum series: rate(rungar_job_runner_wait_seconds_sum[1h]) / rate(rungar_job_runner_wait_seconds_count[1h]) is the mean wait for a runner over the last hour, and histogram_quantile(0.95, sum by (le) (rate(rungar_job_runner_wait_seconds_bucket[1h]))) the wait 95 jobs in 100 were within.

Providers

MetricTypeLabelsDescription
rungar_provider_reachablegaugeprovider1 if the provider’s last listing succeeded, or it has not been listed yet; 0 if not.
rungar_provider_runnersgaugescale_set, providerRunners currently alive on the provider, by scale set. A scale set with none of its runners on the provider is absent.
rungar_provider_calls_totalcounterprovider, call, resultCalls made to the provider, by call and result. call is list, create or delete. result is ok; no_capacity, when the provider said it was full; or error. Calls the daemon itself cancelled are not counted.
rungar_provider_call_duration_secondshistogramprovider, callHow long each call to the provider took, whatever its result. call is as for rungar_provider_calls_total. Buckets, in seconds: 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, 60, 120, 300.

rungar_provider_reachable is updated every reconcile_interval: a provider is reachable if its last listing succeeded. rungar_provider_runners is updated with the scale sets’ runner gauges, on every reconciliation. The calls are counted and timed as they are made.