Failure Alerting

Get an email — or a webhook — when a run fails.

A scheduled forecast that fails silently overnight is the worst failure mode of an operational pipeline. Rebase tells you about it two ways, and they are deliberately different in scope:

  • Email to the owner — the person who deployed the workflow gets a mail when it starts failing, and another when it recovers. On by default; nothing to configure.
  • Workspace webhook — every failed run is POSTed to a URL you register. Off by default; one delivery per failure, for feeding your own tooling.

The same webhook also carries dataset staleness events (dataset.stale / dataset.fresh) when enabled with --on-stale — see Data Quality for freshness SLAs, which catch pipelines that stop producing without ever failing.

Email to the Owner

Every workflow records who deployed it, and every run records who asked for it. When an unattended run fails — one a cron schedule or an upstream trigger started — Rebase mails that person:

Subject: Rebase: accounting-sync failed

accounting-sync failed.

Project:   accounting
Workspace: acme
Trigger:   schedule
Run:       8f4c2f9e-...
Started:   2026-08-20 20:46:03 UTC
Finished:  2026-08-20 20:47:11 UTC

What happened
  the run was stopped because it exceeded its 512 MiB memory limit after using ~531 MiB

What to try
  Give the function more memory — for example memory="1Gi" on the function
  definition — or reduce how much the run holds in memory at once.

Error
  MemoryError: ...

Logs
  rebase run logs 8f4c2f9e-...

"What happened" and "What to try" come from the same structured diagnosis the SDK renders and the webhook carries as failure_reason.

One mail per outage

A workflow on a ten-minute schedule that breaks produces 144 failed runs a day. It does not produce 144 mails:

EventMail
First failure"failed"
Every failure after it(silent)
Still failing 24h later"still failing (N runs)"
First success again"recovered"

What is never mailed

  • Runs you triggered yourself (rebase run, the SDK, an endpoint call). The error is already in front of you — mailing it tells you what you just watched happen.
  • Ephemeral runs. They have no deployed definition to be "currently broken".

Turning it off

rebase workspace notifications set --no-email    # per workspace
rebase workspace notifications show

Mail is sent from no-reply@rebase.energy and needs a RESEND_API_KEY on the deployment; without one, no mail is sent regardless of this setting.

Who gets it

The run's owner, resolved in this order:

  1. The person who deployed the workflow, for a scheduled or triggered run.
  2. The person who made the request, for anything else — the signed-in user, or the person who created the API key it used.
  3. Failing both, whoever created the workspace.

A workflow deployed before ownership was recorded has no owner until someone redeploys it; the redeploy adopts it. Ownership is never reassigned by a later redeploy — a colleague redeploying your workflow does not inherit your alerts.

The Workspace Webhook

Configure the Webhook

rebase workspace notifications set \
  --webhook-url https://hooks.example.com/rebase \
  --webhook-secret <secret> \
  --on-failure

rebase workspace notifications show

Slack and Microsoft Teams incoming-webhook URLs work directly. To stop notifications, run rebase workspace notifications set --no-on-failure, or --clear-webhook to remove the stored URL and secret.

Payload

The webhook receives a POST with Content-Type: application/json and an X-Rebase-Event: run.failed header:

{
  "event": "run.failed",
  "sent_at": "2026-07-11T03:01:02+00:00",
  "run": {
    "id": "8f4c2f9e-...",
    "status": "failed",
    "error": "Cloud Run execution failed: ...",
    "failure_reason": {
      "code": "out_of_memory",
      "message": "the run was stopped because it exceeded its 512 MiB memory limit after using ~531 MiB",
      "hint": "Give the function more memory — for example memory=\"1Gi\" on the function definition ...",
      "observed": { "memory_limit_mib": 512, "memory_used_mib": 531, "exit_code": 137 }
    },
    "target_type": "workflow",
    "target_id": "1391f0d3-...",
    "target_name": "accounting-sync",
    "project_id": "b7a3c1d4-...",
    "workspace_id": "acme",
    "execution_backend": "prefect_cloud_run_service",
    "trigger_source": "schedule",
    "created_by_profile_id": "5c1e77a2-...",
    "created_at": "2026-07-11T03:00:00+00:00",
    "started_at": "2026-07-11T03:00:05+00:00",
    "finished_at": "2026-07-11T03:01:00+00:00"
  }
}

trigger_source distinguishes runs fired by a cron schedule ("schedule"), by a trigger ("trigger"), and manual or API-triggered runs ("api").

failure_reason is Rebase's own diagnosis of why the container died — code is one of out_of_memory, nonzero_exit, container_lost, step_timeout, run_timeout, or function_exception (the function body raised) — and is null when the run reported its own error normally. Prefer it over parsing error.

Verify the Signature

When a --webhook-secret is configured, every delivery carries an X-Rebase-Signature header containing an HMAC-SHA256 of the raw request body:

X-Rebase-Signature: sha256=<hex digest>

Verify it before trusting the payload:

import hashlib
import hmac


def verify(body: bytes, signature_header: str, secret: str) -> bool:
    expected = "sha256=" + hmac.new(secret.encode(), body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, signature_header)

Delivery Semantics

  • One attempt per failed run, with a 5-second timeout. Delivery failures are logged server-side and never affect the run itself.
  • Notifications are workspace-wide. Besides run.failed, the webhook can carry dataset.stale and dataset.fresh events (enable with --on-stale); staleness events are edge-triggered — one per transition, not one per check.
  • Unlike owner email, the webhook fires on every failed run, including ones you triggered yourself. It is a feed, not a notification.

On this page