System status

Webhook Delivery degraded — all other systems operational.

Past incidents

  1. · minor · 74 minutes · Webhook Delivery

    Elevated webhook delivery latency

    A misconfigured connection pool limit in one delivery worker group caused webhook deliveries to queue for up to 11 minutes between 14:06 and 15:20 UTC. No deliveries were dropped and all queued webhooks were sent in order once the limit was corrected. We have added an alert on delivery queue age at the 2-minute mark.

  2. · minor · 38 minutes · Dashboard

    Dashboard unavailable for some sessions

    An expired internal certificate caused the dashboard to return errors for roughly one in five sessions. The API and scheduler were unaffected and all jobs ran normally. The certificate was rotated and we now track expiry dates for all internal certificates with a 30-day warning.

  3. · major · 112 minutes · Scheduler, Job Execution

    Delayed cron fires in ap-southeast

    A leader election bug in the ap-southeast scheduler caused cron fires to be delayed by up to 6 minutes between 03:40 and 05:32 UTC. Affected schedules fired late but exactly once. We fixed the lease renewal logic and added a check that simulates clock drift between scheduler nodes.

  4. · minor · 163 minutes · Run Logs

    Run log search returning stale results

    An indexing backlog following a storage migration meant recent runs did not appear in log search for up to 40 minutes. Underlying run data and webhook results were never affected. We increased indexer capacity and now page on indexing lag over 5 minutes.

  5. · major · 21 minutes · REST API, Scheduler

    API errors during database failover

    A planned primary failover in us-east took longer than expected, and roughly 9% of API requests returned 503 responses for 21 minutes. Scheduled jobs continued to fire from cached state. We have revised the failover runbook and shortened the connection retry window in the API tier.

  6. · minor · 52 minutes · Webhook Delivery

    Signature header missing on retried webhooks

    For 52 minutes, webhooks retried after a first failed attempt were sent without the Tend-Signature header due to a regression in a release that morning. Affected customers were notified directly and endpoints that enforce signatures rejected those attempts, which were then retried correctly after the fix. A contract check on signature presence now runs before every release.

  7. · minor · 97 minutes · Job Execution

    Concurrency limits not enforced across regions

    A caching issue caused concurrency keys used from both us-east and eu-central to be counted per region rather than globally for 97 minutes. A small number of runs briefly exceeded configured limits. Counters are now stored in a single authoritative service and reconciled every 10 seconds.

  8. · minor · 29 minutes · REST API

    Increased 429 responses on Pro plan

    A deployment applied the Hobby rate limit table to a subset of Pro projects, returning 429 responses at 60 requests per minute instead of 600. The deployment was rolled back after 29 minutes and no jobs were lost. Rate limit configuration now ships with a validation step that compares limits to each project's plan.