Docs / Guides

Retries and backoff

How Tend decides when a run has failed, how the exponential-backoff-with-jitter delay is computed, and how to tune attempts, base, and cap per job.

Last updated March 19, 2026

#What counts as a failure

A run attempt fails when Tend cannot get a successful answer from your target. Specifically, an attempt is considered failed if the connection cannot be established, the response is a 5xx status, the response is 408 or 429, or the attempt exceeds the maximum job runtime of 15 minutes. Any other 4xx status is treated as a permanent failure: retrying the same request against the same endpoint is unlikely to change the answer, so Tend stops immediately and marks the run failed.

Any 2xx response marks the attempt, and the run, as succeeded. Redirects are followed up to three hops. If your endpoint returns a Retry-After header on a 429 or 503 response, Tend uses it as a lower bound for the next delay, capped by the job's backoff cap.

Target responseResultRetried?
200-299Run succeededNo
408, 429Attempt failedYes
500-599Attempt failedYes
Other 4xxRun failed permanentlyNo
Connection error or timeoutAttempt failedYes

#The default retry policy

Unless you specify otherwise, every job is allowed 5 retry attempts after the initial attempt, using exponential backoff with full jitter. The delay before retry n is a random value between 0 and min(cap, base * 2^n), where base is 10 seconds and cap is 3600 seconds. The maximum you can request is 25 retry attempts.

Full jitter means the delay is drawn uniformly from the whole range rather than clustered near the ceiling. This is deliberate: if a downstream service fails at 14:00 and a thousand runs fail with it, their retries spread across the window instead of arriving together and knocking the service over a second time.

Retry nUpper bound (base 10s, cap 3600s)Actual delay
120 secondsRandom 0-20 s
240 secondsRandom 0-40 s
380 secondsRandom 0-80 s
5320 secondsRandom 0-320 s
103,600 seconds (capped from 10,240)Random 0-3,600 s

#Configuring retries per job

Attach a retry object when you create a job or a schedule. max_attempts sets the number of retries (0 to 25), base_seconds sets the base, and cap_seconds sets the ceiling. The cap can be any value from 1 to 86400 seconds. Setting max_attempts to 0 disables retries entirely, which is appropriate for jobs that are unsafe to repeat.

cURL
curl https://api.tendcomputer.com/v2/jobs \
  -H "Authorization: Bearer tnd_live_8f2c1a9d4e7b" \
  -H "Content-Type: application/json" \
  -d '{
    "target": {"url": "https://api.example-shop.dev/internal/invoices/send"},
    "payload": {"invoice_id": "inv_20482"},
    "retry": {
      "max_attempts": 12,
      "base_seconds": 30,
      "cap_seconds": 7200
    }
  }'

With the policy above, the upper bound for the delay before retry 1 is 60 seconds (30 * 2), before retry 4 it is 480 seconds, and from retry 8 onward it is pinned at 7200 seconds.

#Run states and attempt history

Every run moves through a small set of states: queued, running, retrying, succeeded, failed, and skipped. A run in retrying has at least one failed attempt and a scheduled next_attempt_at. Fetch a run to see every attempt, including the response status, latency, and a truncated copy of the response body.

JSON
{
  "id": "run_01J9M4D2R7QK",
  "job_id": "job_01J9M4C8P1HD",
  "status": "retrying",
  "attempts": [
    {"n": 0, "started_at": "2026-03-19T14:00:02Z", "http_status": 503, "duration_ms": 412},
    {"n": 1, "started_at": "2026-03-19T14:00:31Z", "http_status": 503, "duration_ms": 388}
  ],
  "next_attempt_at": "2026-03-19T14:01:24Z"
}

Run logs are retained for 3 days on Hobby, 30 days on Pro, and 90 days on Scale. Attempt history disappears with the run log, so export anything you need for long-term auditing before it ages out.

#Designing handlers that survive retries

Retries give you at-least-once delivery, which means your endpoint may occasionally receive the same job more than once, for instance when it completes the work but the response is lost on the way back. Design handlers so that repeating a run is harmless.

  • Return a 2xx as soon as the work is durably accepted, then finish slow work asynchronously. Attempts that run longer than 15 minutes are cut off.
  • Return a 4xx only when the request is genuinely unprocessable; it ends retries immediately.
  • Use the run ID (sent in the Tend-Run-Id request header) as a deduplication key in your own database.
  • Honor Retry-After when you are overloaded. Tend treats it as a floor for the next delay.

#Retrying a failed run manually

When a run ends in failed, you can start a new attempt sequence with POST /v2/runs/{id}/retry. This creates a fresh run linked to the original through retried_from, with a full new retry budget. It is the right tool after you have fixed the underlying bug in your handler.

cURL
curl -X POST https://api.tendcomputer.com/v2/runs/run_01J9M4D2R7QK/retry \
  -H "Authorization: Bearer tnd_live_8f2c1a9d4e7b"