Docs / Guides
Retries and backoff
How Tend decides when a run has failed, how the exponential-backoff-with-jitter delay is computed, and how to tune attempts, base, and cap per job.
Last updated March 19, 2026
#What counts as a failure
A run attempt fails when Tend cannot get a successful answer from your target. Specifically, an attempt is considered failed if the connection cannot be established, the response is a 5xx status, the response is 408 or 429, or the attempt exceeds the maximum job runtime of 15 minutes. Any other 4xx status is treated as a permanent failure: retrying the same request against the same endpoint is unlikely to change the answer, so Tend stops immediately and marks the run failed.
Any 2xx response marks the attempt, and the run, as succeeded. Redirects are followed up to three hops. If your endpoint returns a Retry-After header on a 429 or 503 response, Tend uses it as a lower bound for the next delay, capped by the job's backoff cap.
| Target response | Result | Retried? |
|---|---|---|
200-299 | Run succeeded | No |
408, 429 | Attempt failed | Yes |
500-599 | Attempt failed | Yes |
Other 4xx | Run failed permanently | No |
| Connection error or timeout | Attempt failed | Yes |
#The default retry policy
Unless you specify otherwise, every job is allowed 5 retry attempts after the initial attempt, using exponential backoff with full jitter. The delay before retry n is a random value between 0 and min(cap, base * 2^n), where base is 10 seconds and cap is 3600 seconds. The maximum you can request is 25 retry attempts.
Full jitter means the delay is drawn uniformly from the whole range rather than clustered near the ceiling. This is deliberate: if a downstream service fails at 14:00 and a thousand runs fail with it, their retries spread across the window instead of arriving together and knocking the service over a second time.
| Retry n | Upper bound (base 10s, cap 3600s) | Actual delay |
|---|---|---|
| 1 | 20 seconds | Random 0-20 s |
| 2 | 40 seconds | Random 0-40 s |
| 3 | 80 seconds | Random 0-80 s |
| 5 | 320 seconds | Random 0-320 s |
| 10 | 3,600 seconds (capped from 10,240) | Random 0-3,600 s |
#Configuring retries per job
Attach a retry object when you create a job or a schedule. max_attempts sets the number of retries (0 to 25), base_seconds sets the base, and cap_seconds sets the ceiling. The cap can be any value from 1 to 86400 seconds. Setting max_attempts to 0 disables retries entirely, which is appropriate for jobs that are unsafe to repeat.
curl https://api.tendcomputer.com/v2/jobs \
-H "Authorization: Bearer tnd_live_8f2c1a9d4e7b" \
-H "Content-Type: application/json" \
-d '{
"target": {"url": "https://api.example-shop.dev/internal/invoices/send"},
"payload": {"invoice_id": "inv_20482"},
"retry": {
"max_attempts": 12,
"base_seconds": 30,
"cap_seconds": 7200
}
}'job = client.jobs.create(
target={"url": "https://api.example-shop.dev/internal/invoices/send"},
payload={"invoice_id": "inv_20482"},
retry={"max_attempts": 12, "base_seconds": 30, "cap_seconds": 7200},
)
print(job.id) # job_01J9M4C8P1HDconst job = await client.jobs.create({
target: { url: "https://api.example-shop.dev/internal/invoices/send" },
payload: { invoice_id: "inv_20482" },
retry: { max_attempts: 12, base_seconds: 30, cap_seconds: 7200 },
});
console.log(job.id); // job_01J9M4C8P1HDjob, err := client.Jobs.Create(ctx, &tend.JobParams{
Target: tend.Target{URL: "https://api.example-shop.dev/internal/invoices/send"},
Payload: map[string]any{"invoice_id": "inv_20482"},
Retry: &tend.RetryPolicy{
MaxAttempts: 12,
BaseSeconds: 30,
CapSeconds: 7200,
},
})With the policy above, the upper bound for the delay before retry 1 is 60 seconds (30 * 2), before retry 4 it is 480 seconds, and from retry 8 onward it is pinned at 7200 seconds.
#Run states and attempt history
Every run moves through a small set of states: queued, running, retrying, succeeded, failed, and skipped. A run in retrying has at least one failed attempt and a scheduled next_attempt_at. Fetch a run to see every attempt, including the response status, latency, and a truncated copy of the response body.
{
"id": "run_01J9M4D2R7QK",
"job_id": "job_01J9M4C8P1HD",
"status": "retrying",
"attempts": [
{"n": 0, "started_at": "2026-03-19T14:00:02Z", "http_status": 503, "duration_ms": 412},
{"n": 1, "started_at": "2026-03-19T14:00:31Z", "http_status": 503, "duration_ms": 388}
],
"next_attempt_at": "2026-03-19T14:01:24Z"
}Run logs are retained for 3 days on Hobby, 30 days on Pro, and 90 days on Scale. Attempt history disappears with the run log, so export anything you need for long-term auditing before it ages out.
#Designing handlers that survive retries
Retries give you at-least-once delivery, which means your endpoint may occasionally receive the same job more than once, for instance when it completes the work but the response is lost on the way back. Design handlers so that repeating a run is harmless.
- Return a
2xxas soon as the work is durably accepted, then finish slow work asynchronously. Attempts that run longer than 15 minutes are cut off. - Return a
4xxonly when the request is genuinely unprocessable; it ends retries immediately. - Use the run ID (sent in the
Tend-Run-Idrequest header) as a deduplication key in your own database. - Honor
Retry-Afterwhen you are overloaded. Tend treats it as a floor for the next delay.
#Retrying a failed run manually
When a run ends in failed, you can start a new attempt sequence with POST /v2/runs/{id}/retry. This creates a fresh run linked to the original through retried_from, with a full new retry budget. It is the right tool after you have fixed the underlying bug in your handler.
curl -X POST https://api.tendcomputer.com/v2/runs/run_01J9M4D2R7QK/retry \
-H "Authorization: Bearer tnd_live_8f2c1a9d4e7b"new_run = client.runs.retry("run_01J9M4D2R7QK")
print(new_run.id, new_run.retried_from)const newRun = await client.runs.retry("run_01J9M4D2R7QK");
console.log(newRun.id, newRun.retried_from);