Docs / Tools
Monitoring, alerts and run logs
How to inspect runs, understand their lifecycle, search logs within your retention window, and get told when something breaks.
Last updated July 8, 2026
#The run lifecycle
A job is a request to do something. A run is one attempt to do it. A single job that fails twice and then succeeds has three runs, each with its own ID, timestamps and response details. Monitoring in Tend is mostly the study of runs.
| Status | Meaning | Terminal |
|---|---|---|
scheduled | Waiting for its run_at time or next cron tick. | No |
queued | Due to run; waiting for a concurrency slot. | No |
running | Request in flight to your endpoint. Maximum runtime is 15 minutes. | No |
retrying | The last attempt failed; another is scheduled after a backoff delay. | No |
succeeded | Your endpoint returned a 2xx response. | Yes |
failed | All retry attempts were exhausted, or a non-retryable error occurred. | Yes |
cancelled | Cancelled through the API or dashboard before completing. | Yes |
Retry delays use exponential backoff with full jitter: the delay before retry n is a random value between 0 and min(cap, base * 2^n), with a default base of 10 seconds and a default cap of 3,600 seconds. Jitter means two jobs that failed at the same moment will not retry at the same moment, which is kinder to whatever you are calling.
#Inspecting a run
Fetching a run returns the attempt number, the response status your endpoint sent, timing, the region that executed it, and a truncated excerpt of the response body. The excerpt is deliberately short: run logs record what happened, not what your service printed.
curl https://api.tendcomputer.com/v2/runs/run_01J9E7QH4T2XK8M5V3RNW6BZCD \
-H "Authorization: Bearer tnd_live_9fK2xQ7mVb3LpR8dWc1Zh"run = client.runs.retrieve("run_01J9E7QH4T2XK8M5V3RNW6BZCD")
print(run.status, run.attempt, run.response_status, run.duration_ms)const run = await client.runs.retrieve("run_01J9E7QH4T2XK8M5V3RNW6BZCD");
console.log(run.status, run.attempt, run.responseStatus, run.durationMs);{
"id": "run_01J9E7QH4T2XK8M5V3RNW6BZCD",
"job_id": "job_01J9E7QGX9M3P5V2K8RTB4HWNA",
"schedule_id": "sch_01J9C2VN5R8H3K7YQ4WTB6DXMA",
"status": "retrying",
"attempt": 3,
"max_attempts": 5,
"region": "us-west",
"started_at": "2026-07-08T08:30:02Z",
"duration_ms": 15004,
"response_status": 504,
"response_excerpt": "upstream request timeout",
"next_attempt_at": "2026-07-08T08:31:47Z"
}In the example above, the third attempt timed out at roughly 15 seconds and a fourth is scheduled about 105 seconds later. That delay falls between 0 and 80 seconds for attempt three under the defaults, plus queueing time, which is why looking at next_attempt_at beats doing the arithmetic yourself.
#Searching run logs
Run logs are searchable by job ID, schedule ID, status, region and time range, using the same filters and cursor pagination as every other list endpoint. How far back you can look depends on your plan.
| Plan | Run log retention | Requests per minute |
|---|---|---|
| Hobby | 3 days | 60 |
| Pro | 30 days | 600 |
| Scale | 90 days | 3,000 |
| Enterprise | Up to 365 days, optional export to your own storage | 10,000 |
# Everything that ultimately failed for one schedule in the past week
curl -G https://api.tendcomputer.com/v2/runs \
-H "Authorization: Bearer tnd_live_9fK2xQ7mVb3LpR8dWc1Zh" \
--data-urlencode "schedule_id=sch_01J9C2VN5R8H3K7YQ4WTB6DXMA" \
--data-urlencode "status=failed" \
--data-urlencode "created_after=2026-07-01T00:00:00Z" \
--data-urlencode "limit=200"failed = client.runs.list(
schedule_id="sch_01J9C2VN5R8H3K7YQ4WTB6DXMA",
status="failed",
created_after="2026-07-01T00:00:00Z",
limit=200,
)
for run in failed.auto_paging_iter():
print(run.id, run.response_status, run.response_excerpt)const failed = client.runs.list({
scheduleId: "sch_01J9C2VN5R8H3K7YQ4WTB6DXMA",
status: "failed",
createdAfter: "2026-07-01T00:00:00Z",
limit: 200,
});
for await (const run of failed) {
console.log(run.id, run.responseStatus, run.responseExcerpt);
}#Alerts
Alert rules watch a schedule, a tag or the whole project and notify you when a condition holds. Notifications go to an email address, a Slack incoming webhook or any HTTPS endpoint that accepts a signed POST. The most useful rules are the boring ones: a schedule failed, or a schedule did not run.
| Condition | Fires when | Typical use |
|---|---|---|
run_failed | A run reaches terminal status failed after exhausting retries. | Page someone for critical nightly jobs. |
consecutive_failures | N attempts in a row fail for the same schedule (N from 2 to 50). | Catch a flapping dependency before retries are exhausted. |
missed_run | No run started within a grace period after the expected tick. | Detect a paused schedule or a quota problem. |
duration_exceeded | A run takes longer than a threshold you set, up to 15 minutes. | Notice a job slowing down before it times out. |
queue_delay | Runs wait in queued longer than a threshold because concurrency slots are exhausted. | Signal that your concurrency limit needs raising. |
curl -X POST https://api.tendcomputer.com/v2/alert_rules \
-H "Authorization: Bearer tnd_live_9fK2xQ7mVb3LpR8dWc1Zh" \
-H "Content-Type: application/json" \
-d '{
"name": "reconcile-consecutive-failures",
"schedule_id": "sch_01J9C2VN5R8H3K7YQ4WTB6DXMA",
"condition": "consecutive_failures",
"threshold": 3,
"notify": [
{"type": "email", "to": "oncall@northfork.example"},
{"type": "slack", "url": "https://hooks.slack.com/services/T0000/B0000/XXXX"}
]
}'rule = client.alert_rules.create(
name="reconcile-consecutive-failures",
schedule_id="sch_01J9C2VN5R8H3K7YQ4WTB6DXMA",
condition="consecutive_failures",
threshold=3,
notify=[
{"type": "email", "to": "oncall@northfork.example"},
{"type": "slack", "url": "https://hooks.slack.com/services/T0000/B0000/XXXX"},
],
)const rule = await client.alertRules.create({
name: "reconcile-consecutive-failures",
scheduleId: "sch_01J9C2VN5R8H3K7YQ4WTB6DXMA",
condition: "consecutive_failures",
threshold: 3,
notify: [
{ type: "email", to: "oncall@northfork.example" },
{ type: "slack", url: "https://hooks.slack.com/services/T0000/B0000/XXXX" },
],
});#Monitoring webhook deliveries
When a job produces a result, Tend delivers it to your webhook endpoint with a Tend-Signature header. Each delivery is its own record, separate from the run that produced it, with its own status and attempt history. A run can succeed while its delivery is still failing, and the reverse is not possible.
- Your endpoint must answer within 15 seconds. Slower responses are counted as failures.
- Failed deliveries are retried for up to 48 hours with the same backoff scheme used for jobs.
- Inspect deliveries with
GET /v2/webhooks/deliveries, filterable bystatus,run_idand time range. - Redeliver a specific delivery on demand with
POST /v2/webhooks/deliveries/{id}/redeliveronce you have fixed your endpoint.
#Quotas, limits and troubleshooting
Two error responses account for most monitoring surprises. A 429 rate_limit_exceeded means your project exceeded its requests-per-minute limit; wait for the number of seconds in Retry-After. A 429 quota_exceeded applies only to Hobby projects and means the 10,000-run monthly cap has been reached; paid plans are billed for overage instead and never see it. If schedules stop firing on a Hobby project near the end of a month, check for this before anything else.
| Symptom | First thing to check |
|---|---|
Runs stuck in queued | Concurrency limit for your plan (5, 50, 500 or 2,000) and any per-schedule concurrency setting. |
| Runs fail with a 401 or 403 from your endpoint | The auth header or token your endpoint expects. Tend calls exactly the URL and headers you configured. |
region_unavailable when creating jobs | The region is temporarily not accepting jobs. Retry, or omit region to route to the default. |
| No alerts arriving | The notification target in the rule, and whether the condition is met. A job stuck retrying has not yet failed, so run_failed will not fire. |
| Old runs missing from search | Your plan's log retention window: 3, 30 or 90 days, up to 365 on Enterprise. |
When you contact support, include the run ID and the Tend-Request-Id header from the relevant API response. With those two values we can usually find your problem, and occasionally our own, in a single query.