> ## Documentation Index
> Fetch the complete documentation index at: https://agno-v2-docs-agentos-durable-background-execution.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Operations and Monitoring

> The /queue endpoints: dead-letter listing, requeue, queue stats, retention, admin gating, and deployment affinity.

The durable queue mounts an operator surface under `/queue`. With the default `max_attempts=1`, a crashed run stays failed until an operator acts, so these endpoints list failed jobs, grant them another attempt, and report queue depth.

```bash theme={null}
curl localhost:7777/queue/stats
```

```json theme={null}
{"counts": {"completed": 3, "cancelled": 1, "running": 1}, "oldest_queued_age_seconds": null}
```

## Endpoints

| Method | Path                           | Purpose                                                                                       |
| ------ | ------------------------------ | --------------------------------------------------------------------------------------------- |
| `GET`  | `/queue/stats`                 | Job counts by status and the oldest queued job's age in seconds.                              |
| `GET`  | `/queue/jobs`                  | Paginated job listing. `status` is repeatable. `status=failed` alone is the dead-letter list. |
| `GET`  | `/queue/jobs/{job_id}`         | One job. `job_id` equals the run's `run_id`.                                                  |
| `POST` | `/queue/jobs/{job_id}/requeue` | Grant a `failed` or `cancelled` job one more execution.                                       |

All four answer `503` when the durable queue is not enabled on the replica that served the request.

### Listing parameters

| Parameter    | Default      | Description                                                                                                                          |
| ------------ | ------------ | ------------------------------------------------------------------------------------------------------------------------------------ |
| `status`     | all          | One of `queued`, `running`, `completed`, `failed`, `cancelled`, `paused`. Repeat to match several: `status=failed&status=cancelled`. |
| `limit`      | `20`         | Jobs per page, 1 to 1000.                                                                                                            |
| `page`       | `1`          | Page number.                                                                                                                         |
| `sort_by`    | `created_at` | Also `updated_at`, `completed_at`, `status`, `attempt`, `component_type`, `component_id`, `user_id`. Unknown fields are ignored.     |
| `sort_order` | `desc`       | `asc` or `desc`.                                                                                                                     |

### Job fields

| Field                                      | Description                                                           |
| ------------------------------------------ | --------------------------------------------------------------------- |
| `id`                                       | The run's `run_id`.                                                   |
| `component_type`, `component_id`           | `agent`, `team`, or `workflow`, and which one.                        |
| `session_id`, `user_id`                    | Session and user the run was submitted for.                           |
| `status`                                   | `queued`, `running`, `completed`, `failed`, `cancelled`, or `paused`. |
| `attempt`, `max_attempts`                  | Executions started so far, and the budget.                            |
| `error`                                    | Terminal error of the last attempt.                                   |
| `locked_by`, `locked_at`                   | Worker holding the lease and when it was last refreshed.              |
| `idempotency_key`                          | Client-provided key, if any.                                          |
| `payload`                                  | Serialized run parameters. Contains verbatim user input.              |
| `created_at`, `updated_at`, `completed_at` | Epoch seconds.                                                        |

## Requeue

```bash theme={null}
curl -X POST localhost:7777/queue/jobs/{job_id}/requeue
```

Requeue raises `max_attempts` to `attempt + 1` and moves the job back to `queued`. Only `failed` and `cancelled` jobs qualify; anything else answers `400`. The replica that served the request claims the job immediately if it has capacity.

| Query parameter      | Default | When to set it                                                                                                                                                                                                                           |
| -------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `clear_cancellation` | `false` | Requeueing a job that was cancelled. A recorded cancel is never cleared automatically. Without this flag the re-driven attempt is cancelled again at its first checkpoint, visibly. Set `true` as the explicit operator override.        |
| `force`              | `false` | Requeueing a job that failed within the last `lock_grace_seconds`. Its worker may still be executing (a sweep proves heartbeats stopped, not that execution did). The request answers `409` until the grace elapses unless `force=true`. |

Requeue does not bypass attempt fencing. If the presumed-dead worker finishes after all, the later attempt owns the job and the earlier one's writes are discarded.

## Stats and alerting

| Signal                              | Meaning                                                                                                                                 |
| ----------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
| `counts.queued` trending up         | Submissions outpace execution. Add replicas or raise `max_concurrency`. Submissions get `429` at `max_queue_depth`.                     |
| `oldest_queued_age_seconds` growing | Something claimable is not being claimed. Usually a [deployment affinity](#deployment-affinity) mismatch, or every replica at capacity. |
| `counts.failed`                     | The dead-letter backlog.                                                                                                                |
| `counts.paused`                     | Runs waiting on a human. Each one holds its session's line.                                                                             |

## Retention

Once an hour, the worker deletes terminal jobs (`completed`, `failed`, `cancelled`) older than `retention_seconds` (default 24 hours). Paused jobs are exempt: a paused run's job is what a later continue re-queues, and it must outlive human latency. Cancel a paused run to release its job. Run rows and sessions are not touched by queue retention.

## Access control

Job rows expose payloads and user IDs across tenants, and requeue grants execution budget, so `/queue` is an operator surface. Any request carrying a JWT identity must hold the admin scope, or it gets `403`, regardless of whether user isolation is on. Deployments without JWT enforcement (security key or open) pass through, matching how the run routes treat scope enforcement. See [JWT middleware](/agent-os/middleware/jwt) and [Scopes](/agent-os/security/authorization/scopes).

## Deployment affinity

`deployment_id` pins jobs to a subset of workers in a heterogeneous fleet:

```python theme={null}
queue=QueueConfig(durable=True, deployment_id="gpu-pool")
```

| Replica setting                | Enqueues                             | Claims                                                                |
| ------------------------------ | ------------------------------------ | --------------------------------------------------------------------- |
| `deployment_id=None` (default) | Jobs with no `deployment_id`         | Jobs with no `deployment_id`                                          |
| `deployment_id="gpu-pool"`     | Jobs with `deployment_id="gpu-pool"` | Jobs with no `deployment_id` and jobs with `deployment_id="gpu-pool"` |

A mixed fleet needs no extra configuration. A continuation inherits the submitting job's `deployment_id`.

<Warning>
  Jobs with a `deployment_id` that no live worker matches wait forever. They are queued, not stale, so no sweep touches them and no error is raised. The symptom is `oldest_queued_age_seconds` climbing in `/queue/stats`. Sweeping is not affinity-filtered: any replica can fail a stale job, because that only records a failure that already happened.
</Warning>

## Fleet-wide settings

Every replica sharing one queue table must use the same values for `lock_grace_seconds`, `stop_timeout_seconds`, `retention_seconds`, `max_attempts`, and `timeout_seconds`. Each is applied by whichever replica performs the action. The replica that claims a job heartbeats on its own `lock_grace_seconds`, while the replica that sweeps judges staleness on its own value. `max_attempts` is written onto each job by the replica that accepted it. The smallest `retention_seconds` in the fleet wins the hourly cleanup. Divergent values, including transiently during a rolling deploy, can falsely sweep a healthy peer's runs or delete jobs early.

When changing `lock_grace_seconds` on a live fleet, only ever raise it, and roll the sweeping replicas first. A replica sweeping with a smaller grace than its peers heartbeat with judges their live leases stale.

## Shutdown and restart

On graceful shutdown the worker stops claiming and gives in-flight runs `stop_timeout_seconds` (default 30, always below `lock_grace_seconds`) to finish. Heartbeats continue through the drain, so peers do not reclaim a run that is still finishing. Stragglers are cancelled once: with attempts remaining they return to `queued`, otherwise they fail with `interrupted by worker shutdown`. Queued jobs are untouched and execute after the restart, on whichever replica claims them.

## Next Steps

| Task                                               | Guide                                                                                          |
| -------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| Choose `max_attempts` and understand the guarantee | [Durable queue](/agent-os/background-execution/durable-queue#the-guarantee)                    |
| Wire cancels and streams across replicas           | [Multi-replica deployments](/agent-os/background-execution/multi-replica)                      |
| Release a stuck paused run                         | [Cancel and retention](/agent-os/background-execution/hitl-continuations#cancel-and-retention) |
