Skip to main content

docs Quotas and rate limits

Quotas & rate limits

Every organization is bound to a plan. The plan is the single source of truth for resource quotas and per-minute rate budgets — the number a request is enforced at is the same number the API reports back. A new tier is one row in the plans table; nothing in the enforcement path changes.

The seeded plans

Three plans ship seeded: free, founding (a more generous free tier granted to early accounts and kept for life), and pro. Only free is listed; the other two are assigned by signup or billing. On the hosted service free is sold as Standard, founding as Founding, and pro as Pro; see the pricing page.

QuotafreefoundingproMeaning
max_targets2050150Monitored targets in the org
min_check_interval_secs1806030Plan-side floor on a target's check interval. The effective floor is max(this, kind_min)kind_min is 43200 for domain_expiry, 3600 for tls_cert, 300 for flow, 60 for heartbeat, and 10 for http / tcp / dns / ping.
retention_days3090395History window the UI and API will read
raw_days303030Per-check detail retention, stamped onto each ClickHouse row at write time
evidence_days777How long a failed browser-flow run keeps the page it captured. Clamped to raw_days, since the run it explains goes then
max_flow_steps303030Steps one flow monitor may declare. Clamped to the engine ceiling of 30, so a larger value has no effect
max_flow_checks000Browser flow monitors the org can create; 0 doubles as the feature gate, so a create returns 403 FLOW_CHECKS_DISABLED rather than a quota error. Seeded at 0 because a flow also needs flow.enabled on the process that runs it; raise it on the plan row to switch flow on. The hosted service sets its own values, listed on Plans and limits
max_regions3Regions a single monitor can be assigned to
max_members3515Active members in the org
max_pending_invitations101525Outstanding (unaccepted) invitations
max_api_tokens_per_user5710API tokens a single user may hold
max_status_pages125Public status pages the org can run
max_public_components153075Distinct monitors published across all of the org's pages (a monitor on several pages counts once)
max_share_links_per_monitor135Live share links on one monitor
max_shared_monitors2510Monitors with at least one share link
max_maintenance_windows203050Scheduled maintenance windows
max_notification_channels203050Notification channels (Slack/webhook/Telegram/WhatsApp/SMS/…) in the org
max_escalation_policies101050Escalation policies
max_on_call_schedules5525On-call schedules
max_logo_size_bytes104857610485761048576Status-page logo upload ceiling (1 MiB)

Feature flags ride on the same row: custom_domain_enabled, white_label_enabled, sms_alerts_enabled, incident_narration_enabled, on_call_enabled. white_label_enabled is what makes the status-page "powered by" toggle real: on a plan without it the badge always renders, whatever the page setting says.

Rate budget (per minute)freefoundingproCategory
api_writes_per_minute6009001200POST/PATCH/DELETE on /api/v1/*
api_reads_per_minute6000900012000GET/HEAD/OPTIONS on /api/v1/*
bulk_ops_per_minute304560/api/v1/targets/bulk*
test_now_per_minute6090120POST /api/v1/targets/test + the notification-channel test endpoints
check_now_per_minute6090120POST /api/v1/targets/{id}/check-now

One category sits outside the plan: support (POST /api/v1/support, the in-app help form) is capped at a fixed 2 per minute on every tier. It spends the operator's mail budget rather than a tenant resource, so paying more does not buy a larger share of it.

How quotas are enforced

A resource quota is checked atomically at the write, not by a check-then-act in the handler. The friendly handler-side pre-check exists only to produce a clean error on the common, uncontended path; the race-safe guarantee is in the store:

  • Targets — the count bound is inside the INSERT (single and bulk), handed the same max_targets. Concurrent creates at limit - 1 settle at exactly limit, never more.
  • Members — the membership insert runs under a per-org advisory lock, counts, and rolls itself back if it crossed max_members. Re-adding an existing member stays a no-op.
  • Pending invitations — dedupe and the pending cap are enforced in one transaction under the same per-org lock; parallel duplicate-email invites yield exactly one row.
  • Public components — the cap is enforced when a monitor is added as a status-page component, counting distinct monitors across all of the org's pages in the same transaction as the insert.
  • API tokens — count-in-INSERT, scoped per user, handed max_api_tokens_per_user.

Exceeding a resource quota returns 422:

{
  "error": {
    "code": "QUOTA_EXCEEDED",
    "message": "max_targets limit reached: 20 of 20 used on the free plan.",
    "field": null,
    "details": { "quota": "max_targets", "current": 20, "limit": 20, "plan": "free" },
    "trace_id": null
  }
}

The pending-invitation cap is the one exception to the code: it predates the unified envelope and returns 409 INVITATIONS_LIMIT. The cap itself is enforced identically (atomic, never overshoot).

A sub-minimum check interval is its own 422, MIN_CHECK_INTERVAL, enforced on create and PATCH, single and bulk — a target created at the floor cannot be edited below it. The floor is max(plan.min_check_interval_secs, kind_min): the per-kind value (43200 for domain_expiry, 3600 for tls_cert, 300 for flow, 60 for heartbeat, 10 for the rest) applies regardless of plan tier — polling an expiry probe faster yields no signal, and domain_expiry reads RDAP, which rate-limits by source address.

Rate limiting

Two app-side tiers, both keyed on the authenticated subject (never the TCP peer): (org, category) and (user, category). Both are checked; the org tier fires first because it protects shared resources. The per-minute budget comes from the org's plan, except for support, which is fixed. The request category is derived from the path and method:

  • path contains /bulkbulk_ops
  • path ends /testtest_now
  • path ends /check-nowcheck_now
  • path ends /supportsupport
  • any path under /mcpapi_reads, whatever the method (the JSON-RPC body hides the tool name from the middleware; probe-spawning and write tools re-check the stricter category inside the tool)
  • otherwise GET/HEAD/OPTIONSapi_reads, else → api_writes

Exceeding a budget returns 429 with a Retry-After header:

{
  "error": {
    "code": "RATE_LIMITED",
    "message": "Too many requests.",
    "field": null,
    "details": { "scope": "per_org_api_writes", "retry_after_secs": 30 },
    "trace_id": null
  }
}

The limiter is a governor cell per (scope, category) key in a DashMap. A janitor evicts entries idle past the threshold so the map stays bounded by the number of active tenants, not by request volume; its lifetime is bound to the limiter so a refactor cannot silently drop the sweep and leak the map. Unauthenticated requests fall through untouched — per-IP limiting for those (auth endpoints, org creation, the public status surface) is the reverse proxy's job; see Deployment.

Checks themselves are not rate-limited — the scheduler path never enters this middleware, so monitoring throughput is unaffected.

Every quota / rate-limit / abuse rejection is recorded to the append-only quota_events table (event, quota_name, details, hashed IP) as fire-and-forget — it never blocks the response. It is the data source for abuse review.

Usage transparency

EndpointReturns
GET /api/v1/orgs/{id}/usagePlan + current vs limit for every org-scoped quota, policy values, rate budgets, feature flags. Member-gated (a non-member gets the same 404 as GET /orgs/{id}).
GET /api/v1/me/usageThe caller's api_tokens and owned_orgs current/limit.

The operator UI surfaces the same numbers at /settings/usage as progress bars (an unlimited self-host limit renders as ∞). Reported limit == enforced limit by construction: both read the same plan and the same count query.

Anti-abuse

Two deny-lists, applied when a target is created, bulk-created, updated, or test-run. A block is a 400, audited to quota_events with event = abuse_blocked.

  • URL patterns — a case-insensitive regex set of attack-recon paths (exposed VCS dirs, .env, credential paths, admin panels, WordPress xmlrpc pingback, Spring actuator, backup/dump extensions, …). A match is 400 URL_PATTERN_BLOCKED / ABUSE_BLOCKED. The shipped patterns and the compiled fallback are kept byte-identical by a drift guard.
  • Domains — a YAML deny-list (config/abuse_denylist.yaml) matched hierarchically: listing example.com also blocks eu.status.example.com. It carries the operator's own domain (don't monitor yourself) and competing uptime/status providers (monitoring another monitor forms a load-amplification chain). A match is 400 DOMAIN_DENYLISTED. Dedicated monitoring SaaS are listed at the apex; multi-tenant status-page hosts are listed narrowly so legitimate vendor-status checks are not over-blocked.

The lists load at startup. With abuse.hot_reload_enabled set, sending SIGHUP re-reads and validates them and swaps them in atomically; a bad edit keeps the old rules. Without it, changes need a restart. A bad regex or malformed YAML at startup is a clean config error, never a crash loop.

Configuration

[quotas]
plan_cache_ttl_secs  = 300   # org→plan cache; a plans-table edit takes
usage_cache_ttl_secs = 10    #   effect within this window
default_plan         = "pro"   # plan the boot-seeded owner org is placed on

A plans-table change is invisible until the plan cache's TTL elapses (a cache hit is zero DB round-trips on the hot path), then the next lookup refetches.

free is priced for a shared platform: its ceilings bound what one tenant can cost the host. On your own hardware there is no such cost, so a self-hosted install needs no setup here — default_plan already defaults to pro, giving the seeded owner org 150 monitors, a 30s check floor and 13-month retention.

Only boot-time seeding reads it, so it applies to the owner org that first run creates and to nothing else. The operator CLI (bootstrap-owner) does not use it, and an org you re-plan later is never moved back on the next boot. Orgs created afterwards go through signup, which grants founding until that tier's cutoff and free after it — so a second org on a self-hosted box lands on founding, not on pro. Set default_plan = "free" to opt out entirely.

The plan id is resolved against the plans table before the seed writes anything, so a typo fails the boot with the name quoted and leaves no half-seeded account behind.

Quota values still live only in Postgres — default_plan chooses a plan, it does not override any number in one. Raise limits the way SaaS does: edit (or INSERT) the plans row the org is assigned to, or attach a plan_overrides row with the cap fields you want to raise, so the audit trail covers both modes.

Every numeric quota / rate / interval is validated at config load — < 1 is rejected with the offending field named, never a panic in router or limiter construction.

The reverse-proxy per-IP tiers (auth endpoints, org creation, public surface) are documented in Deployment.