Split Horizon into priority-based supervisors, add balanceCooldown

and notification routing

All 15 queues previously ran through one auto-balanced supervisor.
Horizon's `balance: auto` does not honor queue array order for
priority, so despite queue names implying priority ('high' vs 'low'),
a burst on any one queue could starve any other sharing that
supervisor - e.g. a burst of mmo (image/video optimization, 23
dispatch sites, CPU/IO heavy) could delay high-queue DM/follow
delivery just as easily as it could delay low-queue background work.

Split into 4 supervisors grouped by actual job characteristics
(checked via grep across every ->onQueue() call site, not guessed):
- supervisor-priority: high, inbox, pushnotify, follow, default,
  shared - user-facing federation/DM/notification delivery.
- supervisor-fanout: feed, story, groups - bursty timeline/story
  fanout writes triggered by posts, likes, and follows.
- supervisor-media: mmo - image/video optimize/resize/thumbnail.
  Runs a fixed worker pool (balance: false) instead of auto-scaling,
  so it can't claim workers away from the other pools under load.
- supervisor-background: low, delete, adelete, move, intbg - imports,
  crawling, account deletion/migration; not time-sensitive.

Moved the shared supervisor shape into `defaults` (keyed per
supervisor name, per Horizon's own merge behavior) so `environments`
only needs to override what actually differs, instead of each
environment fully redefining supervisor-1 from scratch. Existing env
vars (HORIZON_MAX_PROCESSES, HORIZON_MIN_PROCESSES,
HORIZON_BALANCE_STRATEGY, HORIZON_SUPERVISOR_*) keep governing the
priority supervisor for continuity with existing deployments; the
three new supervisors get their own HORIZON_*_MAX_PROCESSES vars
with conservative defaults.

Also:
- Added balanceCooldown: 3 explicitly (previously relied on
  SupervisorOptions' own constructor default of the same value -
  behavior is unchanged, just no longer implicit).
- Wired LongWaitDetected notification routing
  (Horizon::routeMailNotificationsTo/routeSlackNotificationsTo) to
  new optional config('horizon.notification_routing') keys, sourced
  from env vars. Previously these were hardcoded, commented-out
  examples with nowhere to actually alert on the `waits` thresholds
  already configured below.

Verified by actually starting `php artisan horizon` and inspecting
`horizon:supervisors`: all 4 supervisors registered with exactly the
intended queues, supervisor-media correctly running fixed (non-auto)
balancing. Cross-checked every ->onQueue() call site in app/ against
the new supervisor queue lists - exact match, no queue dropped or
duplicated. Full test suite (715/715) and Larastan clean.
pull/7198/head
Your Name 2 weeks ago
parent cedfe1a0c1
commit bbe7cfa8e1

@ -17,9 +17,16 @@ class HorizonServiceProvider extends HorizonApplicationServiceProvider
{
parent::boot();
// Horizon::routeSmsNotificationsTo('15556667777');
// Horizon::routeMailNotificationsTo('example@example.com');
// Horizon::routeSlackNotificationsTo('slack-webhook-url', '#channel');
if ($mailTo = config('horizon.notification_routing.mail_to')) {
Horizon::routeMailNotificationsTo($mailTo);
}
if ($slackWebhook = config('horizon.notification_routing.slack_webhook_url')) {
Horizon::routeSlackNotificationsTo(
$slackWebhook,
config('horizon.notification_routing.slack_channel')
);
}
if (config('horizon.darkmode') == true) {
Horizon::night();

@ -67,6 +67,24 @@ return [
'middleware' => ['web'],
/*
|--------------------------------------------------------------------------
| Notification Routing
|--------------------------------------------------------------------------
|
| Horizon can notify you when a queue's LongWaitDetected threshold (see
| the `waits` option below) is exceeded. These are read by
| HorizonServiceProvider::boot() and are all optional; leave them unset
| to disable notification delivery entirely.
|
*/
'notification_routing' => [
'mail_to' => env('HORIZON_NOTIFICATIONS_MAIL'),
'slack_webhook_url' => env('HORIZON_NOTIFICATIONS_SLACK_WEBHOOK'),
'slack_channel' => env('HORIZON_NOTIFICATIONS_SLACK_CHANNEL'),
],
/*
|--------------------------------------------------------------------------
| Queue Wait Time Thresholds
@ -171,35 +189,97 @@ return [
| in all environments. These supervisors and settings handle all your
| queued jobs and will be provisioned by Horizon during deployment.
|
| Queues are split across supervisors by workload rather than lumped into
| one auto-balanced pool, because Horizon's `balance: auto` does not
| honor queue array order for priority - a burst on one queue can starve
| another sharing the same supervisor regardless of their names:
|
| - supervisor-priority: user-facing federation delivery (follows,
| deletes, DMs), inbox processing, and push notifications. These
| should never wait behind a burst of media processing or feed fanout.
| - supervisor-fanout: timeline/story/group fanout writes triggered by
| posts, likes, and follows. Bursty (one post can fan out to many
| followers) but not CPU-heavy.
| - supervisor-media: image/video optimization, resizing, and thumbnail
| generation (the `mmo` queue). CPU/IO heavy, so it runs a fixed
| worker pool (`balance: false`) instead of auto-scaling, which would
| otherwise let it claim workers away from the other pools under load.
| - supervisor-background: imports, instance crawling, account
| deletion/migration, and other maintenance work that isn't
| user-facing time-sensitive.
|
*/
'defaults' => [
'supervisor-priority' => [
'connection' => 'redis',
'queue' => ['high', 'inbox', 'pushnotify', 'follow', 'default', 'shared'],
'balance' => env('HORIZON_BALANCE_STRATEGY', 'auto'),
'autoScalingStrategy' => 'time',
'balanceMaxShift' => 1,
'balanceCooldown' => 3,
'minProcesses' => env('HORIZON_MIN_PROCESSES', 1),
'maxProcesses' => env('HORIZON_MAX_PROCESSES', 10),
'memory' => env('HORIZON_SUPERVISOR_MEMORY', 64),
'tries' => env('HORIZON_SUPERVISOR_TRIES', 3),
'nice' => env('HORIZON_SUPERVISOR_NICE', 0),
'timeout' => env('HORIZON_SUPERVISOR_TIMEOUT', 300),
],
'supervisor-fanout' => [
'connection' => 'redis',
'queue' => ['feed', 'story', 'groups'],
'balance' => env('HORIZON_BALANCE_STRATEGY', 'auto'),
'autoScalingStrategy' => 'time',
'balanceMaxShift' => 1,
'balanceCooldown' => 3,
'minProcesses' => env('HORIZON_MIN_PROCESSES', 1),
'maxProcesses' => env('HORIZON_FANOUT_MAX_PROCESSES', 6),
'memory' => env('HORIZON_SUPERVISOR_MEMORY', 64),
'tries' => env('HORIZON_SUPERVISOR_TRIES', 3),
'nice' => env('HORIZON_SUPERVISOR_NICE', 0),
'timeout' => env('HORIZON_SUPERVISOR_TIMEOUT', 300),
],
'supervisor-media' => [
'connection' => 'redis',
'queue' => ['mmo'],
'balance' => false,
'minProcesses' => env('HORIZON_MIN_PROCESSES', 1),
'maxProcesses' => env('HORIZON_MEDIA_MAX_PROCESSES', 4),
'memory' => env('HORIZON_SUPERVISOR_MEMORY', 64),
'tries' => env('HORIZON_SUPERVISOR_TRIES', 3),
'nice' => env('HORIZON_SUPERVISOR_NICE', 0),
'timeout' => env('HORIZON_SUPERVISOR_TIMEOUT', 300),
],
'supervisor-background' => [
'connection' => 'redis',
'queue' => ['low', 'delete', 'adelete', 'move', 'intbg'],
'balance' => env('HORIZON_BALANCE_STRATEGY', 'auto'),
'autoScalingStrategy' => 'time',
'balanceMaxShift' => 1,
'balanceCooldown' => 3,
'minProcesses' => env('HORIZON_MIN_PROCESSES', 1),
'maxProcesses' => env('HORIZON_BACKGROUND_MAX_PROCESSES', 4),
'memory' => env('HORIZON_SUPERVISOR_MEMORY', 64),
'tries' => env('HORIZON_SUPERVISOR_TRIES', 3),
'nice' => env('HORIZON_SUPERVISOR_NICE', 0),
'timeout' => env('HORIZON_SUPERVISOR_TIMEOUT', 300),
],
],
'environments' => [
'production' => [
'supervisor-1' => [
'connection' => 'redis',
'queue' => ['high', 'default', 'follow', 'shared', 'inbox', 'feed', 'low', 'story', 'delete', 'mmo', 'intbg', 'groups', 'adelete', 'move', 'pushnotify'],
'balance' => env('HORIZON_BALANCE_STRATEGY', 'auto'),
'minProcesses' => env('HORIZON_MIN_PROCESSES', 1),
'maxProcesses' => env('HORIZON_MAX_PROCESSES', 20),
'memory' => env('HORIZON_SUPERVISOR_MEMORY', 64),
'tries' => env('HORIZON_SUPERVISOR_TRIES', 3),
'nice' => env('HORIZON_SUPERVISOR_NICE', 0),
'timeout' => env('HORIZON_SUPERVISOR_TIMEOUT', 300),
],
// All values come from `defaults` above (env-configurable);
// nothing production-specific to override at this time.
],
'local' => [
'supervisor-1' => [
'connection' => 'redis',
'queue' => ['high', 'default', 'follow', 'shared', 'inbox', 'feed', 'low', 'story', 'delete', 'mmo', 'intbg', 'groups', 'adelete', 'move', 'pushnotify'],
'balance' => 'auto',
'minProcesses' => 1,
'maxProcesses' => 20,
'memory' => 128,
'tries' => 3,
'nice' => 0,
'timeout' => 300,
],
'supervisor-priority' => ['maxProcesses' => 4],
'supervisor-fanout' => ['maxProcesses' => 2],
'supervisor-media' => ['maxProcesses' => 2],
'supervisor-background' => ['maxProcesses' => 2],
],
],

Loading…
Cancel
Save