On September 28 we shipped a routine backend release. The pipeline went green, and for the next 46 minutes production processed no queued jobs at all. Customers noticed first. Restaurants saw their kitchen tickets stop printing and started reprinting them by hand.

What broke

Every asynchronous job in that backend runs through Laravel Horizon, including the job that sends kitchen tickets to the printer. Horizon runs as a systemd service on two worker VMs, and every deploy reloads systemd and then runs a stop followed by a start on that service.

The unit's stop command isn't a plain signal from systemd. It calls horizon:terminate:

[Service]
ExecStart=/usr/bin/php /var/app/current/artisan horizon
ExecStop=/usr/bin/php /var/app/current/artisan horizon:terminate --wait
Restart=always
RestartSec=5
TimeoutStopSec=30

horizon:terminate is a graceful shutdown, and it reaches further than the unit's own process. It looks up every Horizon master registered in Redis, sends each one SIGTERM, and sets a restart marker in the shared cache.

What the logs show

Both workers restarted within about two seconds of each other. Horizon's logs show two TERM signals and a single successful start, and the second TERM went out 4 milliseconds after that start. Horizon's log volume on the workers then dropped to zero at about 15:05 (UTC-5) and stayed there until we started it by hand at 15:51. The deploy tool reported every step as successful, and CI marked the deploy green.

The root cause in our post-mortem is a race between that graceful stop and the new start: after the stop/start sequence, neither worker was left processing jobs. We're still correlating the exact order of signals on each host. I also can't explain yet why Restart=always didn't bring Horizon back, because we didn't capture the unit's state before the manual start. Until that's answered, the mechanism is an open question.

The code in the release wasn't the problem. Starting Horizon by hand on the same release brought processing back.

Why it took 46 minutes

  • Nothing checked Horizon after the deploy. The deploy ran zero health checks, so a release that left the queues stopped still finished green.
  • Nothing alerted on Horizon or on queue growth. The first customer report came 11 minutes after processing stopped, and infrastructure got involved at minute 33.
  • The symptom looked like a different problem. A ticket with zero print attempts looks a lot like a restaurant without a printer configured, so support and development checked POS and printer settings first.
  • Getting onto the servers took longer than it should have. The authorized SSH keys were set up for the ubuntu user on every server, while our Azure VMs use azure.
  • The escalation to me over chat didn't get an immediate answer and ended in a phone call. Seven minutes passed between the first question and my reply.

Recovery

We started Horizon by hand, first on the web server and then on both Azure workers, and processing resumed at about 15:51. Log volume peaked near 5,000 lines while the backlog drained, then settled back to its usual range. We didn't need to roll back the release. Backend deploys were frozen that afternoon.

What we're changing

None of this is done yet. The plan:

  • Put Horizon under the servers' process monitor, so it gets restarted, or someone gets alerted, whenever it isn't running.
  • Change the unit so systemd stops Horizon with a plain SIGTERM to its main process, and set TimeoutStopSec to fit the longest job.
  • Add a post-deploy check that fails the deploy when Horizon isn't active.
  • Restart workers one at a time and check each one before moving on.
  • Alert on queue growth and on Horizon going quiet.
  • Fix the SSH user per provider, and confirm that development and production don't share Horizon's Redis prefix.
  • Review how production incidents are escalated, so reaching the person on call doesn't take a phone call.

The post-deploy check can be small. Something like this, run on each worker after the restart:

systemctl is-active --quiet horizon@default || exit 1
php artisan horizon:status | grep -q "is running" || exit 1

What I took from it

A deploy that finishes without errors only tells you its steps ran. When a deploy restarts a long-running process, it should check that process afterward and fail if the process isn't doing its job. And a queue that stops should page someone well before a customer calls support.