•5 min read

Laravel Jobs Running Twice: retry_after vs timeout

A customer forwarded me two copies of the same invoice email, four minutes apart. Our Laravel queue log showed the job finishing successfully twice: one payload, one job ID, and two lines ending with a cheerful "invoice sent" message. No exception, nothing in failed_jobs, no alert anywhere.

I spent the first hour hunting for a double dispatch in the controller. It wasn't there. The job took a little over three minutes on that account because of a slow SMTP handshake, and the connection had retry_after sitting at the framework default of 90 seconds. Worker A was still holding the job when Redis decided it had died, put it back on the queue, and worker B ran the whole thing a second time. Worker A then finished normally. Two successes, one email too many.

A minimal round wall clock with black hands and an orange second hand on a plain white wall

One connection, two clocks

The setting that leaked the job is retry_after, defined per connection in config/queue.php. The Laravel docs describe it plainly: with a value of 90, "the job will be released back onto the queue if it has been processing for 90 seconds without being released or deleted". It is not a timeout in the killing sense. Nothing gets stopped. The job is simply assumed dead and handed to somebody else.

The setting that actually kills a job is the worker's --timeout, 60 seconds by default, and it only does anything if the PCNTL extension is loaded. You can also put a ceiling on a single job class, and that value wins over the command line.

#[Timeout(300)]
class SendInvoice implements ShouldQueue
{
    // ...
}

Two numbers, two owners, and one rule connecting them. Straight from the docs: the --timeout value "should always be at least several seconds shorter than your retry_after configuration value", because otherwise "your jobs may be processed twice". That is exactly what we had done. Ninety seconds on the connection, ninety on the CLI flag, no gap at all.

Auditing my own server

Reading config/queue.php is not enough, because the worker timeout can live in a Supervisor block, a Docker command, or a Horizon supervisor. So I printed the effective values instead. First retry_after per connection, since env vars get involved:

php artisan tinker --execute="
foreach (config('queue.connections') as \$name => \$connection) {
    if (! is_array(\$connection)) {
        continue; // 'default' holds a connection name, not a config array
    }
    dump(\$name.' => '.(\$connection['retry_after'] ?? 'no retry_after on this driver'));
}"

Then the workers actually running:

ps -eo pid,etime,args | grep '[q]ueue:work'

Ours printed --timeout=90, copied from the connection value years ago by someone trying to make the two numbers agree. They agreed, and that was the bug. Both commands live in Snippet Ark now so the next person does not have to fight the escaping in that tinker one-liner.

The fix that stuck

Pushing retry_after to 300 everywhere is the wrong move, because that number doubles as the delay before a genuinely dead job becomes available again. Fast jobs should come back fast. Since the setting belongs to the connection and not to the job, the clean fix is a second connection for slow work.

// config/queue.php
'redis' => [
    'driver' => 'redis',
    'connection' => 'default',
    'queue' => 'default',
    'retry_after' => 90,
],

'redis-slow' => [
    'driver' => 'redis',
    'connection' => 'default',
    'queue' => 'slow',
    'retry_after' => 320,
],

Then the slow job asks for that connection and declares its own ceiling, leaving the gap the rule asks for:

SendInvoice::dispatch($invoice)->onConnection('redis-slow');

#[Timeout(300)]
class SendInvoice implements ShouldQueue { /* ... */ }

Three hundred under three hundred and twenty leaves the worker twenty seconds to release the job properly. Do not forget the HTTP client either. The docs are explicit that sockets and outgoing calls do not always respect the job timeout, so the Guzzle request needs its own connect and request timeout. Ours now fails at ten seconds instead of holding a worker for three minutes.

Assume the job will run twice anyway

Even with perfect numbers, a deploy, an OOM kill, or a hard reboot will eventually hand one job to two workers. WithoutOverlapping is the built-in defense, keyed on whatever identifies the work, and expireAfter() matters more than people expect because the lock is not always released when a worker dies. One caveat from the docs bit me during testing: releasing an overlapping job back onto the queue still increments its attempt count, so a job left at one attempt never gets a second chance. Raise Tries or MaxExceptions if you use that middleware.

The other half lives at the business layer: our invoice job now checks a sent_at column inside the transaction and returns early if it is set. Not elegant, but no queue setting can promise single delivery and one line of PHP can.

If the job was slow because of the queries it ran rather than the provider on the other end, the N+1 walkthrough covers that end of it.

What I check first now

When a job looks like it ran twice, the sequence is short. Compare retry_after on the connection against the worker's real --timeout and look for a few seconds of gap. Then check that the job's own #[Timeout] still sits under retry_after. Then confirm the duplicate is even a duplicate, because a log line saying a job has "been attempted too many times" next to a job that logged success is the same bug with tries at one, not a retry loop.

Reading this back out of the worker output is the tedious part. A busy queue writes thousands of lines an hour and the two starts of a single job ID are rarely near each other, so I filter by ID and read the whole thing in Streamlog rather than scrolling a terminal. The log grep post has the shell version if you would rather stay in the terminal.

The uncomfortable part: the job really was slow. The queue settings were the wound, not the illness. A three-minute job that should take two seconds deserves its own ticket, and aligning these numbers only buys time to reach it.