[{"data":1,"prerenderedAt":6},["ShallowReactive",2],{"post-content-laravel-job-ran-twice-retry-after-timeout":3},{"content":4,"lastModified":5},"\u003Cp>A customer forwarded me two copies of the same invoice email, four minutes apart. Our Laravel queue log showed the job finishing successfully twice: one payload, one job ID, and two lines ending with a cheerful \"invoice sent\" message. No exception, nothing in \u003Ccode>failed_jobs\u003C\u002Fcode>, no alert anywhere.\u003C\u002Fp>\n\n\u003Cp>I spent the first hour hunting for a double dispatch in the controller. It wasn't there. The job took a little over three minutes on that account because of a slow SMTP handshake, and the connection had \u003Ccode>retry_after\u003C\u002Fcode> sitting at the framework default of 90 seconds. Worker A was still holding the job when Redis decided it had died, put it back on the queue, and worker B ran the whole thing a second time. Worker A then finished normally. Two successes, one email too many.\u003C\u002Fp>\n\n\u003Cimg src=\"https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1563861826100-9cb868fdbe1c?w=1200&amp;q=80\" alt=\"A minimal round wall clock with black hands and an orange second hand on a plain white wall\" loading=\"lazy\" \u002F>\n\n\u003Ch2>One connection, two clocks\u003C\u002Fh2>\n\n\u003Cp>The setting that leaked the job is \u003Ccode>retry_after\u003C\u002Fcode>, defined per connection in \u003Ccode>config\u002Fqueue.php\u003C\u002Fcode>. The Laravel docs describe it plainly: with a value of 90, \"the job will be released back onto the queue if it has been processing for 90 seconds without being released or deleted\". It is not a timeout in the killing sense. Nothing gets stopped. The job is simply assumed dead and handed to somebody else.\u003C\u002Fp>\n\n\u003Cp>The setting that actually kills a job is the worker's \u003Ccode>--timeout\u003C\u002Fcode>, 60 seconds by default, and it only does anything if the PCNTL extension is loaded. You can also put a ceiling on a single job class, and that value wins over the command line.\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-php\">#[Timeout(300)]\nclass SendInvoice implements ShouldQueue\n{\n    \u002F\u002F ...\n}\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Two numbers, two owners, and one rule connecting them. Straight from the docs: the \u003Ccode>--timeout\u003C\u002Fcode> value \"should always be at least several seconds shorter than your \u003Ccode>retry_after\u003C\u002Fcode> configuration value\", because otherwise \"your jobs may be processed twice\". That is exactly what we had done. Ninety seconds on the connection, ninety on the CLI flag, no gap at all.\u003C\u002Fp>\n\n\u003Ch2>Auditing my own server\u003C\u002Fh2>\n\n\u003Cp>Reading \u003Ccode>config\u002Fqueue.php\u003C\u002Fcode> is not enough, because the worker timeout can live in a Supervisor block, a Docker command, or a Horizon supervisor. So I printed the effective values instead. First \u003Ccode>retry_after\u003C\u002Fcode> per connection, since env vars get involved:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">php artisan tinker --execute=\"\nforeach (config('queue.connections') as \\$name => \\$connection) {\n    if (! is_array(\\$connection)) {\n        continue; \u002F\u002F 'default' holds a connection name, not a config array\n    }\n    dump(\\$name.' => '.(\\$connection['retry_after'] ?? 'no retry_after on this driver'));\n}\"\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Then the workers actually running:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">ps -eo pid,etime,args | grep '[q]ueue:work'\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Ours printed \u003Ccode>--timeout=90\u003C\u002Fcode>, copied from the connection value years ago by someone trying to make the two numbers agree. They agreed, and that was the bug. Both commands live in \u003Ca href=\"\u002Fsnippetark\u002F\">Snippet Ark\u003C\u002Fa> now so the next person does not have to fight the escaping in that tinker one-liner.\u003C\u002Fp>\n\n\u003Ch2>The fix that stuck\u003C\u002Fh2>\n\n\u003Cp>Pushing \u003Ccode>retry_after\u003C\u002Fcode> to 300 everywhere is the wrong move, because that number doubles as the delay before a genuinely dead job becomes available again. Fast jobs should come back fast. Since the setting belongs to the connection and not to the job, the clean fix is a second connection for slow work.\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-php\">\u002F\u002F config\u002Fqueue.php\n'redis' => [\n    'driver' => 'redis',\n    'connection' => 'default',\n    'queue' => 'default',\n    'retry_after' => 90,\n],\n\n'redis-slow' => [\n    'driver' => 'redis',\n    'connection' => 'default',\n    'queue' => 'slow',\n    'retry_after' => 320,\n],\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Then the slow job asks for that connection and declares its own ceiling, leaving the gap the rule asks for:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-php\">SendInvoice::dispatch($invoice)->onConnection('redis-slow');\n\n#[Timeout(300)]\nclass SendInvoice implements ShouldQueue { \u002F* ... *\u002F }\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Three hundred under three hundred and twenty leaves the worker twenty seconds to release the job properly. Do not forget the HTTP client either. The docs are explicit that sockets and outgoing calls do not always respect the job timeout, so the Guzzle request needs its own connect and request timeout. Ours now fails at ten seconds instead of holding a worker for three minutes.\u003C\u002Fp>\n\n\u003Ch2>Assume the job will run twice anyway\u003C\u002Fh2>\n\n\u003Cp>Even with perfect numbers, a deploy, an OOM kill, or a hard reboot will eventually hand one job to two workers. \u003Ccode>WithoutOverlapping\u003C\u002Fcode> is the built-in defense, keyed on whatever identifies the work, and \u003Ccode>expireAfter()\u003C\u002Fcode> matters more than people expect because the lock is not always released when a worker dies. One caveat from the docs bit me during testing: releasing an overlapping job back onto the queue still increments its attempt count, so a job left at one attempt never gets a second chance. Raise \u003Ccode>Tries\u003C\u002Fcode> or \u003Ccode>MaxExceptions\u003C\u002Fcode> if you use that middleware.\u003C\u002Fp>\n\n\u003Cp>The other half lives at the business layer: our invoice job now checks a \u003Ccode>sent_at\u003C\u002Fcode> column inside the transaction and returns early if it is set. Not elegant, but no queue setting can promise single delivery and one line of PHP can.\u003C\u002Fp>\n\n\u003Cp>If the job was slow because of the queries it ran rather than the provider on the other end, \u003Ca href=\"\u002Fposts\u002Flaravel-n-plus-one-queries-find-fix\u002F\">the N+1 walkthrough\u003C\u002Fa> covers that end of it.\u003C\u002Fp>\n\n\u003Ch2>What I check first now\u003C\u002Fh2>\n\n\u003Cp>When a job looks like it ran twice, the sequence is short. Compare \u003Ccode>retry_after\u003C\u002Fcode> on the connection against the worker's real \u003Ccode>--timeout\u003C\u002Fcode> and look for a few seconds of gap. Then check that the job's own \u003Ccode>#[Timeout]\u003C\u002Fcode> still sits under \u003Ccode>retry_after\u003C\u002Fcode>. Then confirm the duplicate is even a duplicate, because a log line saying a job has \"been attempted too many times\" next to a job that logged success is the same bug with \u003Ccode>tries\u003C\u002Fcode> at one, not a retry loop.\u003C\u002Fp>\n\n\u003Cp>Reading this back out of the worker output is the tedious part. A busy queue writes thousands of lines an hour and the two starts of a single job ID are rarely near each other, so I filter by ID and read the whole thing in \u003Ca href=\"\u002Fstreamlog\u002F\">Streamlog\u003C\u002Fa> rather than scrolling a terminal. \u003Ca href=\"\u002Fposts\u002Fgrep-alternatives-for-large-logs\u002F\">The log grep post\u003C\u002Fa> has the shell version if you would rather stay in the terminal.\u003C\u002Fp>\n\n\u003Cp>The uncomfortable part: the job really was slow. The queue settings were the wound, not the illness. A three-minute job that should take two seconds deserves its own ticket, and aligning these numbers only buys time to reach it.\u003C\u002Fp>\n","2026-09-27",1790578737626]