[{"data":1,"prerenderedAt":6},["ShallowReactive",2],{"post-content-xargs-p-parallel-shell-jobs":3},{"content":4,"lastModified":5},"\u003Cp>Last month the nightly archive job on our log box crossed 41,000 files and became something I had to babysit. Every rotated log under \u003Ccode>\u002Fvar\u002Flog\u002Farchive\u003C\u002Fcode> gets compressed at level 6 before it moves to the cold tier, and the loop I wrote two years ago does it one file at a time. On a box with twelve idle cores.\u003C\u002Fp>\n\n\u003Cp>Here is that loop. The fix is short; the traps around it cost me an afternoon.\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">#!\u002Fusr\u002Fbin\u002Fenv bash\nfind \u002Fvar\u002Flog\u002Farchive -name '*.log' -print0 |\nwhile IFS= read -r -d '' f; do\n  gzip -6 \"$f\" 2&gt;\u002Fdev\u002Fnull\ndone\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>It works, and it is 41,000 process launches in a row: about six hours.\u003C\u002Fp>\n\n\u003Ch2>Why You Can't Just Pipe Into a Command\u003C\u002Fh2>\n\n\u003Cp>My first attempt at it was worse:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">find \u002Fvar\u002Flog\u002Farchive -name '*.log' | rm\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>That exits 0 and deletes nothing. \u003Ccode>rm\u003C\u002Fcode> never reads standard input, and neither do \u003Ccode>gzip\u003C\u002Fcode>, \u003Ccode>cp\u003C\u002Fcode> or \u003Ccode>chmod\u003C\u002Fcode>. They take paths as arguments. A pipe hands them a stream of bytes, which they ignore.\u003C\u002Fp>\n\n\u003Cp>\u003Ccode>xargs\u003C\u002Fcode> is the adapter: it reads items off the pipe and builds command lines, because the argument list has a size limit. On this Mac, \u003Ccode>sysctl -n kern.argmax\u003C\u002Fcode> reports 1048576, and going over it looks like this:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">$ \u002Fbin\u002Fecho $(printf 'x%.0s' $(seq 1 2000000))\nbash: \u002Fbin\u002Fecho: Argument list too long\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>That is E2BIG, the error that sends people to the xargs manual at 2am.\u003C\u002Fp>\n\n\u003Ch2>The Whitespace Split\u003C\u002Fh2>\n\n\u003Cp>By default xargs splits input on any whitespace. Two files with spaces in their names:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">$ find . -name '*.txt' | xargs -n 1 echo TOKEN:\nTOKEN: .\u002Fmy\nTOKEN: file\nTOKEN: 1.txt\nTOKEN: .\u002Fmy\nTOKEN: file\nTOKEN: 2.txt\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Two files in, six tokens out. Had the command been \u003Ccode>mv\u003C\u002Fcode> or \u003Ccode>chmod\u003C\u002Fcode>, three of those tokens would be pointing at paths that do not exist. The fix is to make find emit NUL bytes instead of newlines, since filenames cannot contain a NUL byte:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">$ find . -name '*.txt' -print0 | xargs -0 -n 1 echo TOKEN:\nTOKEN: .\u002Fmy file 1.txt\nTOKEN: .\u002Fmy file 2.txt\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>\u003Ccode>-print0\u003C\u002Fcode> has been in GNU find for ages and POSIX added it in Issue 8 in 2024, so \u003Ccode>-print0 | xargs -0\u003C\u002Fcode> is not the careful version of a command, it is the default one. It also stops xargs from treating quotes and backslashes as special.\u003C\u002Fp>\n\n\u003Cp>Dry runs deserve a warning. This one wasted twenty minutes of my life:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">$ find . -name '*.txt' -print0 | xargs -0 echo rm\nrm .\u002Fmy file 1.txt .\u002Fmy file 2.txt\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Drop the \u003Ccode>-print0 | xargs -0\u003C\u002Fcode> and the line is identical, while the real command gets six broken paths. \u003Ccode>echo\u003C\u002Fcode> joins its arguments with spaces, so it flattens the exact problem you are trying to spot. \u003Ccode>-t\u003C\u002Fcode> is the honest flag: it prints each command line before running it.\u003C\u002Fp>\n\n\u003Ch2>Parallel, Which Is the Actual Point\u003C\u002Fh2>\n\n\u003Cp>\u003Ccode>-P\u003C\u002Fcode> is why I rewrote the job. The manual says to pair it with \u003Ccode>-n\u003C\u002Fcode> or \u003Ccode>-L\u003C\u002Fcode>, otherwise chances are only one exec happens, which is a sentence you only understand after watching a \"parallel\" run take exactly as long as the serial one. \u003Ccode>-P 0\u003C\u002Fcode> means as many processes as possible, which I would not do on a shared box.\u003C\u002Fp>\n\n\u003Cimg src=\"https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1518432031352-d6fc5c10da5a?w=1200&amp;q=80\" alt=\"Terminal screen showing a package download in progress next to a process list with many worker processes\" loading=\"lazy\" \u002F>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">JOBS=$(getconf _NPROCESSORS_ONLN 2&gt;\u002Fdev\u002Fnull || echo 4)\nARCHIVE=\u002Fvar\u002Flog\u002Farchive\n\nfind \"$ARCHIVE\" -name '*.log' -print0 |\n  xargs -0 -r -n 1 -P \"$JOBS\" gzip -6\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Six hours became forty minutes. Not the twelve times speedup the core count suggests, since all twelve workers push through the same disk, but I will take it. \u003Ccode>getconf _NPROCESSORS_ONLN\u003C\u002Fcode> prints 12 here and on the Linux runners. \u003Ccode>nproc\u003C\u002Fcode> is the Linux shortcut and \u003Ccode>sysctl -n hw.ncpu\u003C\u002Fcode> the macOS one; neither exists on both.\u003C\u002Fp>\n\n\u003Cp>One more thing about \u003Ccode>-P\u003C\u002Fcode>. If two children print to stdout, the manual says the output arrives in an indeterminate order and very likely mixed up. The gzip warnings survived that. A script printing a status line per file did not, and the fix was one output file per child.\u003C\u002Fp>\n\n\u003Ch2>The -I Flag Quietly Undoes the Batching\u003C\u002Fh2>\n\n\u003Cp>\u003Ccode>-I {}\u003C\u002Fcode> reads nicely and costs more than it looks like, because it implies \u003Ccode>-L 1\u003C\u002Fcode>: one invocation per input line, batching gone. GNU xargs also treats \u003Ccode>-L\u003C\u002Fcode>, \u003Ccode>-I\u003C\u002Fcode> and \u003Ccode>-n\u003C\u002Fcode> as mutually exclusive, keeps whichever came last, and warns on stderr, easy to miss in CI output. The exception is \u003Ccode>-n1\u003C\u002Fcode> after \u003Ccode>-I\u003C\u002Fcode>, ignored because it would not conflict.\u003C\u002Fp>\n\n\u003Cp>When a command wants the item somewhere other than the end of the line, look for a flag that takes a destination before reaching for \u003Ccode>-I\u003C\u002Fcode>. \u003Ccode>mv\u003C\u002Fcode> has \u003Ccode>-t\u003C\u002Fcode>, and this keeps the batching:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-bash\">find . -name '*.bak' -print0 | xargs -0 mv -t \u002Ftmp\u002Fbackups\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Ch2>Exit Codes, and the Failure You Don't See\u003C\u002Fh2>\n\n\u003Cp>GNU documents 123 when an invocation exits with anything other than 0 or 255, 124 for 255, 125 for a signal kill, 126 for cannot run, 127 for not found. My Mac's BSD xargs returned 1 when I ran a child that exited 1, not 123, so the portable reading of that number is \"non-zero\" and checking for 123 is a Linux habit.\u003C\u002Fp>\n\n\u003Cp>The empty input case is the one that bit me. On Linux, xargs runs the command once even when there is no input at all, unless you pass \u003Ccode>-r\u003C\u002Fcode>. macOS skips it. Same script, two behaviors, and the Linux one is dangerous, because your command runs with no arguments whatsoever. My wrapper read \u003Ccode>input=\"$1\"\u003C\u002Fcode>, and when the runner called it with an empty find result the variable was blank and it globbed the working directory instead. Nothing was lost that day. The log line said success, which is what annoyed me.\u003C\u002Fp>\n\n\u003Cp>\u003Ccode>-r\u003C\u002Fcode> is accepted without complaint by the macOS xargs I tested and is not optional on Linux, so leave it in. Pair it with \u003Ccode>set -euo pipefail\u003C\u002Fcode>, otherwise the pipeline's exit status is just the last command's and the xargs number never reaches you. I wrote about the rest of \u003Ca href=\"\u002Fposts\u002Fbash-strict-mode-set-euo-pipefail\u002F\">strict mode and its sharp edges\u003C\u002Fa> earlier.\u003C\u002Fp>\n\n\u003Ch2>What I Keep Now\u003C\u002Fh2>\n\n\u003Cp>For per-item work where items are independent, xargs wins. It runs several at a time, and with \u003Ccode>-print0\u003C\u002Fcode> and \u003Ccode>-r\u003C\u002Fcode> it fails loudly instead of guessing. When the loop body needs branching, shared counters, or variables that survive between iterations, a \u003Ccode>while read\u003C\u002Fcode> loop is still the right tool and I will not pretend otherwise.\u003C\u002Fp>\n\n\u003Cp>The version above lives in \u003Ca href=\"\u002Fsnippetark\u002F\">Snippet Ark\u003C\u002Fa>, JOBS line and \u003Ccode>-r\u003C\u002Fcode> already in it, so the next batch job starts from the script that works, not the one that hides its own errors.\u003C\u002Fp>\n","2026-09-22",1790289963468]