jq -e in Bash Scripts: The Check That Passes When Nothing Matched
The cron job that ruined my Tuesday did not fail. It exited 0, printed that the queue was idle, and restarted the worker on top of two jobs that were mid-flight.
I wrote that script a year ago and had not opened it since. It polls our queue API for running jobs and restarts the worker when there are none. Tuesday there were two, and it restarted anyway. I spent the morning replaying a batch that should have taken twenty minutes.
The culprit was a jq command. If you use jq to drive bash conditionals, the exit code is doing more work than you think.
What -e actually tests
jq on its own exits 0 whenever it managed to parse the input, no matter what it found. A missing key, a null, an empty array: all 0. That is the entire reason -e exists.
The manual is brief about it. With -e, jq exits 0 if the last value it printed was neither false nor null, 1 if that value was false or null, and 4 if it printed nothing at all. The other codes are plumbing: 2 for a usage problem, 3 for a compile error in your filter, 5 for a runtime error like indexing a string.
$ jq -e '.timeout' <<< '{"retries":3}'; echo "exit=$?"
null
exit=1
$ jq -e '.[] | select(.state == "running")' <<< '[{"state":"queued"}]'; echo "exit=$?"
exit=4
Four is the code people do not expect. A bare jq check that finds nothing exits 4, so under set -e bash stops your script right there. That was not my bug, but I see it in review constantly.
The pair of brackets that flips the answer
Here is my actual bug, reduced to its essentials.
if jq -e '[.jobs[] | select(.state == "running")]' jobs.json > /dev/null; then
echo "queue is idle, restarting the worker"
fi
Wrapping a filter in [ ] makes jq collect the results into an array, and that array is a value. It is not false and it is not null, even when empty. So the exit code is 0, the if takes the true branch, and the script reports an idle queue with two jobs running.
$ jq -e '[.jobs[] | select(.state == "running")]' jobs.json; echo "exit=$?"
[]
exit=0
I add brackets reflexively, because collected output is nicer to read than a run of separate objects. The bare version is the one that tests something. When I want a yes or no, jq -e 'any(.jobs[]; .state == "running")' prints true or false and sets the exit code to match.
Missing, null, and false
A key that does not exist and a key whose value is null both make jq print null, and both exit 1 under -e. To tell them apart, ask the object rather than the value.
$ jq -e 'has("timeout")' <<< '{"timeout":null}'; echo "exit=$?"
true
exit=0
$ jq -e '.timeout != null' <<< '{"timeout":null}'; echo "exit=$?"
false
exit=1
I use has() for config validation, where a deleted key and an explicit null mean different things downstream.
// is the other half of this, and it is quietly dangerous. It produces the values on the left that are neither false nor null, which means a flag explicitly set to false counts as absent.
$ jq -c '.enabled // true' <<< '{"enabled":false}'
true
Numbers and empty strings survive, so the bug is asymmetric. 0 // 5 keeps the 0, and a boolean flag ends up true.
Reading jq output one item at a time
Most of my jq loops want raw strings. read -r protects spaces. It does not protect newlines, and one value containing a newline becomes two iterations.
$ jq -r '.names[]' names.json
web 01
line
break
plain
Those are three names, not four. The middle one is a single string with a newline inside it. The for name in $(...) form is worse, since word splitting also breaks on the space in web 01.
jq has a flag for this. --raw-output0 writes a NUL byte after each value instead of a newline, and the manual says outright that it exists for values containing newlines.
while IFS= read -r -d '' name; do
printf '[%s]\n' "$name"
done < <(jq --raw-output0 '.names[]' names.json)
One caveat: if a value contains a NUL byte, jq gives up and exits non-zero. That is the right call, since a NUL cannot survive in a shell variable anyway. The same framing also lets you hand the items to xargs -0.
Three ways I have misread the exit code
A check is only worth writing if the code reading it gets the status.
jq -e '.jobs[] | select(.state == "failed")' jobs.json | wc -l
echo $? # 0, and not jq's 0
A pipeline reports the status of its last command, so jq's 4 never reaches you. set -o pipefail fixes that, which is why strict mode and jq checks turn up in the same script.
port=$(jq -e '.port' config.json) # fails here under set -e
local port=$(jq -e '.port' config.json) # does not
local and declare are builtins. They run, they succeed, and their success overwrites the status of the substitution inside them. I have shipped that bug twice. Assign on one line, mark it local on the next.
The third one comes from a habit. jq -e . is a bad JSON validator, because a valid document whose top level is false or null exits 1 and your validation step rejects good input. Use jq empty, which parses and prints nothing: 0 for valid input, 5 for anything else.
What I do now
jq gets asked questions in a form that has only one answer. any() instead of a bracketed array, has() when missing and null differ, --raw-output0 when the values come from a system I do not control, and -e only where something reads the exit code.
The wrapper I reach for most is four lines:
# jq -er, with a message when the value is absent or null
need() {
jq -er "$2" "$1" || { echo "config: $2 missing in $1" >&2; return 1; }
}
port=$(need config.json '.port') || exit 1
It reads strictly, so a null port stops the deploy instead of arriving as the four-character string "null" in a connection string. Strict reading also rejects false, which is what I want for ports and exactly wrong for boolean flags.
The queue script now uses any(...), and I left a comment explaining why there are no brackets in it. Past me is precisely the person who would add them back. The jq recipes I stop retyping live in Snippet Ark.