[{"data":1,"prerenderedAt":6},["ShallowReactive",2],{"post-content-postgres-not-using-my-index-what-i-check":3},{"content":4,"lastModified":5},"\u003Cp>A support ticket landed on my desk last month. A login endpoint had gone from 45ms to 2.6s. Postgres, twelve million rows in \u003Ccode>users\u003C\u002Fcode>, and a btree index on \u003Ccode>email\u003C\u002Fcode> that had been there for three years.\u003C\u002Fp>\n\n\u003Cp>So the index existed. I ran \u003Ccode>EXPLAIN ANALYZE\u003C\u002Fcode> and got a Seq Scan. Then I did what you do when you do not know what else to do: dropped the index and rebuilt it. Same plan.\u003C\u002Fp>\n\n\u003Cp>Once I read the plan instead of glaring at it, the answer was obvious. Here is the order I check things in now.\u003C\u002Fp>\n\n\u003Ch2>Start with the plan, not the index\u003C\u002Fh2>\n\n\u003Cp>Plain \u003Ccode>EXPLAIN\u003C\u002Fcode> shows what the planner intends to do. \u003Ccode>EXPLAIN (ANALYZE, BUFFERS)\u003C\u002Fcode> runs the query and shows what actually happened.\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-sql\">EXPLAIN (ANALYZE, BUFFERS)\nSELECT id, email, last_login_at\nFROM users\nWHERE LOWER(email) = LOWER($1);\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Two things matter most. Estimated rows versus actual rows: if the planner guessed 50 and the query returned 180,000, everything downstream sits on a bad number. And \u003Cstrong>Rows Removed by Filter\u003C\u002Fstrong>: a Seq Scan whose filter discards 99% of what it read is Postgres telling you, in arithmetic, that it is doing work you did not ask for.\u003C\u002Fp>\n\n\u003Cp>Mine said \u003Ccode>Rows Removed by Filter: 11999999\u003C\u002Fcode>.\u003C\u002Fp>\n\n\u003Cimg src=\"https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1558494949-ef010cbdcc31?auto=format&fit=crop&w=1200&q=80\" alt=\"A dark server rack with tangled network cables and rows of blinking status lights\" loading=\"lazy\" \u002F>\n\n\u003Ch2>The column is wrapped in something\u003C\u002Fh2>\n\n\u003Cp>A btree index stores the raw value of the column and nothing else. If your \u003Ccode>WHERE\u003C\u002Fcode> clause calls \u003Ccode>LOWER(email)\u003C\u002Fcode>, the planner has to evaluate that function on every row before it can compare anything, so an index on \u003Ccode>email\u003C\u002Fcode> is useless to it. Same story for \u003Ccode>date_trunc()\u003C\u002Fcode> and \u003Ccode>CAST()\u003C\u002Fcode>.\u003C\u002Fp>\n\n\u003Cp>So the fix is not \"add an index\", it is \"index the expression\":\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-sql\">CREATE INDEX users_email_lower_idx ON users (LOWER(email));\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>Two details bite. The query must use the exact same expression, since the planner matches expression trees structurally. And the function has to be \u003Ccode>IMMUTABLE\u003C\u002Fcode>, because the index layout is fixed at build time. \u003Ccode>LOWER()\u003C\u002Fcode> qualifies; \u003Ccode>to_char()\u003C\u002Fcode> is only \u003Ccode>STABLE\u003C\u002Fcode> and gets rejected.\u003C\u002Fp>\n\n\u003Cp>If the expression gets messy, a generated column is cleaner, but write \u003Ccode>STORED\u003C\u002Fcode> explicitly. Postgres 18 made \u003Ccode>VIRTUAL\u003C\u002Fcode> the default when you omit the keyword, and virtual columns cannot be indexed.\u003C\u002Fp>\n\n\u003Cp>My ORM called \u003Ccode>lower(email)\u003C\u002Fcode>; a hand-written reporting query used \u003Ccode>email = $1\u003C\u002Fcode>. One used the index. Guess which one got reported.\u003C\u002Fp>\n\n\u003Ch2>Maybe it is just not selective enough\u003C\u002Fh2>\n\n\u003Cp>Indexes are not free at read time. Every matching row is a potential random page read, and when your condition matches 40% of the table, walking the index and then visiting most of the heap costs more than reading the table straight through.\u003C\u002Fp>\n\n\u003Cp>What helps is an index on the part you actually query, not the whole column:\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-sql\">CREATE INDEX orders_pending_recent_idx\n  ON orders (created_at DESC)\n  WHERE status = 'pending';\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>A partial index only holds rows matching the predicate, so it stays small when one of a handful of \u003Ccode>status\u003C\u002Fcode> values owns the table. The catch: the query needs a matching \u003Ccode>WHERE\u003C\u002Fcode> clause, easy to forget six months later.\u003C\u002Fp>\n\n\u003Ch2>Statistics from last Tuesday\u003C\u002Fh2>\n\n\u003Cp>Every row-count estimate comes from \u003Ccode>pg_statistic\u003C\u002Fcode>, filled in by \u003Ccode>ANALYZE\u003C\u002Fcode> on autovacuum's schedule. Load five million rows in a nightly batch and the planner may still think the table is small, so \"this thing is tiny, just scan it\" is a reasonable decision made on stale data.\u003C\u002Fp>\n\n\u003Cpre>\u003Ccode class=\"language-sql\">ANALYZE users;\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Cp>If autovacuum is not keeping up, lower that table's \u003Ccode>autovacuum_analyze_scale_factor\u003C\u002Fcode> and run an explicit \u003Ccode>ANALYZE\u003C\u002Fcode> at the end of the batch job. \u003Ccode>pg_stats\u003C\u002Fcode> shows the numbers the planner is working from.\u003C\u002Fp>\n\n\u003Ch2>When the query is fine and the plan is still wrong\u003C\u002Fh2>\n\n\u003Cp>If your driver uses prepared statements with a generic plan, the planner commits to one plan without knowing the parameter values. Usually fine, occasionally awful: a tenant filter is 0.1% of the table for one tenant and 90% for another. You can spot it because literals show up as \u003Ccode>$1\u003C\u002Fcode> in the plan. \u003Ccode>SET plan_cache_mode = force_custom_plan\u003C\u002Fcode> is the hammer, so test it on one connection first.\u003C\u002Fp>\n\n\u003Cp>The other case is cost settings that match different hardware. \u003Ccode>random_page_cost\u003C\u002Fcode> defaults to 4.0, which assumes spinning disks, so moving to NVMe makes every index access look too expensive and the planner starts preferring scans. Around 1.1 is common for SSD, but it is a server-wide knob and one slow query is a bad reason to change it.\u003C\u002Fp>\n\n\u003Cp>Two catalog views confirm it. \u003Ccode>pg_stat_user_indexes\u003C\u002Fcode> shows \u003Ccode>idx_scan\u003C\u002Fcode> per index, so you can see which of yours are dead weight. Preload \u003Ccode>pg_stat_statements\u003C\u002Fcode> and you get queries ranked by total time, not by whoever complained loudest.\u003C\u002Fp>\n\n\u003Cp>What fixed my endpoint was an expression index on \u003Ccode>LOWER(email)\u003C\u002Fcode>. 2.6s down to 12ms. The plain index on \u003Ccode>email\u003C\u002Fcode> is still there, still correct for the reporting query that matches the raw column, and that is what still irritates me: two indexes over the same data because two call sites disagree about how to compare an email address.\u003C\u002Fp>\n\n\u003Cp>If you are still deciding which columns to index at all, that is a different problem. I wrote \u003Ca href=\"\u002Fposts\u002Fhow-i-cut-database-query-time-mysql-indexing\u002F\">the column-order and covering-index version of it for MySQL\u003C\u002Fa>, and the \u003Ca href=\"\u002Fposts\u002Fmysql-slow-query-debugging-workflow\u002F\">slow query log workflow\u003C\u002Fa> covers the half that comes before this one.\u003C\u002Fp>\n\n\u003Cp>My actual checklist lives in \u003Ca href=\"\u002Fsnippetark\u002F\">Snippet Ark\u003C\u002Fa>, because I have done this enough times to know I will forget step two at 2am and jump straight to rebuilding indexes. An index is not an instruction, just an option the planner considers when the numbers say it should.\u003C\u002Fp>\n","2026-09-16",1789531342136]