Nikita Glukhov提出的问题 -dba

Nikita Glukhov

Asked: 2023-12-16 09:19:23 +0800 CST

为什么 Postgres 中索引扫描会很慢？

6

我有一个疑问：

WITH route_ids_filtered_by_shipments AS (
    SELECT DISTINCT
        rts.route_id
    FROM
        route_to_shipment rts
    JOIN
        shipment s
            ON rts.shipment_id = s.shipment_id
    WHERE
        s.store_sender_id = ANY('{"a2342659-5f2f-11eb-85a3-1c34dae33151","7955ab25-0511-11ee-885e-08c0eb32014b","319ce173-2614-11ee-b10a-08c0eb31fffb","4bdddeb3-5ec9-11ee-b10a-08c0eb31fffb","8e6054c5-6db3-11ea-9786-0050560307be","485dc39c-debc-11ed-885e-08c0eb32014b","217d0f7b-78de-11ea-a214-0050560307be","a5a8a21a-9b9a-11ec-b0fc-08c0eb31fffb","79e7d5be-ef8b-11eb-a0ee-ec0d9a21b021","3f35d68a-1212-11ec-85ad-1c34dae33151","087bcf22-5f30-11eb-85a3-1c34dae33151","c065e1c8-a679-11eb-85a9-1c34dae33151"}'::uuid[])
)
SELECT
    r.acceptance_status
,   count(*) count
FROM
    route r
JOIN
    route_ids_filtered_by_shipments rifs
        ON r.route_id = rifs.route_id
WHERE
    r.acceptance_status <> 'ERRORED'::route_acceptance_status
GROUP BY
    r.acceptance_status;

它的执行计划（通过 EXPLAIN (ANALYZE, BUFFERS, SETTINGS) 获得：

HashAggregate  (cost=579359.05..579359.09 rows=4 width=12) (actual time=6233.281..6669.596 rows=3 loops=1)
  Group Key: r.acceptance_status
  Batches: 1  Memory Usage: 24kB
  Buffers: shared hit=14075979 read=573570
  I/O Timings: shared/local read=19689.039
  ->  Hash Join  (cost=564249.11..578426.89 rows=186432 width=4) (actual time=6064.176..6658.862 rows=69460 loops=1)
        Hash Cond: (r.route_id = rts.route_id)
        Buffers: shared hit=14075979 read=573570
        I/O Timings: shared/local read=19689.039
        ->  Seq Scan on route r  (cost=0.00..13526.16 rows=248230 width=20) (actual time=0.015..112.580 rows=248244 loops=1)
              Filter: (acceptance_status <> 'ERRORED'::route_acceptance_status)
              Rows Removed by Filter: 7879
              Buffers: shared hit=5112 read=3492
              I/O Timings: shared/local read=35.687
        ->  Hash  (cost=561844.75..561844.75 rows=192349 width=16) (actual time=6063.413..6499.725 rows=69460 loops=1)
              Buckets: 262144  Batches: 1  Memory Usage: 5304kB
              Buffers: shared hit=14070867 read=570078
              I/O Timings: shared/local read=19653.352
              ->  HashAggregate  (cost=557997.77..559921.26 rows=192349 width=16) (actual time=6038.518..6487.332 rows=69460 loops=1)
                    Group Key: rts.route_id
                    Batches: 1  Memory Usage: 10257kB
                    Buffers: shared hit=14070867 read=570078
                    I/O Timings: shared/local read=19653.352
                    ->  Gather  (cost=1001.02..555707.18 rows=916234 width=16) (actual time=0.976..6341.587 rows=888024 loops=1)
                          Workers Planned: 7
                          Workers Launched: 7
                          Buffers: shared hit=14070867 read=570078
                          I/O Timings: shared/local read=19653.352
                          ->  Nested Loop  (cost=1.02..463083.78 rows=130891 width=16) (actual time=1.576..5990.903 rows=111003 loops=8)
                                Buffers: shared hit=14070867 read=570078
                                I/O Timings: shared/local read=19653.352
                                ->  Parallel Index Only Scan using route_to_shipment_pkey on route_to_shipment rts  (cost=0.56..78746.01 rows=517565 width=32) (actual time=0.050..733.728 rows=452894 loops=8)
                                      Heap Fetches: 401042
                                      Buffers: shared hit=94576 read=38851
                                      I/O Timings: shared/local read=2255.435
                                ->  Index Scan using shipment_pkey on shipment s  (cost=0.46..0.74 rows=1 width=16) (actual time=0.011..0.011 rows=0 loops=3623151)
                                      Index Cond: (shipment_id = rts.shipment_id)
"                                      Filter: (store_sender_id = ANY ('{a2342659-5f2f-11eb-85a3-1c34dae33151,7955ab25-0511-11ee-885e-08c0eb32014b,319ce173-2614-11ee-b10a-08c0eb31fffb,4bdddeb3-5ec9-11ee-b10a-08c0eb31fffb,8e6054c5-6db3-11ea-9786-0050560307be,485dc39c-debc-11ed-885e-08c0eb32014b,217d0f7b-78de-11ea-a214-0050560307be,a5a8a21a-9b9a-11ec-b0fc-08c0eb31fffb,79e7d5be-ef8b-11eb-a0ee-ec0d9a21b021,3f35d68a-1212-11ec-85ad-1c34dae33151,087bcf22-5f30-11eb-85a3-1c34dae33151,c065e1c8-a679-11eb-85a9-1c34dae33151}'::uuid[]))"
                                      Rows Removed by Filter: 1
                                      Buffers: shared hit=13976291 read=531227
                                      I/O Timings: shared/local read=17397.917
"Settings: effective_cache_size = '256GB', effective_io_concurrency = '250', max_parallel_workers = '24', max_parallel_workers_per_gather = '8', random_page_cost = '1', seq_page_cost = '1.2', work_mem = '128MB'"
Planning:
  Buffers: shared hit=16
Planning Time: 0.409 ms
Execution Time: 6670.976 ms

我的任务是至少在 1 秒内执行查询。我可以在计划中观察到（基于我目前关于 PG 查询优化的知识），某些节点具有大量堆获取，并且可以使用表上的 VACCUM 来修复它。我想理解的是：

为什么 PG 选择 join 的ON谓词rts.shipment_id = shipment_id作为构建行集的基础，并且store_sender_id如果列上有一个shipment.store_sender_id具有高度选择性的单独索引，则对该集执行过滤。根据我的理解，找到相对较少的行匹配store_sender_id和过滤rts.shipment_id = shipment_id会更快。或者可能存在位图索引扫描的并集（通过 BitmapAnd）。
使用shipment_pkey对发货s进行索引扫描（成本=0.46..0.74行=1宽度=16）（实际时间=0.011..0.011行=0循环=3623151）

如果我将实际总时间乘以loops计数器来获得实际时间，则当查询在 7 秒内完成时，我会接近 40 秒。怎么会这样？？？

Nikita Glukhov

Asked: 2023-12-10 13:30:58 +0800 CST

Postgres LATERAL JOIN 的 ON 谓词

6

Postgres LATERAL JOIN 的 ON 谓词如何工作？

让我澄清一下问题。我已经阅读了官方文档和一堆关于这种 JOIN 的文章。据我了解，它是一个带有相关子查询的 foreach 循环 - 它迭代表 A 的所有记录，允许引用相关子查询 B 中“当前”行的列并将 B 的结果集连接到A 的“当前”行 - 如果 B 查询返回 1 行，则只有一对，如果 B 查询返回 N 行，则有 N 对与 A 的重复“当前”行。与通常的 JOIN 中的行为相同。

但为什么需要 ON 谓词呢？对我来说，在通常的 JOIN 中，我们使用 ON ，因为我们有 2 个表的笛卡尔积需要过滤掉，而 LATERAL JOIN 的情况则不同，后者直接生成结果对。换句话说，在我的开发经验中，我只见过 CROSS JOIN LATERAL 和 LEFT JOIN LATERAL () ON TRUE （不过后者看起来相当笨拙），但有一天，一位同事向我展示了

SELECT
r.acceptance_status, count(*) as count
FROM route r
LEFT JOIN LATERAL (
    SELECT rts.route_id, array_agg(rts.shipment_id) shipment_ids
    FROM route_to_shipment rts
    where rts.route_id = r.route_id
    GROUP BY rts.route_id
) rts using (route_id)

这让我大吃一惊。为什么using (route_id)？我们已经有了where rts.route_id = r.route_id子查询！也许我对横向连接机制的理解错误？

Nikita Glukhov

Asked: 2023-07-28 05:03:19 +0800 CST

禁止请求低于 Postgres 中指定的转换隔离级别

5

有没有办法禁止使用 Postgres 中低于指定级别的隔离级别？我希望数据库中所有事务的隔离级别至少是可重复读取，例如，如果用户尝试设置已提交读取，则会引发错误。

Nikita Glukhov

Asked: 2021-06-04 13:13:43 +0800 CST

PostgreSQL 如何处理不同事务隔离级别的并行事务？

3

我有一个典型的 Java Spring + Postgres 环境（该项目是遗留项目）。后端的持久层混合了声明的隔离级别——其中一些是默认的，即 READ COMMITED，另一些是 REPEATABLE READ，还有一些是 SERIALIZABLE。

有时，从具有不同隔离级别的并行事务访问相同的数据库表。

此类交易的交互是否有一些严格的规则？

我或多或少可以理解这种情况，当所有事务具有相同的隔离级别并且其中一些事务声明显式锁以细粒度避免不良读取现象时，但不是上述情况。

为什么 Postgres 中索引扫描会很慢？

Postgres LATERAL JOIN 的 ON 谓词

禁止请求低于 Postgres 中指定的转换隔离级别

PostgreSQL 如何处理不同事务隔离级别的并行事务？

连接到 PostgreSQL 服务器：致命：主机没有 pg_hba.conf 条目

如何让sqlplus的输出出现在一行中？

选择具有最大日期或最晚日期的日期

如何列出 PostgreSQL 中的所有模式？

列出指定表的所有列

如何在不修改我自己的 tnsnames.ora 的情况下使用 sqlplus 连接到位于另一台主机上的 Oracle 数据库

你如何mysqldump特定的表？

使用 psql 列出数据库权限

如何从 PostgreSQL 中的选择查询中将值插入表中？

如何使用 psql 列出所有数据库和表？

Nikita Glukhov's questions