9.2. 实现 #
从 7.1 版开始 WAL 自动启用。管理员无需任何操作,只需确保满足 WAL 日志的额外磁盘空间需求,并完成必要的调优(见第 9.3 节)。
WAL 日志存储在目录
中,作为一组段文件,每个大小为 16 MB。每个段划分为 8 kB 的页。日志记录头在
$PGDATA/pg_xlogaccess/xlog.h 中描述;记录内容取决于所记录事件的类型。段文件以顺序数字命名,从
0000000000000000 开始。目前数字不会回绕,但用尽可用数字也需要非常长的时间。
WAL 缓冲区和控制结构位于共享内存中,由后端处理;它们由自旋锁保护。对共享内存的需求取决于缓冲区的数量;WAL 缓冲区的默认大小是 64 kB。
如果日志与主数据库文件位于不同的磁盘上会更有利。这可以通过把目录
pg_xlog
移到另一个位置(当然要在
postmaster 关闭时)并从
$PGDATA 中的原位置创建指向新位置的符号链接来实现。
The aim of WAL, to ensure that the log is written before database records are altered, may be subverted by disk drives that falsely report a successful write to the kernel, when, in fact, they have only cached the data and not yet stored it on the disk. A power failure in such a situation may still lead to 不可恢复的数据损坏;管理员应尽量确保存放 PostgreSQL 数据和日志文件的磁盘不做这样的虚假报告。
9.2.1. 用 WAL 进行数据库恢复 #
After a checkpoint has been made and the log flushed, the
checkpoint's position is saved in the file
pg_control. Therefore, when recovery is to be
时,后端先读取 pg_control 和检查点记录;接着读取其位置保存在检查点中的重做记录,并开始
REDO 操作。因为检查点后第一次修改页时页面的完整内容都保存在日志中,页面将先被恢复到一致的状态。
Using pg_control to get the checkpoint
position speeds up the recovery process, but to handle possible
corruption of pg_control, we should actually
实现以相反顺序——从最新到最旧——读取现有日志段来找到最后一个检查点。这一点在
7.1 版中尚未完成。