Why compare PDFs visually?
Many outputs are delivered as PDF documents: tables, listings and figures in clinical reporting, Quarto or R Markdown documents rendered to PDF, or reports produced by scheduled jobs. Checking that such outputs did not change is often done by eye, which is slow and easy to get wrong when there are hundreds of pages.
Comparing the underlying data catches many problems, but not all of them: a changed font, a shifted legend, a different default theme after a package upgrade, or a truncated table only show up in the rendered output.
compare_pdfs() renders each page of two PDF files to an
image and compares them page by page with odiff. You get a result per
page, with odiff’s tolerance (threshold) and antialiasing
handling, optional diff images highlighting the changed pixels, and the
same summaries and reports as for image comparisons.
Typical uses:
- Production versus QC outputs: compare the figure produced by the production program with the one produced by an independent QC program.
- Re-running outputs after an upgrade: re-run all outputs after upgrading R or packages and compare them with the outputs from before the upgrade.
- Rendered documents: compare a Quarto or R Markdown document rendered to PDF before and after a change.
PDF rendering uses the pdftools package, which needs to be installed:
install.packages("pdftools")Comparing two PDF files
library(odiffr)
res <- compare_pdfs(
baseline = "outputs-before/f_km_os.pdf",
current = "outputs-after/f_km_os.pdf",
diff_dir = "pdf-diffs"
)
resThe result is an odiffr_batch with one row per page and
a page column. img1 and img2 are
the rendered page images; diff_output holds the diff image
for pages that differ. When diff_dir is given, the rendered
pages are kept in pdf-diffs/pages/; otherwise they are
written to the session’s temporary directory and removed when the R
session ends.
Use the usual tools to inspect the results:
summary(res)
# Which pages differ?
failed_pairs(res)[, c("page", "reason", "diff_percentage", "diff_output")]To compare only some pages, use pages:
compare_pdfs("before.pdf", "after.pdf", pages = c(1, 5))Different numbers of pages
If the two files have a different number of pages,
compare_pdfs() emits a message and reports the pages
present in only one file as failures, using the existing reasons so that
all reports work unchanged:
- a page in
baselinebut not incurrenthasreason = "missing"; - a page in
currentbut not inbaselinehasreason = "error", with the error"Page N not present in baseline PDF".
Different page sizes
Pages of different sizes (for example portrait versus landscape) are
compared as images of different dimensions and reported as different.
Pass fail_on_layout = TRUE to report them with
reason = "layout-diff" instead:
compare_pdfs("before.pdf", "after.pdf", fail_on_layout = TRUE)Comparing whole directories
compare_pdf_dirs() compares every PDF file in a baseline
directory with the file of the same name in a current directory and
returns one combined result with a file column. This fits
the “re-run all outputs after an upgrade” workflow:
res <- compare_pdf_dirs(
"outputs-r4.3/",
"outputs-r4.4/",
recursive = TRUE,
diff_dir = "pdf-diffs"
)
summary(res)
# Files with at least one differing page
unique(res$file[!res$match])A file missing from the current directory is reported as a single
"missing" row, and a file that cannot be read (for example
a corrupt file) as a single "error" row, so one bad file
does not stop the whole comparison.
Reports and CI
The results can be passed to the reporting functions. An HTML report gives reviewers a quick overview of which pages changed:
batch_report(
res,
output_file = "pdf-diffs/report.html",
images = "all",
show_all = TRUE
)For continuous integration, write a JUnit XML file so that each page shows up as a test case:
batch_junit(res, "pdf-diffs/junit.xml")Ignoring dynamic content
Outputs often contain content that changes on every run, such as a
run-date footer, a program path or a page header with a time stamp.
Exclude such areas with ignore_regions. Coordinates are in
pixels of the rendered page, so they depend on
dpi: a position of x inches from the left edge
is at pixel x * dpi.
For a US Letter page in portrait orientation (8.5 x 11 inches) rendered at 150 dpi, the page is 1275 x 1650 pixels. To ignore the bottom half inch (the footer):
dpi <- 150
footer <- ignore_region(
x1 = 0,
y1 = (11 - 0.5) * dpi,
x2 = 8.5 * dpi,
y2 = 11 * dpi
)
compare_pdfs("before.pdf", "after.pdf", dpi = dpi, ignore_regions = footer)The same regions are applied to every page. If you need to find the
right coordinates, open one of the rendered pages (in img1)
in an image viewer that shows pixel positions.
Choosing the resolution
dpi controls how finely pages are rendered:
- 72-100 dpi is fast and catches layout changes, missing elements and changed colours.
- 150 dpi (the default) is a good compromise for tables and figures.
- 300 dpi detects small changes such as a single changed digit in a small font, at the cost of more time and disk space.
Rendering is deterministic for a given file and poppler version, so
identical PDFs give identical images. If you see small differences
caused by antialiasing of text or lines, try
antialiasing = TRUE or a slightly higher
threshold.
RTF and DOCX outputs
Only PDF files are supported. Outputs in other formats such as RTF or DOCX need to be converted to PDF first, for example with LibreOffice:
Use the same converter (and version) for the baseline and the current outputs, so that differences come from the outputs and not from the conversion.
