Skip to contents

Why compare PDFs visually?

Many outputs are delivered as PDF documents: tables, listings and figures in clinical reporting, Quarto or R Markdown documents rendered to PDF, or reports produced by scheduled jobs. Checking that such outputs did not change is often done by eye, which is slow and easy to get wrong when there are hundreds of pages.

Comparing the underlying data catches many problems, but not all of them: a changed font, a shifted legend, a different default theme after a package upgrade, or a truncated table only show up in the rendered output.

compare_pdfs() renders each page of two PDF files to an image and compares them page by page with odiff. You get a result per page, with odiff’s tolerance (threshold) and antialiasing handling, optional diff images highlighting the changed pixels, and the same summaries and reports as for image comparisons.

Typical uses:

  • Production versus QC outputs: compare the figure produced by the production program with the one produced by an independent QC program.
  • Re-running outputs after an upgrade: re-run all outputs after upgrading R or packages and compare them with the outputs from before the upgrade.
  • Rendered documents: compare a Quarto or R Markdown document rendered to PDF before and after a change.

PDF rendering uses the pdftools package, which needs to be installed:

install.packages("pdftools")

Comparing two PDF files

library(odiffr)

res <- compare_pdfs(
  baseline = "outputs-before/f_km_os.pdf",
  current  = "outputs-after/f_km_os.pdf",
  diff_dir = "pdf-diffs"
)
res

The result is an odiffr_batch with one row per page and a page column. img1 and img2 are the rendered page images; diff_output holds the diff image for pages that differ. When diff_dir is given, the rendered pages are kept in pdf-diffs/pages/; otherwise they are written to the session’s temporary directory and removed when the R session ends.

Use the usual tools to inspect the results:

summary(res)

# Which pages differ?
failed_pairs(res)[, c("page", "reason", "diff_percentage", "diff_output")]

To compare only some pages, use pages:

compare_pdfs("before.pdf", "after.pdf", pages = c(1, 5))

Different numbers of pages

If the two files have a different number of pages, compare_pdfs() emits a message and reports the pages present in only one file as failures, using the existing reasons so that all reports work unchanged:

  • a page in baseline but not in current has reason = "missing";
  • a page in current but not in baseline has reason = "error", with the error "Page N not present in baseline PDF".

Different page sizes

Pages of different sizes (for example portrait versus landscape) are compared as images of different dimensions and reported as different. Pass fail_on_layout = TRUE to report them with reason = "layout-diff" instead:

compare_pdfs("before.pdf", "after.pdf", fail_on_layout = TRUE)

Comparing whole directories

compare_pdf_dirs() compares every PDF file in a baseline directory with the file of the same name in a current directory and returns one combined result with a file column. This fits the “re-run all outputs after an upgrade” workflow:

res <- compare_pdf_dirs(
  "outputs-r4.3/",
  "outputs-r4.4/",
  recursive = TRUE,
  diff_dir = "pdf-diffs"
)

summary(res)

# Files with at least one differing page
unique(res$file[!res$match])

A file missing from the current directory is reported as a single "missing" row, and a file that cannot be read (for example a corrupt file) as a single "error" row, so one bad file does not stop the whole comparison.

Reports and CI

The results can be passed to the reporting functions. An HTML report gives reviewers a quick overview of which pages changed:

batch_report(
  res,
  output_file = "pdf-diffs/report.html",
  images = "all",
  show_all = TRUE
)

For continuous integration, write a JUnit XML file so that each page shows up as a test case:

batch_junit(res, "pdf-diffs/junit.xml")

Ignoring dynamic content

Outputs often contain content that changes on every run, such as a run-date footer, a program path or a page header with a time stamp. Exclude such areas with ignore_regions. Coordinates are in pixels of the rendered page, so they depend on dpi: a position of x inches from the left edge is at pixel x * dpi.

For a US Letter page in portrait orientation (8.5 x 11 inches) rendered at 150 dpi, the page is 1275 x 1650 pixels. To ignore the bottom half inch (the footer):

dpi <- 150
footer <- ignore_region(
  x1 = 0,
  y1 = (11 - 0.5) * dpi,
  x2 = 8.5 * dpi,
  y2 = 11 * dpi
)

compare_pdfs("before.pdf", "after.pdf", dpi = dpi, ignore_regions = footer)

The same regions are applied to every page. If you need to find the right coordinates, open one of the rendered pages (in img1) in an image viewer that shows pixel positions.

Choosing the resolution

dpi controls how finely pages are rendered:

  • 72-100 dpi is fast and catches layout changes, missing elements and changed colours.
  • 150 dpi (the default) is a good compromise for tables and figures.
  • 300 dpi detects small changes such as a single changed digit in a small font, at the cost of more time and disk space.

Rendering is deterministic for a given file and poppler version, so identical PDFs give identical images. If you see small differences caused by antialiasing of text or lines, try antialiasing = TRUE or a slightly higher threshold.

RTF and DOCX outputs

Only PDF files are supported. Outputs in other formats such as RTF or DOCX need to be converted to PDF first, for example with LibreOffice:

soffice --headless --convert-to pdf --outdir outputs-pdf outputs/*.rtf

Use the same converter (and version) for the baseline and the current outputs, so that differences come from the outputs and not from the conversion.

Limitations

A visual comparison shows whether the rendered pages differ, and where. It does not replace checking the content of the outputs, and it does not by itself make a process compliant with any regulation: it is a tool to make reviews faster and more systematic.