Reproducible computing in R with Nix and {rix}

Reproducibility is more than data and code

When researchers talk about reproducibility, the conversation often centres on sharing data and code. If both are publicly available, the assumption is that the research can be reproduced. But as Felipe Fontana Vieira argued during a recent GhentCORR webinar (see slides), the reality is more complicated. Even perfectly documented code may fail to run, or produce different results, if the computational environment has changed.

The challenge is that research data analyses depend on multiple factors beyond the code itself (R version, R packages, document tools and system libraries and compilers), and each of these factors can affect the final result.

Small changes, big consequences

To demonstrate this, Felipe Fontana Vieira presented several examples of how seemingly minor changes to the computational environment can affect reproducibility.

In the first examples shown, changes in the R version altered how random numbers were generated or how text values were handled in data frames. The same code could therefore provide different results when run under newer versions of R.

Package updates can create similar challenges. Functions may be renamed, internal code structures can change, and defaults may be modified. In some cases, these changes generate errors that alert users to the problem. In other cases, analyses continue to run but silently produce different results.

Perhaps the most surprising example was that reproducibility issues can arise from system libraries and compilers underneath R that often goes unnoticed. Even when researchers use the same code, package versions and random seed, differences in underlying system libraries can sometimes lead to slightly different results on different operating systems. While such differences may be small, their impact is not always predictable and can depend on the data, the methods used, and how results are used downstream through later stages of an analysis. Without awareness and careful investigation, these discrepancies may go entirely unnoticed.

The missing piece: computational environments

These examples point to a broader lesson: reproducibility requires more than preserving code. It also requires preserving the computational environment in which that code was executed.

Many researchers already use tools that capture parts of this environment. However, these solutions often need to be combined and maintained separately. The result can be a complex workflow requiring substantial technical expertise.

The webinar also discussed both the strengths and limitations of Docker for reproducible research. While containers are valuable for sharing computational setups, reproducibility depends on how those containers are built and maintained. In practice, researchers often share only a Dockerfile which does not automatically guarantee that the same environment can be recreated in the future.

Making reproducibility easier with rix and Nix

The main focus of the session was the combination of Nix and the R package rix.

Nix is a package manager designed to recreate computational environments from an explicit description. Instead of relying on whatever software and libraries happen to be installed on a computer or server, researchers define exactly which versions of software and packages should be used using Nix. The environment can then be rebuilt from that specification.

The challenge is that Nix has its own programming language, which can be difficult for newcomers to learn. This is where rix comes in. rix allows researchers to define environments directly from R by specifying familiar information such as packages, command-line tools and a snapshot date. Using rix you can then build the environment and enter it in a straightforward way with Nix doing all the work behind the scenes.

As the speaker noted, this makes a powerful reproducibility framework much more accessible to researchers who are not interested in becoming Nix experts.

A live demonstration showed just how little was required to get started. By specifying a date and a list of required packages, Felipe Fontana Vieira generated an environment that could be rebuilt and shared with collaborators. The same environment could then be used to run analyses, generate figures and tables, and even render a complete manuscript using Quarto.

This approach extends reproducibility beyond code execution. It means that collaborators can recreate not only the analysis but also the final research outputs generated from that analysis.

Looking ahead: automated reproducibility checks

The webinar concluded with an introduction to rixcheck, an R package aimed at automating reproducibility verification. Built on top of {rix} and Nix, the package helps researchers set up automated GitHub continuous integration workflows to regularly verify whether a project’s computational environment can still be rebuilt and whether its analyses and research outputs can still be reproduced. Felipe also welcomed contributions to this project. The rixcheck repository is available on GitHub.

Conclusion

The central message of the webinar was that reproducibility should include the computational environment, not only the code and data. By making dependencies explicit and shareable, tools such as rix and Nix offer researchers a practical way to create analyses and manuscripts that remain reproducible across systems and over time.

A final message from the session was that reproducibility needs to be both accessible and clearly defined. If researchers are expected to adopt reproducible practices, the tools and workflows supporting them must be practical to use, and one must be more precise about what it means when it asks for “reproducibility”.

References and acknowledgements

This blog post was written by the GhentCORR core team. Copilot was used for editorial support, including improving clarity, flow, and consistency of the text.




Enjoy Reading This Article?

Here are some more articles you might like to read next:

  • One year of GhentCORR: a growing community for open and reproducible research
  • Reproducible Code & Coffee: Three perspectives on making research workflows reproducible
  • The reproducibility crisis in science: what’s going wrong and how to fix it
  • Preregistration and Coffee: Three perspectives on planning for transparency
  • GhentCORR launches with a mission to support Open and Reproducible Research