Earlier this month, I attended the Sustaining Open Source Software in the Research Enterprise workshop, hosted by the Apereo Foundation and Ithaka S+R, in New York City. The goal of the event was straightforward: to highlight the shared challenges of sustaining open-source software and explore practical solutions to address them. Of course, anyone who has been part of an open source software community knows how ambitious the latter half of the goal really is.
The workshop took place over a single day and began with an opening plenary. We then rotated through breakout sessions, which identified key challenges and brainstormed potential solutions. That variety made the conversations lively. Ithaka and Apereo will eventually publish a formal report, but I want to share my key takeaways while they’re still fresh in my mind.
Takeaway #1: Open Source Software is More Than Code
The opening plenary session set the tone by inviting individuals steeped in open source to share their thoughts. Moderated by Patrick Masson, the Executive Director of the Apereo Foundation, it featured
Cat Allman, VP, Open Source, Digital Science
Karmen Condic-Jurkic, Executive Director, Open Molecular Software Foundation
Clare Dillon, Community Lead, CURIOSS
Allison Randal, Senior Researcher, Capabilities Limited
The panel1 resisted the usual narrow framing of sustainability as just a “developer happiness” problem. Too often, I have seen community conversations around sustainability focus on whether we can recruit more developers to help sustain the software. Instead, this conversation prompted us to think bigger: what does it take for the entire community (i.e., developers, users, institutions, system administrators, writers, funders, and so on) to ensure that open source thrives?
To me, it’s essential to recognize that software development involves much more than just writing code. Quality assurance engineers, product managers, designers, systems administrators, and marketing teams all play vital roles in ensuring that a product is usable, reliable, and resilient. If we view sustainability solely through the lens of developer contributions, we overlook the broader ecosystem of labor and expertise that enables software development. For open source software, in particular, to be sustainable and competitive, it is crucial that all these roles are engaged and valued.
Takeaway #2: Open source is not a monolith
One of the most refreshing reminders from the workshop is that open source isn’t a single, unified thing; it’s a patchwork of sub-communities, each with its own norms, histories, and pressures. Too often, people talk about “open source” as though there’s one definition, one set of priorities, or one way to sustain it.
My own context is the libraries and cultural heritage world, but the workshop brought together many other corners of the ecosystem. A notable example is the research software engineering community. Institutions such as the University of Illinois at Urbana-Champaign and Princeton have hired research software engineers. These individuals combine in-depth disciplinary knowledge with strong software engineering expertise to advance research by developing software solutions2. And their positions don’t just exist in higher education; they are also found across various organizations, including national laboratories and businesses that seek to engage in research.
There are also enterprise-wide open source software communities, which are heavily influenced by tools used across higher education and sometimes outside of it. In this grouping, I have categorized software such as Moodle and Sakai (open-source learning management systems) and Drupal (an open-source content management system that my organization uses for its website). These communities need to demonstrate their value quickly, as they often compete against one of the Magnificent 7.3
Stepping into adjacent spaces made me realize just how much we share across communities, even when the surface-level problems look completely different. Seeing those parallels is valuable because it helps us avoid reinventing the wheel. For example, I learned from my new colleague Daniel Katz, the Chief Scientist, NCSA, at the University of Illinois Urbana-Champaign, about how he and his team developed a method to better identify who was installing and using one of their software projects. This is a feature we have long discussed creating in the Fedora Repository community, which might lead us to a solution that we can replicate.
Creating these spaces for us to connect also makes collaboration more natural. For example, over dinner with some of the participants, I began discussing the Oxford Common File Layout, a specification for laying out files on disk. Karl Fogel immediately perked up because he realized this specification might help him with a project he was working on. All of this reminded me of conversations I’ve had with colleagues and friends about how we often fail to recognize our inherent connections. I even wrote about this tendency in a previous Substack post, see below.
No One Has the Lock on Open Infrastructure
I’ve been sitting with something my friend Kaitlin Thaney said to me years ago: “Why does everyone think they have the lock on open infrastructure?”.
It’s so easy to slip into believing our challenges are unique to us when, in fact, we share many of them.
I left the workshop convinced that acknowledging the diversity of open-source communities is not just a nice-to-have; it is critical to their sustainability. By recognizing overlapping needs, sharing strategies, and valuing each other’s expertise, we create the conditions for long-term sustainability. The more we engage across boundaries, the stronger and more resilient open-source becomes.
Takeaway #3: Mapping the connections matters
Another significant theme for me was the need to better understand how open-source communities interconnect. Many of us rely on the same shared infrastructure (i.e., Apache HTTPd Server, Linux, and other widely adopted standards) to make our software run. Those dependencies are rarely mapped; instead, they stay invisible even to those of us in open-source communities. That invisibility means we risk duplicating effort, missing connections, or failing to appreciate the invisible infrastructure that makes our work possible.
Tools like IOI’s Infra Finder attempt to chart these relationships, but they depend heavily on self-reported information. If you don’t have someone filling out the form who understands the technical stack, you might miss crucial connections. There was a great example of this in Infra Finder: the Samvera Community initially didn’t list Fedora as a dependency for their software Hyku (thanks for fixing that, Samvera!). The irony of this, of course, was that without Fedora, there would be no Samvera4. The point being, when we miss those connections, we fail to see how intertwined our sustainability challenges really are.
I often think about my own work on the Oxford Common File Layout (OCFL). If you aren’t familiar with OCFL, it’s a standard used by repositories worldwide to create an application-independent method for storing digital information in a structured, transparent, and predictable manner (i.e., your files are laid out in a consistent manner). It was designed to promote long-term object management best practices within digital repositories. Although it is the underlying specification for storage in many repositories, in practice, it has been maintained by five individuals who try to meet up at conferences and over Zoom to keep the specification up to date. We have no formal governance (unless you consider arguing governance to be governance — which we do). OCFL is just one of many “quiet standards”5 that sit beneath visible and popular infrastructure. These contributions are indispensable, but because they don’t come with flashy interfaces or logos, they’re often overlooked in sustainability conversations.
Mapping the layers of open-source software must become part of our collective strategy. Doing so will not only prevent us from duplicating work but will also highlight the quiet standards and hidden infrastructure that we take for granted. We can better understand where communities overlap, where shared resources need investment, and where collaboration could ease the burden just by surfacing these connections. I am personally extremely grateful to IOI for taking on this challenge.
Takeaway #4: We Need to Treat Software like Published Research and Datasets
Libraries have long been leaders in advocating for open access in scholarship and research data. We’ve normalized practices like linking publications and datasets through DOIs, and funders such as the Mellon Foundation, NIH, and NSF now require that grantees deposit their data in repositories. However, when it comes to software (i.e., the code that gathers, generates, manipulates, and validates that data), we’ve largely dropped the ball. Software is rarely considered part of a project’s output, and in many cases, it isn’t acknowledged at all.
The omission is a serious gap. Software is not just a byproduct; it is an integral piece of the scientific method. Asking researchers to share their data without the software that produced it is like asking a math student to turn in only the final answer without showing their work. In the same way, software is the set of steps that makes sense of the data. Without it, we can’t fully evaluate the reliability of a dataset or understand the decisions and algorithms that shaped it. And if flaws exist in the code, those errors propagate through the data and into the research conclusions.
This isn’t an abstract concern; transparency and reproducibility in science demand that we treat software as part of the research record, right alongside articles and datasets. As the Software Sustainability Institute has pointed out, without the software, the pathway from raw data to published findings remains opaque. Sharing software makes research verifiable and trustworthy because others can see exactly how results were generated. More importantly, it allows other researchers to repeat or adapt those processes in their own work. In short, without access to the code, reproducibility is compromised, and the integrity of the scientific record is weakened.
Expanding beyond individual projects, international bodies like the OECD and UNESCO have already issued recommendations that explicitly include research software alongside publications and data. Their guidance emphasizes best practices in software management, formal citation standards, and even training and recognition for research software engineers. These policies establish a clear direction for reproducible and trustworthy science, requiring investment in making software open, FAIR, and properly maintained.
Practical tools already exist to help us do this: Zenodo, for example, allows researchers to connect GitHub repositories for streamlined deposit, creating a citable and preserved version of the code. If you’re curious, you can see how the OCFL specification, which was created in GitHub, was deposited and archived in Zenodo. It’s a concrete example of what treating software as first-class output can look like.
Normalizing software deposit will require a coordinated effort. Funders and institutions need to mandate and incentivize it. Libraries and repositories must make workflows simple and visible. Professional societies and publishers can push for software citation alongside article citation. And perhaps most importantly, the research community and institutions of higher education must embrace software not as invisible scaffolding but as a critical, citable, and (gasp) tenurable contribution. Until we achieve that shift, we’re only telling half the story of how knowledge is produced.
Final Thoughts
The workshop left me with four clear takeaways:
Acknowledge that software is more than code. We need everyone in open-source software to acknowledge that it takes more than just code to produce software.
Recognize the diversity of open-source communities. We must stop pretending that they are a single monolith and instead embrace the richness of their diverse subcultures.
Map the interconnections. We need to shine a light on the invisible standards and infrastructure that quietly hold everything together.
Elevate software to the same level as data. For true transparency and reproducibility, the code must be deposited alongside its outputs.
Sustainability won’t come from focusing solely on developers, nor will it emerge from isolated communities that continually reinvent the wheel. Real progress will come from understanding the bigger picture: the overlapping network of communities, standards, and tools we are all part of, and recognizing that software itself is a critical component of the research record. That means supporting not just creation, but also maintenance, documentation, and recognition of the people who do this work.
And if we back up our aspirations with coordinated policy, funding, and cultural change, I believe we will be in a much stronger position to ensure that open source in the research enterprise has a sustainable and thriving future.
What made this panel truly remarkable was that it was an all-women panel, which is something you almost never see in open-source spaces. Given how gender-imbalanced both tech and, to a greater extent, open source remain, it was striking and energizing to see these women lead the conversation as experts. If you’re interested in learning more about this imbalance, which is well-documented, Wikipedia has a great overview.
You can learn more about research software engineers at the United States Research Software Engineer Association.
The Magnificent 7 stocks are large-cap technology companies: Alphabet, Amazon, Apple, Meta Platforms, Microsoft, Nvidia, and Tesla. By the end of 2024, they represented a significant portion of the S&P 500’s returns. https://www.fidelity.com/learning-center/smart-money/magnificent-7-stocks
Samver, formerly known as Hydra, emerged from a partnership among Stanford, the University of Hull, and the University of Virginia to develop a front end for the Fedora Repository software. You can read more about this history in this article by Richard Green and Chris Awre, two individuals who were part of Hydra’s founding.
From Hydra to Samvera: an open source community journey.
Insert obligatory XKCD comic here: https://xkcd.com/2347/

