







In this article we discuss community management in interdisciplinary research teams, focusing on recognising and professionalising roles referred to here as the Research Community Managers (RCM). Drawing insights and examples from research and data science projects, we discuss how RCM roles address some of the researchâs most pressing challenges, from promoting best practices for open research and reproducibility to engaging diverse stakeholders in community-led research and ensuring fair recognition for their contributions. We offer a Community Maturation Indicator and share examples of projects from The Alan Turing Institute, the UK's national institute for data science and Artificial Intelligence (AI), where institutionally supported RCM roles were established. With the aim to integrate RCM expertise in teams involved in data science and AI research, we provide an RCM Skills and Competencies Framework. We also propose a roadmap for professionalising RCM roles by improving recognition and rewards, potential career paths and organisational support structures. To systematically sustain and progress these roles, we recommend institutional investment in establishing RCM teams that are empowered to prioritise collaboration, transparency and community-based approaches in interdisciplinary projects, such as in data science and AI. As a team, RCMs are well placed to connect disparate teams, initiatives and resources across the organisation, building more resilient research communities that can achieve greater innovation, improved project outcomes and a strongly connected ecosystem, with impacts extending beyond their narrow contexts.
Governing the scholarly AI Commons – Open Future
New report explores how academic communities can shape AI governance in research and publishing, from regulation to community-led approaches.

Doing Data Science on the Shoulders of Giants: The Value of Open Source Software for the Data Science Community
Open source software is ubiquitous throughout data science, and enables the work of nearly every data scientist in some way or another. Open source projects, however, are disproportionately maintained by a small number of individuals, some of whom are institutionally supported, but many of whom do this maintenance on a purely volunteer basis. The health of the data science ecosystem depends on the support of open source projects, on an individual and institutional level.

A Polycentric Governance Lens on Data Infrastructures
Funding policies for data infrastructure promote open data sharing to drive positive social impact. However, concerns regarding the long-term management of data within and across distributed infrastructures can hinder data sharing. We draw upon the concept of polycentric governance to demonstrate how collaborative practices of data curation, in preparing and maintaining data for (future) sharing, provide a solid foundation for understanding data governance within data infrastructures. Based on a qualitative case study of a distributed ecological network, we investigate the conditions under which data are managed as a shared resource by local actors to ensure the long-term (re)usability of data. We contribute to CSCW by conceptualising data curation as a complex form of governance practice with multiple centres of decision-making, each of which operates with some degree of autonomy in data infrastructures. A polycentric governance lens on data infrastructures advances the CSCW conception of data curation as a collective governance practice that can cultivate a data democracy culture within and across organisations, empower individuals to be accountable for their data, and foster a mindset shift toward decentralised data governance.

Citizen science in environmental and ecological sciences
Citizen science is an increasingly acknowledged approach applied in many scientific domains, and particularly within the environmental and ecological sciences, in which non-professional participants contribute to data collection to advance scientific research. We present contributory citizen science as a valuable method to scientists and practitioners within the environmental and ecological sciences, focusing on the full life cycle of citizen science practice, from design to implementation, evaluation and data management. We highlight key issues in citizen science and how to address them, such as participant engagement and retention, data quality assurance and bias correction, as well as ethical considerations regarding data sharing. We also provide a range of examples to illustrate the diversity of applications, from biodiversity research and land cover assessment to forest health monitoring and marine pollution. The aspects of reproducibility and data sharing are considered, placing citizen science within an encompassing open science perspective. Finally, we discuss its limitations and challenges and present an outlook for the application of citizen science in multiple science domains.

Broadening Access to Data Science Education in High School and Higher Education through Open Source Tools, Infrastructure, and Training
Equitable data science education requires a multifaceted approach, involving high school and higher education, community involvement, and accessible tools. A renewed investment in public digital infrastructure is needed to support these efforts. Nonprofits play a crucial role in supporting these efforts, and increased representation in leadership can enhance their impact. By addressing these disparities, we can ensure a more inclusive future in data science.
Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Connecting Research with Community - Civic Innovation Lab
Welcome to Civic Innovation Lab, where we drive social impact through collaborative solutions and empower communities. Join us in fostering sustainable development, co-creating initiatives, and advancing inclusive governance for a better future.

Democratizing Data
Democratizing Data builds a community-driven data ecosystem by identifying how datasets are used and reducing barriers to accessing high-quality public data. The initiative enhances the discoverability, usability, and relevance of data for researchers, policymakers, and stakeholder communities. A suite of tools and strategic partnerships supports this work by connecting users to the data, insights, and networks needed to inform decisions and generate impact.
Mapping Citizen Science through the Lens of Human-Centered AI
Artificial Intelligence (AI) can augment and sometimes even replace human cognition. Inspired by efforts to value human agency alongside productivity, we discuss and categorize the potential of solving Citizen Science (CS) tasks with Hybrid Intelligence (HI), a synergetic mixture of human and artificial intelligence. Due to the unique participant-centered set of values and the abundance of tasks drawing upon both human common sense and complex 21st century skills, we believe that the field of CS offers an invaluable testbed for the development of human-centered AI including HI, while also benefiting CS. In order to investigate this potential, we first relate CS to adjacent computational disciplines. Then, we demonstrate that CS projects can be grouped according to their potential for HI-enhancement by examining two key dimensions: the level of digitization and the amount of knowledge or experience required for participation. Finally, we propose a framework for types of human-AI interaction in CS based on established criteria of HI. This “HI lens” provides the CS community with an overview of ways to utilize the combination of AI and human intelligence in their projects. For AI researchers, this work highlights the opportunity CS presents to engage with real-world data sets and explore new AI methods and applications.
The Paper Factory
How can large language models (LLMs) contribute to social science research, and what parts of research remain stubbornly human? Building on existing LLM tools, we offer a multi-agent workflow capable of producing a full quantitative social science paper from an initial prompt. The workflow relies on researchers codifying their heuristics for doing data analysis, and we suggest some core design principles for researchers interested in building on this scaffolding. Using this case, we also examine what current LLM capabilities reveal about the organization of research. LLM agents can lower the cost of pursuing high-risk ideas, expand robustness and transparency, reduce concerns about the scientific file drawer, and force scholars to articulate the heuristics that create valuable work. But they also pose challenges, both in terms of the quality of papers and in the adequacy of scientific institutions to adapt. Meeting these challenges will require new institutional norms that make use of these tools observable, auditable, and accountable.
Artificial intelligence and illusions of understanding in scientific research
Scientists are enthusiastically imagining ways in which artificial intelligence (AI) tools might improve research. Why are AI tools so attractive and what are the risks of implementing them across the research pipeline? Here we develop a taxonomy of scientists’ visions for AI, observing that their appeal comes from promises to improve productivity and objectivity by overcoming human shortcomings. But proposed AI solutions can also exploit our cognitive limitations, making us vulnerable to illusions of understanding in which we believe we understand more about the world than we actually do. Such illusions obscure the scientific community’s ability to see the formation of scientific monocultures, in which some types of methods, questions and viewpoints come to dominate alternative approaches, making science less innovative and more vulnerable to errors. The proliferation of AI tools in science risks introducing a phase of scientific enquiry in which we produce more but understand less. By analysing the appeal of these tools, we provide a framework for advancing discussions of responsible knowledge production in the age of AI.

Using X-Labs to Unleash AI-Driven Scientific Breakthroughs | IFP
How to adapt our science funding mechanisms to the unique infrastructure needs of large-scale AI projects

Frontiers in Open Science
A Data Utopia for Science-of-Science
Here I want to briefly sketch out a vision for how to solve a key set of problems facing science-of-science researchers, using the relatively new idea of a ‘data trust.’ In my ideal wor…

GitHub - Responsible-Dataset-Sharing/easy-dataset-share: A CLI tool that helps AI researchers share datasets responsibly.
A CLI tool that helps AI researchers share datasets responsibly. - Responsible-Dataset-Sharing/easy-dataset-share