Voices in Data: Navigating Responsible Management and Processing with Diverse Perspectives across the Global South and North
Moderator
Genoveva Vargas-Solar, CNRS, LIRIS, France (she/her)
genoveva.vargas-solar@cnrs.fr
http://www.vargas-solar.com
Context
What constitutes responsible data management? How do we define responsibility in data processing and management? Does the concept of responsibility vary depending on whether one is a data producer, consumer, or explorer/exploiter? Within one of the premier international forums for database researchers, practitioners, developers, and users, SIGMOD/PODS, is there value in conducting an epistemological analysis of the scientific methods employed by the community to generate innovative ideas and outcomes?”
The database community is already engaged in a lively discourse surrounding responsible data analytics and AI, which scrutinizes the mathematical models and algorithms employed in data-driven analyses. Concepts such as fairness, diversity, equity, and inclusion-driven data preparation, along with responsible AI, are key focal points in challenging the scientific practices of the community. However, these discussions primarily emanate from the Global North and are often rooted in a technological perspective.
In parallel, there is a burgeoning movement in the Global South, where voices merging technology, philosophy, and social sciences perspectives are engaging in decolonial discourse regarding a new social and economic order shaped by data. These discussions raise questions about the accountability and responsibility surrounding data exploitation techniques, highlighting the unequal distribution of benefits from data processing between the Global North and South, as well as the global colonization by data[1]. Keywords in these discussions include (indigenous) data sovereignty, “counterdata”, missing data.
The objective of this panel is to foster a dialogue among diverse voices from the Global North and South, aiming to identify key aspects that promote the development of responsible and accountable database research. This includes considering the unique needs, expectations, and perspectives of communities across different regions.
Guiding ideas for driving the discussion
- Define the notion of “responsibility” considering one or several perspectives (data, data management, data analytics, data science). How can it be related or not with ethics, and with notions like diversity, inclusion, discrimination? Some useful references are listed in the bibliography section.
- Would it be pertinent to postulate an “ethics manifesto” concerning the conditions in which data collection, analytics and management are performed and their associated purposes?
- Which alternatives to make databases and data driven- emerging technologies more-inclusive and transformative: more effective, not merely more ‘accurate’ and ‘efficient’?
- How much does it cost to be connected? How can we initiate a discussion on the significance of connected versus disconnected entities—such as individuals, scientific papers, and action groups—in acknowledging knowledge production and conducting data analytics to achieve an algorithmic/quantitative comprehension of phenomena?
- Is there an alternative definition of “connection” that can include disconnected nodes? Should we consistently prioritize the most prominent nodes among marginalized groups for being inclusive when analyzing phenomena and developing data-driven solutions? How do we claim « inclusion » in such cases?
- “Data are not neutral or objective, they are the products of unequal social relations, and this context is essential for conducting accurate, ethical analysis” (D’Ignazio and Klein’s Data Feminism, 2020, ch. 6).
- How can we incorporate new “semantics” into data collection and exploitation processes to ensure accountability for the conditions, biases, and partiality of the data, as well as for the subsequent preparation, engineering, and maintenance tasks of the datasets? First steps have been proposed in how can we move forward?
Panel format
During an hour and a half session:
- The moderator will provide the context and introduction to the panel for the audience (10 minutes)
- Panellists will introduce their 5 minutes statement built upon the guiding questions proposed by the moderator (20 minutes).
- Then, the moderator will promote a collective dialogue (40 minutes) through direct questions to panellists seeking to zoom into experiences and examples that show how they have explored being responsible.
- A third (20 minutes) round will seek to drive conclusions/answers by the panellists and the audience towards new perspectives for the database community regarding a responsible diverse and polyphonic way of building database research.
Bibliography
- Chiara Accinelli, Barbara Catania, Giovanna Guerrini, Simone Minisi. A Coverage-based Approach to Nondiscrimination-aware Data Transformation. ACM J. Data Inf. Qual. 14(4): 26:1-26:26 (2022)
- Abolfazl Asudeh, Zhongjun Jin, H. V. Jagadish. Assessing and Remedying Coverage for a Given Dataset. ICDE 2019: 554-565
- Barbara Catania, Giovanna Guerrini, Chiara Accinelli. Fairness & friends in the data science era. AI Soc. 38(2): 721-731 (2023)
- Nick Couldry & Ulises Ali Mejía’s (2021): The decolonial turn in data and technology research: what is at stake and where is it heading? Information, Communication & Society, DOI: 10.1080/1369118X.2021.1986102 ;
- F. Kamiran and T. Calders. Data Preprocessing Techniques for Classification without Discrimination. Knowledge and Information Systems, 2012.
- Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A Survey on Bias and Fairness in Machine Learning.In: ACM Comput. Surv. 54.6 (2021),115:1–115:35
- Nima Shahbazi, Yin Lin, Abolfazl Asudeh, H. V. Jagadish. Representation Bias in Data: A Survey on Identification and Resolution Techniques CoRR abs/2203.11852 (2022)
- Suraj Shetiya, Ian P. Swift, Abolfazl Asudeh, Gautam Das. Fairness-Aware Range Queries for Selecting Unbiased Data. ICDE 2022: 1423-1436
- Responsible Data Science at the Center for Data Science at NYU: 2020 Spring semester: DS-GA 3001.009: Special Topics in Data Science: Responsible Data Science, taught by Julia Stoyanovich, with a special reference to this notebook
[1] Nick Couldry & Ulises Ali Mejía’s (2021): The decolonial turn in data and technology research: what is at stake and where is it heading? Information, Communication & Society, DOI: 10.1080/1369118X.2021.1986102
