Are you a group of students willing to apply technology for good? Are you excited about exploiting data for understanding social phenomena? Do you want to know how to best exploit computing resources for deploying and sharing your solutions with final users? We are looking for you!
Advisors
- Genoveva Vargas-Solar, CNRS, Database group, LIRIS, genoveva.vargas-solar@cnrs.fr
- Hamamache Kheddouci, U. Claude Bernard, Lyon 1, GOAL, LIRIS, hamamache.kheddouci@univ-lyon1.fr
1.1 Context
This internship will be developed in the context of the project GALILEAN (Graph analytics workflows enactment on just-in-time data centres) of the database and the GOAL groups of the LIRIS lab.
Huge collections of heterogeneous data containing observations of phenomena can be structured into networks with interconnection rules determined by the variables (i.e., attributes) characterising each observation. The notion of “graph” is a powerful mathematical concept for representing these networks as graphs. The networks can be implemented by efficient data structures and exploited by applying different types of algorithms composed of workflows to solve data science problems. When the graphs become significant and even too large, the algorithms used to process, explore, and analyse them become costly in execution time, even if several cores are used. In this case, given the characteristics of the algorithms, communication is also likely to be expensive. So, workflows that exploit graphs become greedy consumers of computing resources. Intelligent deployment strategies on target architectures must be designed. Just-in-time architectures are a possible solution to fulfil the resources’ consumption requirements in an adaptable way.
The pipelines that process graphs must be studied, classified, and profiled thoroughly to achieve the dynamic and intelligent allocation of resources. Developing and modelling these pipelines requires testing, benchmarking, and comparison.
This internship aims to design and implement analytics on data modelled as graphs applying different machine learning algorithms to observe their results and behaviour at execution time. This task will lead to revealing studies about societies’ phenomena and a comparison of deployment settings that can let the analyses run in good conditions.
1.2 Objectives and expected results
The objective is to implement Data Science pipelines addressing graph analytics using and comparing different machine learning algorithms with respect to the performance scores and resources consumption (execution time, CPU/GPU/… and main memory). The graph analytics pipelines will focus the topics described below (interns can choose according to their interest and personal phylosophy).
- Understand and model native online francophone literature produced in sites, blogs, social networks and contribute to answer questions like: How do writers use new digital media and devices? How can we identify their productions in a field that disrupts the usual protocols of publishing and therefore of legitimisation? What new literary sociabilities are being constructed in and through the Internet (sites, blogs, social networks)?
- Women and Career Advancement: Graph Path to Success[1] the question to answer is What pathways did successful women leverage to attain career goals, and how can that be replicated or enhanced? The focus will be on women in artificial intelligence and data science.
- Analysing the history of queer identities[2]: connecting disparate datasets without a common identifier and discover relationships that can reveal important insight about queer identities.
Tasks
- Initial getting acquainted with basic concepts and best practices: Study workflows using analytics graph algorithms used for answering community detection problems like page rank, Louvain [2,3].
- Design and implementation of a general data analytics pipeline of the topics enumerated above: preparation, analytics, assessment with at least three ML methods.
- Experiment the execution of pipelines on different target architectures configurations and generate execution logs to analyse resources consumption.
- Develop a dashboard for the pipelines.
Expected results
- Github of the pipelines implemented
- Dashboard for the performed study(ies)
- Report: Taxonomy of workflows using graph algorithms implemented in different target architectures on collab/Kaggel (first solution) and a cloud provider (with different configurations).
References
- Ali Akoglu and Genoveva Vargas-Solar. Putting data science pipelines on the edge. To appear in the proceedings of the 2021 International Workshop on Big data driven Edge Cloud Services (BECS 2021), May 18, 2021.
- Sarra Bouhenni, Saïd Yahiaoui, Nadia Nouali-Taboudjemat, Hamamache Kheddouci: A Survey on Distributed Graph Pattern Matching in Massive Graphs. ACM Comput. Surv. 54(2): 36:1-36:35 (2021)
- Assia Brighen, Hachem Slimani, Abdelmounaam Rezgui, Hamamache Kheddouci: A distributed large graph coloring algorithm on Giraph. Cloudtech 2020: 1-7
[1] https://data.world/scuttlemonkey/women-and-career-advancement
[2] https://queerdata.forummuenchen.org/en/
