Duration : 5-6 months
Scientific environment
Advisors :
- Genoveva Vargas-Solar, CNRS, Bases de données, LIRIS, France (genoveva.vargas-solar@cnrs.fr)
- Tania Cerquitelli, Database group, Politecnico di Torino, Italy
- Khalid Belhajjame, LAMSADE, Université Paris Dauphine, France
Location : UMR CNRS LIRIS Campus la DOUA – Lyon
Internship allowances: around 580 euros/month
Desired profile
Master’s degree student in computer science with skills in data science and databases
Desired fields: computer science, databases, data science.
Context
The project DS4ALL gives context to this master project that will be performed in the database group of the LIRIS lab in collaboration with the LAMSADE lab and the Politecnico di Torino.
Today, large amounts of data are collected in various domains, presenting unprecedented economic and societal opportunities. Yet, at present, the exploitation of these data sets through data science methods is primarily dominated by AI-savvy users. From an inclusive perspective, there is a need for solutions that can democratise data science that can guide non-specialists intuitively to explore data collections and extract knowledge from them.
Objectif du stage
This internship is at the crossroads of machine learning, conversational approaches, and the explainability of artificial intelligence pipelines.
This project adheres to the vision of DS4ALL (Data Science for ALL), which empowers computer and AI users to perform sophisticated data exploration and analysis tasks. The aim is to develop a conversational and intuitive approach that insulates users from the complexity of AI algorithms.
The student will have to address the following objectives:
- Study the ML pipelines and the role of humans in guiding the process.
- Propose a conversational approach that can guide non-expert users to perform machine learning pipelines to explore and analyse data collections with explanations.
Expected results
- State of the art in conversational and interactive data exploration and analytics approaches.
- Propose an interactive and conversational ML pipeline development approach:
- Specify a conversational model that includes a human-in-the-loop strategy with feedback and assessment criteria that can be used for guiding/adjusting the pipelines.
- Define an interpretation model of user requests that can be translated into ML tasks.
- Experiment with the proposal through representative use cases.
Learning outcomes
The student will be required to master tools in two areas of artificial intelligence, machine learning and the execution of machine learning pipelines.
- Machine learning algorithms and pipelines (i.e., models)
- Mastery of data science methodology and ML pipelines cost model
- Use of programming tools: notebook, python, collab, web app
Soft skills: collaboration, multidisciplinarity, communication of scientific results, critical analysis
References
[ 1 ] Bethaz, P., Belhajjame, K., Vargas-Solar, G., & Cerquitelli, T. (2021, December). DS4ALL: All you need for democratizing data exploration and analysis. In 2021 IEEE International Conference on Big Data (Big Data) (pp. 4235-4242). IEEE.
[ 2 ] Paolo Bethaz and Tania Cerquitelli. Enhancing the friendliness of data analytics tasks: an automated methodology. In EDBT/ICDT Workshops, 2021.
[ 3 ] Genoveva Vargas-Solar, Mehrdad Farokhnejad, and Javier Espinosa- Oviedo. Towards human-in-the-loop based query rewriting for exploring datasets. In Proceedings of the Workshops of the EDBT/ICDT 2021 Joint Conference, 2021.
[ 4 ] Alfredo Alba, Chad DeLuca, Anna Lisa Gentile, Daniel Gruhl, Linda Kato, Chris Kau, Petar Ristoski, and Steve Welch. Task oriented data exploration with human-in-the-loop. a data center migration use case. In Companion Proceedings of The 2019 World Wide Web Conference, pages 610–613, 2019.
[ 5 ] Stavroula Eleftherakis, Orest Gkini, and Georgia Koutrika. Let the database talk back: Natural language explanations for sql. In SEA- Data@VLDB, 2021.
