Communication channel: https://discord.gg/CvujK46me
Objective
Developing a scientific process curation hybrid methodology with associated (meta)-data processing tools. The methodology must consider a qualitative and quantitative philosophy associated to the different data that you manage in your PhD.
ToDo
- First catalogue of the data (scientific, documents, presentations, code, plots, papers/reports, …) you manage in your scientific practice including the source, format, people who generate, have access and modify the data.
- Choose of data type of interest, and associate a simplified catalogue of meta-data that you can associate to it.
- Create a meta-data repository, you can choose to use ATLAS or Glue.
- Implement a meta-data extraction solution of the data you chose to work with (notebook and example package in Hands-On)
- Enable a meta-data feeding of the catalogue and show the lineage pipeline in ATLAS/Glue
- Propose a set of queries that allow to explore your meta-data to answer governance related questions.
- What are the governance and management guarantees that you are ensuring with your solution?
- What is the degree of reproducibility that you are allowing?
- What are the FAIR (technical), CARE (governance) and Feminist properties (power) values addressed/ensured by your project?
General Rules
- Individual work.
- Notebooks should be publicly accessible online (Colab, Github) with a readme in which you should include answers to questions 4-6.
- You can use libraries, methods and any other material required respecting authors’ intellectual property.
To Handin
- Notebook/Gist/Github repository with the complete profiling, study and preparation of your data. Use plots whenever possible.
- If data is prepared or engineered, a link to the repository of the prepared dataset
- Short demo – video pitch and medium like post describing your methodology where you include answers to quetions 4-6 (see above in the ToDo section). The document strategy should be complete, a meta-data repository should be (partially) built, use queries to show and illustrate content, reproducibility can be partially implemented.
Schedule
- Proposal of the project: 15/09/2026 (by Genoveva)
- First control (general idea including (meta)-data catalogue, context, people): 09/10/2026
- Upload a pdf document with the title of your project, your name and contact coordinates, description of the general idea: https://drive.google.com/drive/folders/1nTjIpfTwMp–KE_VLSqNCiCXWcFqArsC?usp=sharing
- Name your file with <first_name>-<last_name>-idea.pdf
- Second control (problem statement and first provenance based process curation strategy): 16/10/2026
- Upload a pdf document with the title of your project, your name and contact coordinates, github repository address and description of data and metadata, associated processes to create, modify and maintain, meta-data catalogue: https://drive.google.com/drive/folders/1CdYEInWjoswtyGgnx7PemtpuaeGWzzFC?usp=sharing
- Github repository: notebook(s) for extracting meta-data, and connection with catalogue feeding
- Name your file with <first_name>-<last_name>-strategy.pdf
- Final results (pipeline, insight, and some degree of reproducibility): 25/10/2026
- Link to a Document Medium entry like that includes a link to your video pitch to be shared in the Discord channel of the course.
