Project

Communication channel: https://discord.gg/CvujK46me

Objective

Developing a scientific process curation hybrid methodology with associated (meta)-data processing tools. The methodology must consider a qualitative and quantitative philosophy associated to the different data that you manage in your PhD.

ToDo

  1. First catalogue of the data (scientific, documents, presentations, code, plots, papers/reports, …) you manage in your scientific practice including the source, format, people who generate, have access and modify the data.
  2. Choose of data type of interest, and associate a simplified catalogue of meta-data that you can associate to it.
  3. Create a meta-data repository, you can choose to use ATLAS or Glue.
    • Implement a meta-data extraction solution of the data you chose to work with (notebook and example package in Hands-On)
    • Enable a meta-data feeding of the catalogue and show the lineage pipeline in ATLAS/Glue
    • Propose a set of queries that allow to explore your meta-data to answer governance related questions.
  4. What are the governance and management guarantees that you are ensuring with your solution?
  5. What is the degree of reproducibility that you are allowing?
  6. What are the FAIR (technical), CARE (governance) and Feminist properties (power) values addressed/ensured by your project?

General Rules

  1. Individual work.
  2. Notebooks should be publicly accessible online (Colab, Github) with a readme in which you should include answers to questions 4-6.
  3. You can use libraries, methods and any other material required respecting authors’ intellectual property.

To Handin

  1. Notebook/Gist/Github repository with the complete profiling, study and preparation of your data. Use plots whenever possible.
  2. If data is prepared or engineered, a link to the repository of the prepared dataset
  3. Short demo – video pitch and medium like post describing your methodology where you include answers to quetions 4-6 (see above in the ToDo section). The document strategy should be complete, a meta-data repository should be (partially) built, use queries to show and illustrate content, reproducibility can be partially implemented.

Schedule

  • Proposal of the project: 15/09/2026 (by Genoveva)
  • First control (general idea including (meta)-data catalogue, context, people): 09/10/2026
  • Second control (problem statement and first provenance based process curation strategy): 16/10/2026
    • Upload a pdf document with the title of your project, your name and contact coordinates, github repository address and description of data and metadata, associated processes to create, modify and maintain, meta-data catalogue: https://drive.google.com/drive/folders/1CdYEInWjoswtyGgnx7PemtpuaeGWzzFC?usp=sharing
    • Github repository: notebook(s) for extracting meta-data, and connection with catalogue feeding
    • Name your file with <first_name>-<last_name>-strategy.pdf
  • Final results (pipeline, insight, and some degree of reproducibility): 25/10/2026
    • Link to a Document Medium entry like that includes a link to your video pitch to be shared in the Discord channel of the course.