You will contribute to the development of infrastructure connecting molecular dynamics simulations with structural biology resources and biological knowledge bases. A major component of the role will be developing AI-driven approaches to mine scientific literature and automatically extract experimental and biological metadata to enrich MD datasets. You will develop and extend SIFTS (Structure Integration with Function, Taxonomy and Sequence), a core PDBe resource that provides residue-level mappings between PDB structures, UniProtKB sequences, and other biological resources, to facilitate the integration of MD-derived insights across the wider life sciences data ecosystem.
This is an interdisciplinary role combining structural bioinformatics, molecular dynamics, and scientific software development. You will apply both scientific understanding and technical expertise to develop data integration workflows, APIs, and biological annotations that improve interoperability and reuse of structural and molecular simulation data across various resources.
Primary responsibilities:
Design and implement data integration pipelines that connect MDDB with major life science resources, including PDBe, UniProt, PDBe-KB, and other relevant knowledge bases and databases
Develop and deploy AI- and machine learning-based approaches for extracting experimental and biological metadata from scientific literature to enrich MDDB datasets and support downstream biological interpretation
Extending and maintaining the SIFTS infrastructure and codebase to support integration of molecular dynamics and other data resources
Develop and maintain software tools, APIs, workflows, and documentation that facilitate FAIR data integration, metadata enrichment, and integrated data access
Collaborating with domain experts, software engineers, and data resource providers to enable the integration of MD-derived biological insights into the wider life sciences data ecosystem
Supporting FAIRification, standardisation, and interoperability of MD datasets and associated annotations
Collaborating with international partners across ELIXIR, Instruct-ERIC, EU-OPENSCREEN, HPC centres, and industry
Participating in community standards development, technical documentation, training, outreach, and dissemination activities
You have
PhD in Bioinformatics, Computational Biology, Structural Biology, Computer Science, Data Science, or a related field
Familiarity with structural biology and molecular simulation data
Experience with NLP/LLM-based scientific literature mining
Demonstrated experience with FAIR data principles, metadata standards, and scientific repositories
Understanding of sequence, structure, and functional annotations of proteins
Experience in scientific software development, preferably in Python
Experience with Linux environments, Git, and CI/CD practices
Scientific publications relevant to structural biology, bioinformatics, or protein annotations
Strong communication, collaboration, and problem-solving skills
You may also have
Postdoctoral research experience in a relevant field
Experience with graph databases (e.g. Neo4J), REST APIs, containerisation technologies, and workflow management systems such as Nextflow
Experience in data visualisation and analysis
Understanding of FAIR data principles and the biological data lifecycle
Experience in reporting and presenting scientific topics
Experience working in international and interdisciplinary teams
Deadline 19 July
Don't forget to mention EuroScienceJobs when applying.