Doctoral researcher · Medical AI & Clinical NLP
Hendrik Damm.
Together with clinicians and colleagues, I develop and evaluate language models for medicine — from literature screening for clinical guidelines to grounded question answering on health records.

- peer-reviewed publications
- 19
- first places with our team in international shared tasks
- 3
- ImageCLEFmedical Caption editions co-organised, two as lead organiser
- 3
- manuscripts in peer review, 3 as first author
- 4
01About
Between computer science and medicine.
I’m a doctoral researcher in the DFG Research Training Group WisPerMed, which studies knowledge- and data-based personalisation of medicine at the point of care. I work at the Department of Computer Science of FH Dortmund and the Institute for Medical Informatics, Biometry and Epidemiology at University Hospital Essen, and I’m pursuing my doctorate (Dr. rer. medic.) at the University of Duisburg-Essen.
My current focus is AI support for clinical guideline development: together with the experts behind the German S3 melanoma guideline, I study whether large language models can act as additional reviewers in literature screening — and how physicians come to trust such systems. Alongside this, I work with colleagues on grounded question answering over health records and on clinical text generation, from discharge summaries to lay-language versions of medical documents.
Since 2024, I have also been part of the organising team of ImageCLEFmedical Caption, an international benchmark for medical image understanding; I led the team in 2025 and 2026. Before my doctorate, I studied medical informatics in Dortmund and worked in data management and data integration at Hannover Medical School and University Hospital Münster.
02Research
What I work on
My work ranges from method development, often in international shared tasks, to studies with physicians and guideline experts.
01 · Evidence synthesis · Abstract screening · Oncology
AI for living clinical guidelines
An AI assistance system for updating the German S3 melanoma guideline: large language models as additional reviewers in title-and-abstract screening, evaluated against the guideline experts’ decisions — a step towards living guidelines.
6 related publications02 · EHR question answering · Attribution · Reliability
Grounded & trustworthy LLMs
Question answering over electronic health records that cites its evidence, claim-level attribution for generated text, and detecting unreliable model answers to medical questions.
6 related publications03 · Summarization · Lay language · Interoperability
Clinical text generation & extraction
From free text to structured data and back: discharge and tumor-board summaries, lay-language versions of clinical documents, and FHIR-ready information extraction with open-weight models.
9 related publications04 · Trust · Eye tracking · Interviews
Human–AI collaboration in medicine
How physicians perceive, trust and work with AI: a multicentre randomised eye-tracking study on AI-generated trust labels and an interview study on AI-assisted literature screening in guideline development.
3 related publications05 · Benchmarking · Multimodal · Image captioning
ImageCLEFmedical Caption
An international benchmark for medical concept detection and caption generation, held at CLEF; its tenth edition took place in 2026.
8 related publications03Selected work
Selected publications
In peer review
- Revision under review · round 2npj Digital Medicine
Criterion guided large language models as additional reviewers for melanoma guideline title and abstract screening
Hendrik Damm, Enis Ömer Doğru, Noëlle Bender, …, Christoph M. Friedrich
- Major revision · Oct 2026JMIR AI
Trust Requirements for AI-Assisted Title-Abstract Screening in Clinical Guideline Development: Qualitative Interview Study
Hendrik Damm, Kilian Elfert, Noëlle Bender, …, Elisabeth Livingstone
- Under reviewPLOS Digital Health
AI-framed evidence labels redirect physician attention and increase expert-concordant selection in multicentre randomised eye-tracking study
Hendrik Damm, Kilian Elfert, Noëlle Bender, …, Christoph M. Friedrich
- Minor revision submitted · round 2npj Digital Medicine
Challenges in AI Based Tumor Board Case Summarization and Recommendations
Wen-wai Yim, Hendrik Damm, Tabea M. G. Pakull, …, Georg Lodde
Published
- 2026JMIR 2026
Extracting Medical Information From Unstructured Clinical Text Using Large Language Models to Enhance Health Care Interoperability: Proof-of-Concept Study
Bahadır Eryılmaz, Kamyar Arzideh, Mikel Bahn, Hendrik Damm, …, René Hosch
- 2026CL4Health @ LREC 20261st place · Subtask 3
WisPerMed at ArchEHR-QA 2026: Retrieval-Augmented Prompting for Grounded EHR Question Answering
Jan-Henning Büns, Tabea M. G. Pakull, Hendrik Damm, …, Norbert Fuhr
- 2025EJC Skin Cancer
Comparative analysis of international melanoma guidelines
Ahmad Idrissi-Yaghir, Henning Schäfer, Hendrik Damm, Georg C. Lodde, Elisabeth Livingstone, Dirk Schadendorf, Christoph M. Friedrich
- 2025CL4Health @ NAACL 20251st place overall
WisPerMed @ PerAnsSumm 2025: Strong Reasoning Through Structured Prompting and Careful Answer Selection Enhances Perspective Extraction and Summarization of Healthcare Forum Threads
Tabea M. G. Pakull, Hendrik Damm*, Henning Schäfer, Peter A. Horn, Christoph M. Friedrich
- 2025CLEF 2025 Working Notes
Overview of ImageCLEFmedical 2025 – Medical Concept Detection and Interpretable Caption Generation
Hendrik Damm, Tabea M. G. Pakull, Helmut Becker, …, Christoph M. Friedrich
- 2024BioNLP @ ACL 20241st place
WisPerMed at “Discharge Me!”: Advancing Text Generation in Healthcare with Large Language Models, Dynamic Expert Selection, and Priming Techniques on MIMIC-IV
Hendrik Damm, Tabea M. G. Pakull, Bahadır Eryılmaz, Helmut Becker, Ahmad Idrissi-Yaghir, Henning Schäfer, Sergej Schultenkämper, Christoph M. Friedrich
04News
Recent updates
- PaperNew preprint led by Bohao Chu: “Locally Sound, Globally Insufficient” — reasoning traces can be supported at every step and still fail the question as a whole, a failure regime we call the local-global gap in multi-hop reasoning.
- ServiceImageCLEFmedical Caption 2026, the task’s 10th edition, concluded at CLEF 2026 in Jena; I led its organising team. The ImageCLEF 2026 overview is out in Springer LNCS.
- TalkTalk at the 36th German Skin Cancer Congress (ADO) in Leipzig: an AI assistance system for updating the S3 melanoma guideline — requirements, early practical testing and the transition to living guidelines.
- PaperPublished in JMIR, led by Bahadır Eryılmaz: fine-tuning open-source LLMs on synthetic FHIR-derived discharge letters to extract interoperable information from clinical text.
- AwardArchEHR-QA 2026: 1st place in Subtask 3 and 2nd place in Subtask 4 with our retrieval-augmented system (CL4Health @ LREC 2026, Palma — 43 teams).
- PaperTwo new preprints led by Bohao Chu: eTracer — traceable text generation via claim-level grounding — and PCoA, a benchmark for medical aspect-based summarization with phrase-level context attribution.
- ServiceImageCLEFmedical Caption 2025 (9th edition) took place at CLEF 2025 in Madrid; I led the organising team and organised the task’s session at the CLEF Labs workshop.