Open-source large language model-based on-premises pipeline for automated data extraction from unstructured electronic health records: a pilot study - Takeaways - MDSpire

On-Premises Pipeline Utilizing Open-Source Large Language Models for Automated Extraction of Data from Unstructured Electronic Health Records: A Pilot Investigation

  • By

  • Vasileios Ntinopoulos

  • Hector Rodriguez Cetina Biefer

  • Laura Rings

  • Rodney Alexander Rosalia

  • Omer Dzemali

  • June 1, 2026

Share

  • 1

    The study evaluated an on-premises, open-source LLM-based data extraction pipeline for unstructured electronic health records.

  • 2

    Fifty medical texts in German were processed using 14 mid-sized LLMs, assessing their performance in various classification tasks.

  • 3

    Qwen3-30b-a3b-q8 achieved the highest overall accuracy of 0.954, with 13 LLMs exceeding 0.90 in overall accuracy.

  • 4

    Twelve LLMs demonstrated perfect accuracy in information extraction, while nine LLMs showed perfect response consistency.

  • 5

    The pilot study indicates the feasibility of LLM-based automated EHR data extraction, warranting further validation in healthcare.

Original Source(s)

Related Content