Open-source large language model-based on-premises pipeline for automated data extraction from unstructured electronic health records: a pilot study - Summary - MDSpire

On-Premises Pipeline Utilizing Open-Source Large Language Models for Automated Extraction of Data from Unstructured Electronic Health Records: A Pilot Investigation

  • By

  • Vasileios Ntinopoulos

  • Hector Rodriguez Cetina Biefer

  • Laura Rings

  • Rodney Alexander Rosalia

  • Omer Dzemali

  • June 1, 2026

Share

Objective:

To evaluate an on-premises, open-source large language model (LLM)-based data extraction pipeline for automated data extraction from unstructured electronic health records (EHRs).

Approach:
  • Data Preprocessing: Automated script-based EHR data preprocessing extracted 50 medical texts in German for evaluation.
  • LLM Evaluation: 14 mid-sized LLMs (30B–90B parameters) were evaluated in various tasks including information extraction, binary classification, and multilevel classification.
  • Performance Assessment: LLM response consistency was assessed over three same-prompt iterations.
Key Findings:
  • In overall accuracy, Qwen3-30b-a3b-q8 presented the highest value (0.954) and 13 LLMs had values over 0.90.
  • In information extraction accuracy, 12 LLMs exhibited a value of 1.0 and all 14 LLMs had values over 0.96.
  • In binary classification accuracy, Llama3.2-vision-90b-q4 exhibited the highest value (0.972), 5 LLMs had values of at least 0.95 and all 14 LLMs showed values over 0.93.
  • In multilevel classification accuracy, Qwen3-30b-a3b-q8 exhibited the highest value (0.940), four LLMs had values over 0.90 and all LLMs presented values over 0.80.
  • Nine LLMs exhibited perfect response consistency and the remaining five LLMs had a Krippendorff’s alpha value of 0.999.
Interpretation:

Multiple LLMs demonstrated high accuracy and response consistency, indicating their potential for reliable automation of data extraction from EHRs.

Limitations:
  • The study was limited to a small dataset of 50 medical texts, which may not be representative of broader clinical scenarios.
  • Further validation in larger studies is needed to confirm findings and assess generalizability.
Conclusion:

This pilot study demonstrates the feasibility of on-premises, privacy-preserving, LLM-based automated EHR data extraction pipelines.

Original Source(s)

Related Content