Open-source large language model-based on-premises pipeline for automated data extraction from unstructured electronic health records: a pilot study - Report - MDSpire

On-Premises Pipeline Utilizing Open-Source Large Language Models for Automated Extraction of Data from Unstructured Electronic Health Records: A Pilot Investigation

  • By

  • Vasileios Ntinopoulos

  • Hector Rodriguez Cetina Biefer

  • Laura Rings

  • Rodney Alexander Rosalia

  • Omer Dzemali

  • June 1, 2026

Share

Clinical Report: On-Premises Pipeline Utilizing Open-Source LLMs for EHR Data Extraction

Overview

This pilot study evaluated an on-premises, open-source large language model (LLM)-based data extraction pipeline for automated extraction of data from unstructured electronic health records (EHRs). The results indicated high accuracy across various tasks.

Background

Data extraction from electronic health records (EHRs) is essential for clinical practice and research but is often labor-intensive and prone to errors. This study investigates the performance of LLMs in extracting structured data from unstructured medical texts.

Data Highlights

The study evaluated 14 mid-sized open-source LLMs across multiple tasks, achieving high accuracy rates in information extraction, binary classification, and multilevel classification.

Key Findings

  • Qwen3-30b-a3b-q8 achieved the highest overall accuracy of 0.954.
  • 12 LLMs exhibited perfect accuracy (1.0) in information extraction tasks.
  • Llama3.2-vision-90b-q4 had the highest binary classification accuracy at 0.972.
  • Qwen3-30b-a3b-q8 also led in multilevel classification accuracy with a score of 0.940.
  • Nine LLMs demonstrated perfect response consistency, with others showing a Krippendorff’s alpha value of 0.999.

Clinical Implications

Further validation in larger studies is necessary to confirm these results.

Conclusion

This pilot study demonstrates the feasibility of LLM-based automated data extraction pipelines for EHRs.

Related Resources & Content

  1. JMIR, Journal of Medical Internet Research, 2026 -- Automated Identification of Nursing Diagnoses and Interventions From Nursing Records Using a Retrieval-Augmented Large Language Model Approach: Quantitative Study
  2. npj Digital Medicine, 2026 -- Enhanced Transferability of Predictions from Electronic Health Records Across Different Countries and Coding Frameworks Using Large Language Models
  3. asco ai in oncology, 2026 -- Can AI-Extracted EHR Data Be Trusted? The VALID Framework Takes Aim at a Growing Problem
  4. aace endocrine ai, 2026 -- Selective LLM use may improve electronic health record phenotyping accuracy
  5. WHO, 2024 -- WHO releases AI ethics and governance guidance for large multi-modal models
  6. WHO releases AI ethics and governance guidance for large multi-modal models
  7. https://academic.oup.com/jamia/article/33/3/553/8425815

Original Source(s)

Related Content