Researchers in Taiwan have developed a hierarchical computational algorithm capable of identifying pregnancies and accurately determining gestational age using routinely collected nationwide health data, according to a study published in Springer’s Pharmacoepidemiology and Drug Safety journal. The method integrates diagnostic codes, procedure records, and laboratory measurements from Taiwan’s National Health Insurance Research Database (NHIRD) to systematically classify pregnancies and estimate their duration—a capability that could enhance epidemiological research and drug safety monitoring in pregnancy cohorts.
Key takeaways
- A hierarchical algorithm successfully extracted pregnancy identification and gestational age estimation from Taiwan’s nationwide linked health insurance records without requiring direct access to obstetric ultrasound or medical charts
- The method combines diagnostic codes, procedure records, and laboratory test dates to classify pregnancies and determine duration with structured clinical logic
- This approach enables large-scale pregnancy cohort identification for pharmacoepidemiology research and drug safety assessment in Taiwan’s 23+ million insured population
Study at a Glance
| Source | Pharmacoepidemiology and Drug Safety |
| Study type | Algorithm development and validation study |
| Data source | Taiwan National Health Insurance Research Database (NHIRD) |
| Population | Pregnant women with delivery records in Taiwan national health system |
| Country | Taiwan |
Algorithm data integration hierarchy
Systematic classification of pregnancy identification and gestational age estimation from linked health records
Source: Springer Pharmacoepidemiology and Drug Safety, 2026 | Georgian Medical Journal News
Extracting pregnancy data from administrative records
Taiwan’s healthcare system maintains comprehensive claims data through the NHIRD, which covers approximately 99% of the country’s population and includes diagnostic codes, procedure records, laboratory results, and prescription dispensing information. However, the NHIRD historically lacked direct pointers to pregnancy episodes or their timing, making large-scale pregnancy cohort studies difficult to assemble.
The hierarchical algorithm addresses this gap by systematically mining diagnostic and procedure records to identify pregnancies and infer gestational age, according to the published methodology. The approach sequences multiple data types in order of reliability: pregnancy-related diagnostic codes (such as routine antenatal care or pregnancy complications) serve as primary identifiers, while delivery codes and obstetric procedures provide temporal anchors. Laboratory test dates—including pregnancy tests and routine prenatal investigations—are cross-referenced to estimate gestational duration.
Clinical and research applications in pharmacovigilance
Determining gestational age accurately is critical for drug safety research because fetal and maternal physiology change substantially across pregnancy trimesters, affecting drug metabolism and teratogenic risk. By enabling systematic identification of pregnancy cohorts and their timing from administrative data, the algorithm allows epidemiologists to conduct large-scale observational studies examining medication exposure during pregnancy—a research area where prospective data collection is often impractical.
The algorithm’s application to Taiwan’s 23+ million insured population creates a foundation for real-world evidence studies on drug safety in pregnancy. Research published in journals including Pharmacoepidemiology and Drug Safety has demonstrated the importance of national pharmacy and medical claims databases for pregnancy exposure assessment, particularly for frequently used medications where observational cohorts can be assembled rapidly.
Implications across the healthcare spectrum
A hierarchical algorithm successfully identifies pregnancies and estimates gestational age from routinely collected health insurance data, enabling large-scale pregnancy cohort assembly without requiring direct access to obstetric charts or ultrasound records.
— Researchers, Springer Pharmacoepidemiology and Drug Safety (2026)
What this means
Scaling automated pregnancy identification to other health systems
Taiwan’s experience demonstrates that structured national health insurance databases—common in many countries—contain sufficient coded information to enable computational pregnancy identification. The hierarchical approach may be transferable to other systems using similar diagnostic coding standards (ICD-9, ICD-10) and procedure classification systems, though validation would be required for each jurisdiction’s specific coding practices and data completeness patterns.
As published in recent analyses of pregnancy cohort assembly methods, automated algorithms reduce the manual chart review burden associated with traditional pregnancy identification approaches. This efficiency gain enables researchers in countries with comparable health information systems to conduct large-scale pregnancy safety studies more rapidly, strengthening the global evidence base on medication safety in pregnancy.
Frequently asked questions
Why is identifying pregnancies from administrative data challenging?
Pregnancies are coded as diagnoses or conditions across multiple healthcare encounters rather than as unified episodes with explicit start and end dates. The algorithm sequences diagnostic, procedure, and laboratory records chronologically to reconstruct pregnancy episodes and estimate duration—a task that requires structured logic to avoid misclassification.
How does the algorithm estimate gestational age without ultrasound records?
The algorithm cross-references pregnancy diagnostic codes with delivery date records and estimated conception dates derived from pregnancy tests and clinical markers. This multi-source approach provides approximations of gestational duration suitable for epidemiological classification, though it may lack the precision of obstetric ultrasound dating.
Can this method be used for drug safety research?
Yes. By identifying pregnancies and their timing, the algorithm enables researchers to link medication dispensing records to pregnancy episodes and classify exposure by trimester—the foundational requirement for pharmacoepidemiological studies of medication safety in pregnancy, as described in clinical pharmacology research.
Taiwan’s hierarchical algorithm represents a practical advance in translating routine health data into research infrastructure for pregnancy cohort studies. As healthcare systems globally invest in data integration and standardization, similar approaches may enable other countries to unlock pregnancy safety research from their own administrative databases, accelerating evidence synthesis on medication safety across the globe.
Was this article helpful?
Disclaimer. This article is health journalism intended for general information and education. It is not medical advice and is not a substitute for professional diagnosis or treatment. Always consult a qualified healthcare provider about your individual circumstances. Full disclaimer →
Related Coverage




Medically reviewed by Prof. Giorgi Pkhakadze, MD, MPH, PhD. Spotted an error? Contact the editorial team.




