Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.
Evaluating dimensionality reduction of comorbidities for predictive modeling in individuals with neurofibromatosis type 1
2
Zitationen
8
Autoren
2024
Jahr
Abstract
Objective: Dimensionality reduction techniques aim to enhance the performance of machine learning (ML) models by reducing noise and mitigating overfitting. We sought to compare the effect of different dimensionality reduction methods for comorbidity features extracted from electronic health records (EHRs) on the performance of ML models for predicting the development of various sub-phenotypes in children with Neurofibromatosis type 1 (NF1). Materials and Methods: EHR-derived data from pediatric subjects with a confirmed clinical diagnosis of NF1 were used to create 10 unique comorbidities code-derived feature sets by incorporating dimensionality reduction techniques using raw International Classification of Diseases codes, Clinical Classifications Software Refined, and Phecode mapping schemes. We compared the performance of logistic regression, XGBoost, and random forest models utilizing each feature set. Results: XGBoost-based predictive models were most successful at predicting NF1 sub-phenotypes. Overall, features based on domain knowledge-informed mapping schema performed better than unsupervised feature reduction methods. High-level features exhibited the worst performance across models and outcomes, suggesting excessive information loss with over-aggregation of features. Discussion: Model performance is significantly impacted by dimensionality reduction techniques and varies by specific ML algorithm and outcome being predicted. Automated methods using existing knowledge and ontology databases can effectively aggregate features extracted from EHRs. Conclusion: Dimensionality reduction through feature aggregation can enhance the performance of ML models, particularly in high-dimensional datasets with small sample sizes, commonly found in EHRs health applications. However, if not carefully optimized, it can lead to information loss and data oversimplification, potentially adversely affecting model performance.
Ähnliche Arbeiten
Comprehensive, Integrative Genomic Analysis of Diffuse Lower-Grade Gliomas
2015 · 3.221 Zit.
Oncogenes and signal transduction
1991 · 2.760 Zit.
THE RECURRENCE OF INTRACRANIAL MENINGIOMAS AFTER SURGICAL TREATMENT
1957 · 2.410 Zit.
Gastrointestinal stromal tumors: Pathology and prognosis at different sites
2006 · 2.038 Zit.
Identification and characterization of the tuberous sclerosis gene on chromosome 16
1993 · 1.671 Zit.