관찰된 표현형 및 제외된 표현형을 기반으로 한 희귀 질환 사례와 대조군의 판별
희귀 질환은 개별적으로는 드물지만 집단적으로는 널리 퍼져 있다. 이들의 주요 임상적 과제는 치료가 아닌 진단에 있다. 임상 관리 초기 단계에서는 관찰된 표현형이 희귀 질환과 관련이 있는지 여부가 불분명한 경우가 많다. 조기 진단을 위해 머신러닝 방법을 활용하여 이러한 표현형과 희귀 질환 사이의 잠재적 연관성을 발굴하는 것은 진단적 딜레마를 완화할 수 있는 실행 가능한 전략을 제공한다. 본 연구에서 우리는 머신러닝이 관찰된 표현형과 제외된 표현형을 특징(feature)으로 사용하여 희귀 질환 사례를 대조군으로부터 효과적으로 판별할 수 있음을 입증하였다. 그중 Random Forest 모델이 일정 수준의 일반화 성능과 함께 가장 우수한 분류 성능을 달성하였다. 특징 선택 결과에 대한 추가 분석을 바탕으로, 우리는 특이성(specificity)과 발생 횟수(occurrence count)라는 두 요소가 희귀 질환 판별을 위한 표현형 선택에 중요하며, 특징 구축 과정에서 관찰된 표현형과 제외된 표현형 모두에 대등한 중요성을 부여해야 한다는 결론을 내렸다.
Rare diseases are individually uncommon but collectively prevalent. Their primary clinical challenge lies not in treatment but in diagnosis. In the early stages of clinical management, it is frequently unclear whether the observed phenotypes are associated with a rare disease. Leveraging machine learning methods to mine latent associations between these phenotypes and rare diseases for early diagnosis offers a viable strategy to alleviate this diagnostic dilemma. In the present study, we demonstrated that machine learning can effectively discriminate rare disease cases from their controls using observed and excluded phenotypes as features. Among them, the Random Forest model achieved the best classification performance with a certain degree of generalizability. Based on further analysis of the feature selection results, we conclude that the two factors, specificity and occurrence count, are important for phenotype selection in rare disease discrimination, and comparable importance should be attached to both observed and excluded phenotypes, during feature construction.