Issue 01AI 바이오 논문
AC POST
AI 바이오 논문 목록
arxiv2026년 9월 20일 09:00

언어 모델 기반의 박테리아 유전자 편집 탐지

최근 유전체 편집 기술의 발전으로 박테리아의 유전적 조작이 용이해졌으며, 이를 통해 병원성 강화나 항생제 내성 확대와 같이 위험할 수 있는 새로운 형질을 부여할 수 있게 되었다. 인위적으로 변형된 박테리아를 탐지하는 능력은 잠재적인 생물학적 위협을 식별하는 데 매우 중요하다. 그러나 박테리아 간의 수평적 유전자 전달을 통한 자연적인 유전자 교환으로 인해 악의적인 유전체 편집을 추적하는 것은 어려울 수 있다. 본 연구에서는 천연 유전체와 시뮬레이션된 편집 유전체를 포함한 광범위한 데이터셋을 큐레이션한 후, 자연어 처리 방식을 활용하여 편집된 유전체를 탐지하였다. 우리는 문장의 단어를 분석하는 대신 유전체의 유전자군을 모델링하는 transformer-encoder 기반의 머신러닝 분류기를 개발하였다. 데이터셋을 통해 모델을 학습시킨 결과, 부자연스러운 맥락으로 인해 박테리아 유전체에 인위적으로 추가된 유전자를 정확하게 탐지할 수 있었다. 우리의 접근 방식은 특정 마커 유전자에 의존하지 않고 설계된 서열을 식별할 수 있는 확장 가능한 방법을 제공하며, 생물 보안, 농업, GMO 규제 등 다양한 분야에 적용될 잠재력이 있다.

Recent advances in genome editing allow easy genetic manipulation of bacteria, providing them with new traits, some of which could be hazardous, e.g. enhanced virulence or extended resistance to antibiotics. The ability to detect artificially modified bacteria is crucial for identifying potential bio-threats. However, malicious genome editing could be challenging to trace due to the natural exchange of genes among bacteria through horizontal transfer. After curating extensive datasets including natural genomes and simulated edited genomes, we utilized a natural language processing approach to detect edited genomes. We developed a transformer-encoder-based machine-learning classifier that, instead of analyzing words in sentences, models gene families in genomes. After training the model on our datasets, it is able to accurately detect genes artificially added to bacterial genomes due to their unnatural context. Our approach provides a scalable method for identifying engineered sequences without relying on specific marker genes, with potential applications in biosecurity, agriculture, GMO regulation and more.