Skip to main content
QUICK REVIEW

[논문 리뷰] AptaFind: A lightweight local interface for automated aptamer curation from scientific literature

Geoffrey J. Taghon|arXiv (Cornell University)|2026. 01. 12.
Biomedical Text Mining and Ontologies인용 수 0
한 줄 요약

AptaFind는 로컬 언어 모델과 결정적 정규식 방법을 결합하여 문헌에서 어프타머 데이터를 선별하는 3단계(on-device) 파이프라인을 통해 각 계층의 커버리지 약 ~84% 및 처리 속도 약 ~953 targets/hour를 달성한다.

ABSTRACT

Aptamer researchers face a literature landscape scattered across publications, supplements, and databases, with each search consuming hours that could be spent at the bench. AptaFind transforms this navigation problem through a three-tier intelligence architecture that recognizes research mining is a spectrum, not a binary success or failure. The system delivers direct sequence extraction when possible, curated research leads when extraction fails, and exhaustive literature discovery for additional confidence. By combining local language models for semantic understanding with deterministic algorithms for reliability, AptaFind operates without cloud dependencies or subscription barriers. Validation across 300 University of Texas Aptamer Database targets demonstrates 84 % with some literature found, 84 % with curated research leads, and 79 % with a direct sequence extraction, at a laptop-compute rate of over 900 targets an hour. The platform proves that even when direct sequence extraction fails, automation can still deliver the actionable intelligence researchers need by rapidly narrowing the search to high quality references.

연구 동기 및 목표

  • 산재한 어프타머 문헌 문제를 해결하고 수작업 큐레이션 노력을 줄인다.
  • 신뢰성을 위한 로컬(클라우드 없이) 파이프라인을 개발하고, 언어 모델과 결정론적 구문 분석을 결합한다.
  • 정확도와 커버리지의 균형을 위해 직접 시퀀스, 큐레이션된 리드, 포괄적 문헌의 세 계층 출력을 제공한다.
  • UT 데이터베이스 타깃에서 접근법을 검증하고 계층별 성능을 정량화한다.
  • 프라이버시와 재현성에 중점을 둔 오픈 소스 소프트웨어를 제공한다.

제안 방법

  • 직접 시퀀스 추출(Tier 1), 큐레이션된 리드(Tier 2), 포괄적 문헌 발견(Tier 3)을 포함하는 3단계 인텔리전스 아키텍처를 통합한다.
  • 온-디바이스 처리로 의미 이해를 위해 로컬 1B 파라미터 Llama3.2 모델을 사용한다.
  • 검증, 중복 제거 및 서식을 위한 결정적 정규식 파이프라인(Minimum Agentic Flow, MAF)과 언어 모델 지침을 결합한다.
  • 뉴클레오타이드 서열(20–100 nt)과 결합 데이터(Kd, Ki)를 정규식으로 단위 보존하며 추출하고 LM 맥락과 데이터를 조화시킨다.
  • 로컬 PDF, PubMed/PMC, bioRxiv에서 다중 소스 검색을 수행하고 필요 시 보충 수확과 브라우저 자동화를 수행한다.
  • 서열이 생물학적 제약(길이 20–100 nt, GC 20–80%, 5’→3’ 방향)에 부합하는지 검증하고, 출처 간 100% 동질성으로 중복 제거한다.
Figure 1: AptaFind implements a three-tier research intelligence approach ensuring every search delivers actionable value. (A) System Architecture: Multi-source literature discovery combines PubMed, PMC, and bioRxiv searches with supplement harvesting and browser automation for comprehensive coverag
Figure 1: AptaFind implements a three-tier research intelligence approach ensuring every search delivers actionable value. (A) System Architecture: Multi-source literature discovery combines PubMed, PMC, and bioRxiv searches with supplement harvesting and browser automation for comprehensive coverag

실험 결과

연구 질문

  • RQ1AptaFind의 세 계층 출력이 다양한 문헌 소스에서 어프타머 데이터를 얼마나 효과적으로 포착하는가?
  • RQ2일반 하드웨어에서 계층별 회복률과 처리 속도는 어느 정도인가?
  • RQ3Minimum Agentic Flow 원칙이 LM 중심 접근법이나 정규식 중심 접근법보다 신뢰성을 개선하는가?
  • RQ4유료 콘텐츠, 이미지 기반 서열 추출, 복잡한 표와 관련된 한계는 무엇인가?
  • RQ5이 접근법이 어프타머 외의 다른 문헌 마이닝 도메인으로 일반화될 수 있는가?

주요 결과

  • Tier 3 (Literature Discovery) 구문 across 100-target samples에서 84.0% ± 3.5% 커버리지 달성.
  • Tier 2 (Research Leads) 구문 across 100-target samples에서 84.0% ± 3.5% 커버리지 달성.
  • Tier 1 (Direct Extraction) 구문 across 100-target samples에서 79.3% ± 0.6% 커버리지 달성.
  • Processing speed는 Mac Studio (M2 Max)에서 약 954 ± 43 targets per hour 수준.
  • Validation은 MAF 접근 방식이 의미 이해를 결정적 데이터 처리와 분리하여 신뢰성을 향상시키는 것을 보여준다.
  • 방법론은 클라우드 의존 없이 로컬로 유지되며, 시간당 약 1000 타깃 쿼리를 제공하고 빠른 문헌 선별을 가능하게 한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.