[논문 리뷰] pyBibX -- A Python Library for Bibliometric and Scientometric Analysis Powered with Artificial Intelligence Tools
pyBibX는 Scopus, Web of Science, PubMed 데이터에 대해 AI 도구를 이용한 EDA, 네트워크 분석 및 NLP를 결합한 생물통계적 및 과학계통학적 분석을 수행하는 파이썬 라이브러리입니다.
Bibliometric and Scientometric analyses offer invaluable perspectives on the complex research terrain and collaborative dynamics spanning diverse academic disciplines. This paper presents pyBibX, a python library devised to conduct comprehensive bibliometric and scientometric analyses on raw data files sourced from Scopus, Web of Science, and PubMed, seamlessly integrating state of the art AI capabilities into its core functionality. The library executes a comprehensive EDA, presenting outcomes via visually appealing graphical illustrations. Network capabilities have been deftly integrated, encompassing Citation, Collaboration, and Similarity Analysis. Furthermore, the library incorporates AI capabilities, including Embedding vectors, Topic Modeling, Text Summarization, and other general Natural Language Processing tasks, employing models such as Sentence-BERT, BerTopic, BERT, chatGPT, and PEGASUS. As a demonstration, we have analyzed 184 documents associated with multiple-criteria decision analysis published between 1984 and 2023. The EDA emphasized a growing fascination with decision-making and fuzzy logic methodologies. Next, Network Analysis further accentuated the significance of central authors and intra-continental collaboration, identifying Canada and China as crucial collaboration hubs. Finally, AI Analysis distinguished two primary topics and chatGPT preeminence in Text Summarization. It also proved to be an indispensable instrument for interpreting results, as our library enables researchers to pose inquiries to chatGPT regarding bibliometric outcomes. Even so, data homogeneity remains a daunting challenge due to database inconsistencies. PyBibX is the first application integrating cutting-edge AI capabilities for analyzing scientific publications, enabling researchers to examine and interpret these outcomes more effectively.
연구 동기 및 목표
- 원시 데이터 소스에서 주요 데이터베이스 전반에 걸친 포괄적인 생물통계적 및 과학계통학적 분석 가능.
- 생물통계 워크플로에 최첨단 AI 기능(임베딩, 토픽 모델링, 텍스트 요약)을 통합.
- 그래픽 시각화 및 네트워크 분석(Citation, Collaboration, Similarity)을 통한 탐색적 데이터 분석 제공.
- 실제 데이터셋에서 라이브러리를 시연하여 인사이트와 실용적 활용을 보여줌.
제안 방법
- Scopus, Web of Science, PubMed에서 데이터 수집.
- 그래픽 출력이 포함된 탐색적 데이터 분석 수행.
- Citation, Collaboration, Similarity 분석을 위한 네트워크 분석 모듈 통합.
- 임베딩 벡터, Sentence-BERT, BerTopic, BERT, chatGPT, PEGASUS를 포함한 AI/NLP 도구를 텍스트 작업에 적용.
- chatGPT 질의를 통해 결과 해석의 상호작용적 해석 가능.
실험 결과
연구 질문
- RQ1AI 도구가 생물통계적 및 과학계통학적 결과 해석을 어떻게 향상시킬 수 있는가?
- RQ2분석 데이터세트의 중심 저자 및 협력 패턴은 무엇인가?
- RQ3AI 보조 텍스트 분석에서 어떤 주제가 등장하는가?
- RQ4데이터 이질성과 데이터베이스 불일치가 소스 간 분석에 어떤 영향을 미치는가?
주요 결과
- EDA는 의사 결정 및 퍼지 로직(methodologies)에 대한 관심이 증가하고 있음을 시사한다.
- 네트워크 분석은 중심 저자와 대륙 간 협력의 주목할 만함을 식별하며, 캐나다와 중국이 협력 허브로 나타난다.
- AI 분석은 두 가지 주요 주제를 드러내고 텍스트 요약에서 chatGPT의 우위를 보인다.
- 이 라이브러리는 연구자가 bibliometric 결과에 대해 chatGPT에 질문할 수 있도록 해준다.
- 데이터 동질성은 데이터베이스 간 불일치로 인해 여전히 도전적이다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.