[논문 리뷰] Large Language Models for Patent Classification: Strengths, Trade-offs, and the Long Tail Effect
논문은 CPC 특허 분류를 위해 엔코더 기반 분류기와 오픈-가중치 LLM을 비교하며, 정확도, 효율성, 장-tail 커버리지에서 보완적 강점과 트레이드오프를 강조한다.
Patent classification into CPC codes underpins large scale analyses of technological change but remains challenging due to its hierarchical, multi label, and highly imbalanced structure. While pre Generative AI supervised encoder based models became the de facto standard for large scale patent classification, recent advances in large language models (LLMs) raise questions about whether they can provide complementary capabilities, particularly for rare or weakly represented technological categories. In this work, we perform a systematic comparison of encoder based classifiers (BERT, SciBERT, and PatentSBERTa) and open weight LLMs on a highly imbalanced benchmark dataset (USPTO 70k). We evaluate LLMs under zero shot, few shot, and retrieval augmented prompting, and further assess parameter efficient fine tuning of the best performing model. Our results show that encoder based models achieve higher aggregate performance, driven by strong results on frequent CPC subclasses, but struggle on rare ones. In contrast, LLMs achieve relatively higher performance on infrequent subclasses, often associated with early stage, cross domain, or weakly institutionalised technologies, particularly at higher hierarchical levels. These findings indicate that encoder based and LLM based approaches play complementary roles in patent classification. We additionally quantify inference time and energy consumption, showing that encoder based models are up to three orders of magnitude more efficient than LLMs. Overall, our results inform responsible patentometrics and technology mapping, and motivate hybrid classification approaches that combine encoder efficiency with the long tail coverage of LLMs under computational and environmental constraints.
연구 동기 및 목표
- 대규모, 계층적이며 불균형한 라벨 공간에 대한 자동화된 CPC 특허 분류를 촉진한다.
- 낮은 빈도이거나 약하게 표현된 CPC 하위 클래스에 대해 LLM이 엔코더 기반 모델을 보완할 수 있는지 평가한다.
- 엔코더 방식과 LLM 접근 방식의 추론 효율성과 환경 영향을 평가한다.
- 정확도, 확장성 및 지속 가능성을 균형 있게 달성하는 하이브리드 전략의 가능성을 조사한다.
- 일반화 가능성을 테스트하기 위한 독립적인 특허 데이터 세트에 대한 외부 검증을 제공한다.
제안 방법
- USPTO-70k CPC 분류 과제에서 감독된 엔코더 베이스라인(BERT, SciBERT, PatentSBERTa)과 오픈-가중치 LLM을 비교한다.
- LLM에 대한 zero-shot, few-shot 및 검색 증강 프롬프트를 평가한다.
- 최고 성능의 LLM 구성에 매개변수 효율적인 LoRA 미세조정을 적용한다.
- CodeCarbon을 사용하여 에너지 사용량과 CO2 배출량을 정량화하고 지속 가능성 트레이드오프를 분석한다.
- 레이블 계층에 걸친 부트스트랩 신뢰 구간 및 통계적 테스트와 함께 계층적 및 매크로/마이크로 평균 지표를 평가한다.

실험 결과
연구 질문
- RQ1자주 등장하는 CPC 하위 클래스에서 엔코더 기반 모델은 LLM과 비교하여 어떻게 성능을 보이나?
- RQ2특히 희귀하거나 새로 생성 중인 카테고리에서 LLM이 CPC 계층 전체에 걸쳐 더 균형 잡힌 커버리지를 제공하는가?
- RQ3대규모에서 LLM을 사용하는 것과 엔코더 모델을 사용하는 것의 효율성과 에너지 영향은 무엇인가?
- RQ4두 패러다임의 강점을 활용하는 하이브리드 접근 방식이 강인한 특허 분류에 도움이 되는가?
- RQ5재훈련 없이 외부 데이터셋(예: EPO 특허)에 대한 발견의 일반화 정도는 어느 정도인가?
주요 결과
- 자주 등장하는 CPC 하위 클래스에 의해 주도되어 엔코더 기반 모델이 더 높은 총합 성능을 제공한다.
- LLM은 드문 하위 클래스와 더 높은 계층 수준에서 상대적으로 더 강한 성능을 보인다.
- LLM은 엔코더 모델에 비해 훨씬 높은 계산 및 에너지 비용이 든다.
- 단일한 접근법이 우세하지 않다; 엔코더는 효율성과 확장성에 우수하고, LLM은 보완적인 롱-테일 커버리지를 제공한다.
- 도메인에 특화된 엔코더(예: SciBERT)가 일반 엔코더를 능가하고, 검색 강화 프롬프트는 드문 클래스에 대한 LLM 재현율을 향상시킨다.
- 실제 자원 제약 하에서 엔코더의 효율성과 LLM의 롱-테일 능력을 결합한 하이브리드 워크플로가 동기가 된다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.