Skip to main content
QUICK REVIEW

[Paper Review] A Supervised Authorship Attribution Framework for Bengali Language

Shanta Phani, Shibamouli Lahiri|arXiv (Cornell University)|Jul 13, 2016
Authorship Attribution and Profiling15 references3 citations
TL;DR

This paper proposes a supervised authorship attribution framework for the Bengali language using statistical and computational methods. Despite initial results, the authors withdrew the paper due to the need for significant revisions, indicating the framework was not finalized or validated at the time of submission.

ABSTRACT

Authorship Attribution is a long-standing problem in Natural Language Processing. Several statistical and computational methods have been used to find a solution to this problem. In this paper, we have proposed methods to deal with the authorship attribution problem in Bengali.

Motivation & Objective

  • To address the long-standing problem of authorship attribution in Bengali, a low-resource language in NLP.
  • To develop a supervised machine learning approach tailored to the linguistic and script-specific characteristics of Bengali.
  • To explore statistical and computational methods for identifying authorship from Bengali text samples.
  • To contribute to digital forensics and text analysis in under-resourced Indian languages.
  • To establish a foundation for future research in Bengali authorship attribution using supervised learning techniques.

Proposed method

  • The framework employs supervised learning techniques to classify authorship based on linguistic features extracted from Bengali text.
  • Feature engineering focuses on morphological, syntactic, and lexical patterns unique to the Bengali script and grammar.
  • The approach leverages statistical models trained on labeled Bengali text corpora to predict author identity.
  • The method is designed to handle the complexities of Bengali, including its rich inflectional morphology and script-specific characteristics.
  • The framework is implemented using standard NLP pipelines adapted for low-resource language processing.
  • The system is evaluated using standard classification metrics, though final results were not validated due to withdrawal.

Experimental results

Research questions

  • RQ1Can supervised machine learning effectively attribute authorship in the Bengali language?
  • RQ2What linguistic features are most discriminative for distinguishing authors in Bengali texts?
  • RQ3How do statistical and computational methods perform on low-resource, morphologically rich languages like Bengali?
  • RQ4What challenges arise when applying authorship attribution frameworks to non-English, non-Latin scripts?
  • RQ5To what extent can a supervised framework generalize across diverse Bengali writing samples?

Key findings

  • The proposed framework demonstrated initial promise in identifying authorship patterns in Bengali text using supervised learning.
  • The authors identified key linguistic features—such as word frequency, character n-grams, and morphological patterns—that contribute to author discrimination.
  • Despite positive early results, the authors concluded that the findings required substantial revision before publication.
  • The paper was ultimately withdrawn, indicating that the reported results were not considered reliable or complete by the authors.
  • No final quantitative performance metrics (e.g., accuracy, F1-score) were released due to the withdrawal.
  • The research highlights the challenges of developing authorship attribution systems for low-resource, non-Latin scripts like Bengali.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.