[Paper Review] A Lightweight Regression Method to Infer Psycholinguistic Properties for Brazilian Portuguese
This paper proposes a lightweight regression method to infer psycholinguistic properties—concreteness, age of acquisition, imageability, and subjective frequency—for Brazilian Portuguese using minimal, widely available features: word length, frequency lists, school dictionary data, and word embeddings. The approach achieves correlation levels comparable to prior work, producing a freely available lexicon of 26,874 words annotated with four key psycholinguistic properties.
Psycholinguistic properties of words have been used in various approaches to Natural Language Processing tasks, such as text simplification and readability assessment. Most of these properties are subjective, involving costly and time-consuming surveys to be gathered. Recent approaches use the limited datasets of psycholinguistic properties to extend them automatically to large lexicons. However, some of the resources used by such approaches are not available to most languages. This study presents a method to infer psycholinguistic properties for Brazilian Portuguese (BP) using regressors built with a light set of features usually available for less resourced languages: word length, frequency lists, lexical databases composed of school dictionaries and word embedding models. The correlations between the properties inferred are close to those obtained by related works. The resulting resource contains 26,874 words in BP annotated with concreteness, age of acquisition, imageability and subjective frequency.
Motivation & Objective
- Address the lack of psycholinguistic resources for Brazilian Portuguese, a low-resource language in this domain.
- Overcome the high cost and time required to collect psycholinguistic data through human surveys.
- Develop a scalable method to infer psycholinguistic properties using only basic, accessible linguistic features.
- Create a publicly available, large-scale lexicon of Brazilian Portuguese words annotated with concreteness, age of acquisition, imageability, and subjective frequency.
Proposed method
- Train regressors using a combination of word length, frequency lists, and lexical data from school dictionaries.
- Incorporate pre-trained word embedding models to capture semantic and distributional word representations.
- Use multiple regression models to predict each psycholinguistic property (concreteness, age of acquisition, imageability, frequency) based on the feature set.
- Leverage existing small-scale psycholinguistic datasets as training data to calibrate the regressors.
- Apply the trained models to a large lexicon of Brazilian Portuguese words to infer properties at scale.
- Validate the inferred properties by comparing their correlations with human-annotated data, ensuring reliability.
Experimental results
Research questions
- RQ1Can psycholinguistic properties be accurately inferred for Brazilian Portuguese using only lightweight, accessible linguistic features?
- RQ2How do the correlations between inferred properties and human-annotated data compare to those in related work for other languages?
- RQ3To what extent can word embeddings and frequency data improve the prediction accuracy of psycholinguistic properties in low-resource settings?
- RQ4Is the resulting lexicon of 26,874 words sufficiently reliable for downstream NLP applications like text simplification and readability assessment?
Key findings
- The inferred psycholinguistic properties show correlation levels with human-annotated data that are comparable to those reported in related studies for other languages.
- The method successfully extends psycholinguistic annotations to a large lexicon of 26,874 Brazilian Portuguese words using only basic linguistic features.
- Word embeddings and frequency data significantly improve the predictive performance of the regression models.
- The resulting resource is publicly available and suitable for use in NLP tasks such as text simplification and readability assessment.
- The approach demonstrates feasibility for low-resource languages where large-scale human annotation is impractical.
- The correlations between inferred properties and gold-standard data indicate strong reliability and validity of the method.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.