Skip to main content
QUICK REVIEW

[论文解读] The Computer-Linguistic Analysis of Socio-Demographic Profile of Virtual Community Member

Yuriy Syerov, Соломія Федушко|arXiv (Cornell University)|Jan 21, 2019
Information Systems and Technology Applications参考文献 18被引用 12
一句话总结

本文提出了一套计算机语言学框架,通过分析用户数字通信中的语言模式,推断虚拟社区成员的社会人口统计特征。该框架构建了一个正式模型与结构化框架,从文本数据中推导出性别、年龄和职业领域的语言-交际指标,展示了基于训练的分类方法在自动化社会人口统计画像中的系统性应用。

ABSTRACT

This article considers the current problem of investigation and development of computerlinguistic analysis of socio-demographic profile of virtual community member. Webmembers' socio-demographic characteristics' profile validation based on analysis of sociodemographic characteristics. The topicality of the paper is determined by the necessity to identify the web-community member by means of computer-linguistic analysis of their information track. The formal model of basic socio-demographic characteristics of virtual communities' member is formed. The structural model of lingvo-communicative indicators of socio-demographic characteristics of the web-members and common algorithm of the formation of lingvo-communicative indicators based on processing training sample are developed. Types of the computer-linguistic analysis of indicative characteristics are studied and classifications of lingvo-communicative indicators of gender, age and sphere of activities of web-community member is established. Also, the formal model of the basic socio-demographic characteristics of web-communities' member is introduced.

研究动机与目标

  • 解决从数字通信中自动识别虚拟社区成员社会人口统计特征的挑战。
  • 为网络社区成员的基本社会人口统计特征建立正式模型。
  • 基于用户生成内容中的语言模式,建立语言-交际指标的结构化模型。
  • 通过训练样本对性别、年龄和职业领域的语言指标进行分类与系统化。
  • 提供一种可复现的算法,用于从在线社区的文本数据中生成语言-交际指标。

提出的方法

  • 使用从用户文本中提取的语言特征,对社会人口统计特征进行形式化建模。
  • 基于训练样本的文本处理,开发语言-交际指标的结构化模型。
  • 应用计算机语言学分析技术,识别与性别、年龄和职业相关的模式。
  • 通过代表性训练数据的分析,将语言指标分类到结构化类别中。
  • 设计一种通用算法,用于从处理后的文本输入中生成语言-交际指标。
  • 使用训练样本校准并验证语言特征到人口统计类别的分类。

实验结果

研究问题

  • RQ1如何从语言行为中推断虚拟社区成员的社会人口统计特征?
  • RQ2哪些类型的计算机语言学分析在识别性别、年龄和职业领域方面最为有效?
  • RQ3如何系统地分类和建模语言-交际指标以实现人口统计画像?
  • RQ4哪些结构化与形式化模型能够表示语言特征与社会人口统计属性之间的关系?
  • RQ5基于训练的文本数据处理在多大程度上能产生可靠的人口统计推断?

主要发现

  • 成功开发了虚拟社区成员基本社会人口统计特征的正式模型。
  • 建立了语言-交际指标的结构化模型,实现了对语言线索的系统分类。
  • 研究识别并分类了适用于人口统计推断的计算机语言学分析类型。
  • 建立了性别、年龄和职业领域语言-交际指标的分类系统。
  • 通过训练样本处理,验证了生成语言-交际指标的通用算法。
  • 该框架提供了一种可复现的方法,用于利用在线社区的文本数据实现自动化社会人口统计画像。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。