Yonsei University · Psychology
Professor Kristopher Kyle's research lab specializes in second language (L2) writing and spoken language proficiency, with a strong focus on automated text analysis. The lab develops and validates computational tools—such as TAALES and TAALES 2.0—to measure lexical and syntactic complexity in L2 texts, emphasizing psycholinguistic and usage-based perspectives. Key research directions include advancing text analysis indices for lexical sophistication, syntactic complexity, and lexical diversity, with an emphasis on their validity in relation to human judgment and language learning theories. The lab also investigates how these indices can be used to model holistic language proficiency and support second language acquisition research.
Figures are computed from collected data and may differ slightly.
This study explores the construct of lexical sophistication and its applications for measuring second language lexical and speaking proficiency. In doing so, the study introduces the Tool for the Automatic Analysis of LE xical Sophistication ( TAALES ), which calculates text scores for 135 classic and newly developed lexical indices related to word frequency, range, bigram and trigram frequency, academic language, and psycholinguistic word information. TAALES is freely available; runs on Windows
Abstract Syntactic complexity is an important measure of second language (L2) writing proficiency (Larsen–Freeman, 1978; Lu, 2011). Large‐grained indices such as the mean length of T‐unit (MLTU) have been used with the most consistency in L2 writing studies (Ortega, 2003). Recently, indices such as MLTU have been criticized, both for the difficulty in interpretation (e.g., Norris & Ortega, 2009) and for a potentially misplaced focus on clausal subordination (e.g., Biber, Gray, & Poonpon,
This study introduces the second release of the Tool for the Automatic Analysis of Lexical Sophistication (TAALES 2.0), a freely available and easy-to-use text analysis tool. TAALES 2.0 is housed on a user's hard drive (allowing for secure data processing) and is available on most operating systems (Windows, Mac, and Linux). TAALES 2.0 adds 316 indices to the original tool. These indices are related to word frequency, word range, n-gram frequency, n-gram range, n-gram strength of association, co
Syntactic complexity has been an area of significant interest in L2 writing development studies over the past 45 years. Despite the regularity in which syntactic complexity measures have been employed, the construct is still relatively under-developed, and, as a result, the cumulative results of syntactic complexity studies can appear opaque. At least three reasons exist for the current state of affairs, namely the lack of consistency and clarity by which indices of syntactic complexity have bee
Indices of lexical diversity have been used to estimate the size of a writer’s vocabulary and/or a writer’s lexical proficiency for some time. One issue with many commonly used indices of lexical diversity (e.g., TTR and index) is that they vary as a function of text length. Accordingly, much research has been devoted to the development of indices that are text length independent. However, very little research has investigated the degree to which indices of lexical diversity are reflective of hu
Over the past 45 years, the construct of syntactic sophistication has been assessed in L2 writing using what Bulté and Housen (2012) refer to as absolute complexity (Lu, 2011; Ortega, 2003; Wolfe-Quintero, Inagaki, & Kim, 1998). However, it has been argued that making inferences about learners based on absolute complexity indices (e.g., mean length of t-unit and mean length of clause) may be difficult, both from practical and theoretical perspectives (Norris & Ortega, 2009). Furthermore,
Abstract This study conceptualizes lexical sophistication as a multidimensional phenomenon by reducing numerous lexical features of lexical sophistication into 12 aggregated components (i.e., dimensions) via a principal component analysis approach. These components were then used to predict second language (L2) writing proficiency levels, holistic lexical proficiency scores, and longitudinal lexical growth. The results from regression analyses indicated that 5 lexical components (i.e., bigram an
Abstract Measures of syntactic complexity such as mean length of T-unit have been common measures of language proficiency in studies of second language acquisition. Despite the ubiquity and usefulness of such structure-based measures, they could be complemented with measures based on usage-based theories, which focus on the development of not just syntactic forms but also form-meaning pairs, called constructions (Ellis, 2002). Recent cross-sectional research (Kyle & Crossley, 2017) has indic
Abstract Lexical sophistication has been an important indicator of productive lexical proficiency for almost 30 years. Although lexical sophistication has most often been operationalized as the proportion of low frequency words in a text, a growing body of research has indicated that a number of indices such as concreteness, hypernymy, and n‐gram association strengths meaningfully contribute to the construct. While the increase in available indices has expanded our understanding of the multidime
This study investigates relations between second language (L2) lexical input and output in terms of word information properties (i.e., lexical salience; Ellis, 2006a). The data for this study come from a longitudinal corpus of naturalistic spoken data between L2 learners and first language (L1) interlocutors collected over a year's time. The corpus was analyzed using word information properties related to concreteness, familiarity, and meaningfulness to examine word repetitions between input and
This study explores the construct validity of speaking tasks included in the TOEFL iBT (e.g., integrated and independent speaking tasks). Specifically, advanced natural language processing (NLP) tools, MANOVA difference statistics, and discriminant function analyses (DFA) are used to assess the degree to which and in what ways responses to these tasks differ with regard to linguistic characteristics. The findings lend support to using a variety of speaking tasks to assess speaking proficiency. N
Lexical richness is an umbrella term that refers to the density, sophistication, and variety of vocabulary in a language sample such as an essay or spoken response. In language learning contexts, higher proficiency language users are presumed to produce language that is more lexically dense and sophisticated, and that includes a wider variety of lexical items. The focus of this chapter is describing the ways in which lexical richness has been and can be measured. Each of the three categories (de
Open papers in the app to read, cite, and organize with AI.