[Paper Review] Thematic Analysis with Open-Source Generative AI and Machine Learning: A New Method for Inductive Qualitative Codebook Development
This paper introduces the GATOS workflow, a novel method using open-source generative AI and machine learning to inductively generate qualitative codebooks for thematic analysis. It demonstrates that the workflow reliably identifies themes in synthetic datasets matching known thematic structures, offering a scalable, valid alternative to manual coding in large-scale social science research.
This paper aims to answer one central question: to what extent can open-source generative text models be used in a workflow to approximate thematic analysis in social science research? To answer this question, we present the Generative AI-enabled Theme Organization and Structuring (GATOS) workflow, which uses open-source machine learning techniques, natural language processing tools, and generative text models to facilitate thematic analysis. To establish validity of the method, we present three case studies applying the GATOS workflow, leveraging these models and techniques to inductively create codebooks similar to traditional procedures using thematic analysis. Specifically, we investigate the extent to which a workflow comprising open-source models and tools can inductively produce codebooks that approach the known space of themes and sub-themes. To address the challenge of gleaning insights from these texts, we combine open-source generative text models, retrieval-augmented generation, and prompt engineering to identify codes and themes in large volumes of text, i.e., generate a qualitative codebook. The process mimics an inductive coding process that researchers might use in traditional thematic analysis by reading text one unit of analysis at a time, considering existing codes already in the codebook, and then deciding whether or not to generate a new code based on whether the extant codebook provides adequate thematic coverage. We demonstrate this workflow using three synthetic datasets from hypothetical organizational research settings: a study of teammate feedback in teamwork settings, a study of organizational cultures of ethical behavior, and a study of employee perspectives about returning to their offices after the pandemic. We show that the GATOS workflow is able to identify themes in the text that were used to generate the original synthetic datasets.
Motivation & Objective
- To investigate whether open-source generative AI models can effectively support inductive qualitative codebook development in social science research.
- To address the scalability challenge of manual thematic analysis when dealing with large volumes of textual data, such as thousands of survey responses or hours of interview transcripts.
- To validate the GATOS workflow by comparing its generated codebooks against known thematic structures in synthetic datasets designed to mimic real-world organizational research contexts.
- To establish a reproducible, transparent, and cost-effective method for thematic analysis that reduces reliance on time-intensive manual coding while maintaining thematic accuracy.
- To explore the potential of retrieval-augmented generation and prompt engineering in enhancing the reliability and coherence of AI-generated qualitative codes and themes.
Proposed method
- The GATOS workflow integrates open-source large language models (LLMs), natural language processing (NLP) tools, and retrieval-augmented generation (RAG) to automate inductive codebook generation.
- The method follows a step-by-step process mirroring traditional thematic analysis: familiarizing with data, generating initial codes, identifying themes, reviewing and refining themes, and producing a hierarchical codebook.
- Prompt engineering is used to guide the LLM through structured coding and theme identification tasks, with explicit instructions based on Braun and Clarke’s six-phase thematic analysis framework.
- The system evaluates code relevance and thematic coherence using reflection prompts that assess theme salience, redundancy, and fit with the research question.
- A JSON-structured output format ensures consistency and machine-readability of the final codebook, including theme names, underlying concepts, associated codes, and relationships.
- The workflow is tested on three synthetic datasets simulating real-world organizational research: teammate feedback, ethical organizational culture, and post-pandemic office return perspectives.

Experimental results
Research questions
- RQ1To what extent can open-source generative AI models approximate the inductive process of human thematic analysis in qualitative research?
- RQ2Can the GATOS workflow reliably identify themes and sub-themes that align with the pre-defined thematic structure of synthetic datasets?
- RQ3How does the integration of retrieval-augmented generation and prompt engineering improve the coherence and thematic accuracy of AI-generated codebooks?
- RQ4What is the performance of the GATOS workflow in identifying meaningful, non-redundant themes across diverse organizational research contexts?
- RQ5How does the method compare to traditional manual coding in terms of scalability, consistency, and thematic fidelity when applied to large textual corpora?
Key findings
- The GATOS workflow successfully identified themes in all three synthetic datasets that closely matched the known thematic structures used to generate the data.
- The AI-generated codebooks demonstrated coherence and thematic relevance, with themes reflecting the underlying research questions and data contexts.
- The method effectively reduced redundant or superfluous code creation through structured reflection prompts that evaluated thematic salience and fit.
- The integration of retrieval-augmented generation enhanced the model’s ability to ground code suggestions in relevant data excerpts, improving contextual accuracy.
- The workflow produced hierarchical codebooks with clear relationships between themes and sub-themes, mirroring the structure of traditionally developed codebooks.
- The results suggest that open-source generative AI, when guided by expert-designed prompts and NLP pipelines, can serve as a valid and scalable alternative to manual thematic coding in large-scale qualitative research.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.