[Paper Review] ChatGPT is not all you need. A State of the Art Review of large Generative AI models
A comprehensive state-of-the-art review that taxonomy and analyzes large generative AI models across modalities, outlining key models, developers, applications, and limitations.
During the last two years there has been a plethora of large generative models such as ChatGPT or Stable Diffusion that have been published. Concretely, these models are able to perform tasks such as being a general question and answering system or automatically creating artistic images that are revolutionizing several sectors. Consequently, the implications that these generative models have in the industry and society are enormous, as several job positions may be transformed. For example, Generative AI is capable of transforming effectively and creatively texts to images, like the DALLE-2 model; text to 3D images, like the Dreamfusion model; images to text, like the Flamingo model; texts to video, like the Phenaki model; texts to audio, like the AudioLM model; texts to other texts, like ChatGPT; texts to code, like the Codex model; texts to scientific texts, like the Galactica model or even create algorithms like AlphaTensor. This work consists on an attempt to describe in a concise way the main models are sectors that are affected by generative AI and to provide a taxonomy of the main generative models published recently.
Motivation & Objective
- Provide a concise taxonomy of the main generative AI models
- Analyze each category of models and their applications
- Summarize implications for industry and society across sectors
- Discuss limitations, challenges, and ethical considerations
- Suggest directions for future work and research
Proposed method
- Organize models into a nine-category taxonomy based on input-output mappings
- Describe representative models in each category (e.g., text-to-image, text-to-video, text-to-audio, text-to-text)
- Compare deployment contexts and developer ecosystems across industries
- Highlight non-technical aspects such as data, computation, bias, and ethics
- Exclude deep-dive into underlying architectures to focus on applications and content generation
- Provide a conclusions and future-work section”] ,
- research_questions':['What are the dominant categories of large generative AI models and their input-output mappings?','Which representative models exemplify each category, and who develops them?','What are the key applications and industry/societal implications of these models?','What are the main limitations, risks, and ethical concerns associated with these models?'] ,
- key_findings':['The paper proposes a taxonomy of nine categories of generative AI models organized by input-output mappings.','It covers a wide range of modalities including text-to-image, text-to-3D, image-to-text, text-to-video, text-to-audio, text-to-text, text-to-code, text-to-science, and other models.','Most covered models were released in 2022, with some exceptions (e.g., LaMDA in 2021, Muse in 2023).','Six organizations dominate model deployment, reflecting the need for enormous compute and specialized teams.','Representative models include DALL·E 2, Imagen, Stable Diffusion, Muse, Flamingo, VisualGPT, Dreamfusion, Magic3D, Phenaki, Soundify, AudioLM, Jukebox, Whisper, Codex, Alphacode, Galactica, and Minerva, illustrating broad application areas from art to science.','The paper discusses significant limitations such as data bias, massive data and compute requirements, lack of true understanding, and ethical concerns (e.g., deepfakes in text-to-video).'] ,
- table_headers:[]
- table_rows:[]
Experimental results
Research questions
- RQ1What are the dominant categories of large generative AI models and their input-output mappings?
- RQ2Which representative models exemplify each category, and who develops them?
- RQ3What are the key applications and industry/societal implications of these models?
- RQ4What are the main limitations, risks, and ethical concerns associated with these models?
Key findings
- The paper proposes a taxonomy of nine categories of generative AI models organized by input-output mappings.
- It covers a wide range of modalities including text-to-image, text-to-3D, image-to-text, text-to-video, text-to-audio, text-to-text, text-to-code, text-to-science, and other models.
- Most covered models were released in 2022, with some exceptions (e.g., LaMDA in 2021, Muse in 2023).
- Six organizations dominate model deployment, reflecting the need for enormous compute and specialized teams.
- Representative models include DALL·E 2, Imagen, Stable Diffusion, Muse, Flamingo, VisualGPT, Dreamfusion, Magic3D, Phenaki, Soundify, AudioLM, Jukebox, Whisper, Codex, Alphacode, Galactica, and Minerva, illustrating broad application areas from art to science.
- The paper discusses significant limitations such as data bias, massive data and compute requirements, lack of true understanding, and ethical concerns (e.g., deepfakes in text-to-video).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.