Skip to main content
QUICK REVIEW

[Paper Review] Can I say, now machines can think?

Nitisha Aggarwal, Geetika Jain Saxena|arXiv (Cornell University)|Jul 11, 2023
Computability, Logic, AI Algorithms5 citations
TL;DR

This paper investigates whether modern AI systems, particularly large language models like GPT and Bard, can be considered to 'think' by revisiting Turing's Imitation Game and evaluating current AI capabilities. It argues that while machines do not yet possess human-like consciousness, they exhibit human-like reasoning, adaptability, and performance on cognitive benchmarks—suggesting they are approaching or passing the Turing Test in practice, though ethical and existential risks remain significant.

ABSTRACT

Generative AI techniques have opened the path for new generations of machines in diverse domains. These machines have various capabilities for example, they can produce images, generate answers or stories, and write codes based on the "prompts" only provided by users. These machines are considered 'thinking minds' because they have the ability to generate human-like responses. In this study, we have analyzed and explored the capabilities of artificial intelligence-enabled machines. We have revisited on Turing's concept of thinking machines and compared it with recent technological advancements. The objections and consequences of the thinking machines are also discussed in this study, along with available techniques to evaluate machines' cognitive capabilities. We have concluded that Turing Test is a critical aspect of evaluating machines' ability. However, there are other aspects of intelligence too, and AI machines exhibit most of these aspects.

Motivation & Objective

  • To assess whether contemporary AI systems, especially large language models, can be considered capable of 'thinking' in a manner comparable to humans.
  • To revisit Alan Turing’s Imitation Game and evaluate its relevance in light of recent advancements in generative AI.
  • To analyze the cognitive capabilities of modern AI, including reasoning, language generation, and adaptability, in comparison to human intelligence.
  • To examine societal risks associated with increasingly intelligent AI systems, including deception, bias, and existential threats.
  • To evaluate whether current benchmarks and tests, such as SQuAD and GLUE, provide valid measures of machine intelligence.

Proposed method

  • Re-evaluates Turing’s 1950 Imitation Game as a benchmark for machine intelligence, focusing on the indistinguishability of machine responses from human ones.
  • Analyzes the performance of state-of-the-art large language models (e.g., GPT-3.5, GPT-4, Bard) on natural language understanding and generation tasks.
  • Compares AI capabilities—such as logical inference, memory, attention, and multimodal processing—against human cognitive functions.
  • Reviews empirical results from AI benchmarks like SQuAD, GLUE, and medical/academic exams (e.g., Stanford Medical School, bar exam) to assess reasoning and knowledge retention.
  • Examines the role of generative adversarial networks (GANs) and multimodal models (e.g., DALL-E, Stable Diffusion) in enabling creative and contextually appropriate outputs.
  • Discusses emerging technologies like quantum computing and robotics as enablers of more advanced, physically embodied AI systems.
Figure 1: Conceptual answer by ChatGPT (ChatGPT May 24 Version).
Figure 1: Conceptual answer by ChatGPT (ChatGPT May 24 Version).

Experimental results

Research questions

  • RQ1Can modern large language models pass the Turing Test in practice, despite being explicitly aware of their non-human identity?
  • RQ2To what extent do contemporary AI systems exhibit human-like reasoning, creativity, and adaptability in real-world tasks?
  • RQ3What are the key cognitive capabilities demonstrated by modern AI, and how do they compare to human intelligence across dimensions like memory, attention, and inference?
  • RQ4What are the societal and existential risks posed by increasingly intelligent AI systems, and how do they relate to Turing’s original concerns?
  • RQ5Are existing evaluation frameworks like SQuAD and GLUE sufficient to measure the true intelligence of AI systems, or do they fall short?

Key findings

  • Large language models such as GPT-4 and Bard have achieved performance levels comparable to or exceeding those of human students on standardized exams, including the uniform bar exam and clinical reasoning tests.
  • AI models like GPT-4 have scored highly on benchmarks such as SQuAD and GLUE, indicating strong performance in natural language understanding and generation.
  • Machines can now generate human-like responses in conversation, write stories, compose poetry, and suggest code improvements, often outperforming mediocre human performance.
  • Some AI systems, such as Bard, have been reported by engineers to display sentiment-like behavior, though this remains contested and may reflect sophisticated prompting rather than genuine emotion.
  • Generative AI models demonstrate multimodal capabilities and multitasking, suggesting progress toward artificial general intelligence, though they remain narrow in scope.
  • Despite high performance on benchmarks, no consensus exists on whether machines truly 'think'—highlighting the lack of a universally accepted definition of machine intelligence.
Figure 2: Image generated by the prompt ’A 14th-century girl working on a desktop in her room’.
Figure 2: Image generated by the prompt ’A 14th-century girl working on a desktop in her room’.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.