Skip to main content
QUICK REVIEW

[Paper Review] Brainish: Formalizing A Multimodal Language for Intelligence and Consciousness

Paul Pu Liang|arXiv (Cornell University)|Apr 14, 2022
Topic Modeling4 citations
TL;DR

This paper introduces Brainish, a formalized multimodal language integrating words, images, audio, and sensations for modeling intelligence and consciousness in artificial systems. Built upon the Conscious Turing Machine (CTM), Brainish uses unimodal encoders, a coordinated representation space, and decoders to enable multimodal fusion, translation, and generation—demonstrating superior performance on real-world image, text, and audio tasks compared to unimodal learning.

ABSTRACT

Having a rich multimodal inner language is an important component of human intelligence that enables several necessary core cognitive functions such as multimodal prediction, translation, and generation. Building upon the Conscious Turing Machine (CTM), a machine model for consciousness proposed by Blum and Blum (2021), we describe the desiderata of a multimodal language called Brainish, comprising words, images, audio, and sensations combined in representations that the CTM's processors use to communicate with each other. We define the syntax and semantics of Brainish before operationalizing this language through the lens of multimodal artificial intelligence, a vibrant research area studying the computational tools necessary for processing and relating information from heterogeneous signals. Our general framework for learning Brainish involves designing (1) unimodal encoders to segment and represent unimodal data, (2) a coordinated representation space that relates and composes unimodal features to derive holistic meaning across multimodal inputs, and (3) decoders to map multimodal representations into predictions (for fusion) or raw data (for translation or generation). Through discussing how Brainish is crucial for communication and coordination in order to achieve consciousness in the CTM, and by implementing a simple version of Brainish and evaluating its capability of demonstrating intelligence on multimodal prediction and retrieval tasks on several real-world image, text, and audio datasets, we argue that such an inner language will be important for advances in machine models of intelligence and consciousness.

Motivation & Objective

  • To formalize a multimodal inner language—Brainish—that supports core cognitive functions like prediction, translation, and generation in artificial intelligence.
  • To define the syntax and semantics of Brainish as a unified framework for representing heterogeneous sensory inputs (text, images, audio, sensations).
  • To operationalize Brainish through a machine learning pipeline involving unimodal encoders, a shared representation space, and decoders for multimodal tasks.
  • To demonstrate Brainish’s role in enabling communication and coordination within the Conscious Turing Machine (CTM), a model of machine consciousness.
  • To evaluate Brainish on real-world multimodal datasets, showing its superiority over unimodal learning in fusion and co-learning tasks.

Proposed method

  • Design unimodal encoders to segment and represent individual modalities (text, images, audio) using state-of-the-art neural architectures.
  • Construct a coordinated representation space that aligns and composes features across modalities to derive holistic, multimodal meaning.
  • Implement decoders that map multimodal representations to predictions (for fusion) or raw data (for translation and generation).
  • Apply the framework to the CTM to model how internal communication and coordination could support consciousness in artificial systems.
  • Train and evaluate a simplified Brainish model on real-world datasets for image-text retrieval, multimodal prediction, and co-learning tasks.
  • Use contrastive learning and attention mechanisms to align cross-modal representations and improve zero-shot transfer performance.

Experimental results

Research questions

  • RQ1How can a formal multimodal language like Brainish be designed to support core cognitive functions such as multimodal prediction, translation, and generation?
  • RQ2What syntactic and semantic principles are necessary to unify heterogeneous modalities—words, images, audio, and sensations—into a coherent internal representation?
  • RQ3How does a coordinated representation space enable effective fusion, alignment, and co-learning across unimodal inputs in artificial intelligence?
  • RQ4In what way does Brainish enhance communication and coordination within a computational model of consciousness, such as the CTM?
  • RQ5To what extent does multimodal learning with Brainish outperform unimodal learning in real-world multimodal tasks?

Key findings

  • The implemented Brainish model achieved superior performance on multimodal fusion and co-learning tasks compared to unimodal learning, which could not perform alignment at all.
  • Brainish enabled simultaneous learning of multimodal fusion, alignment, and co-learning, demonstrating its capacity for integrated multimodal reasoning.
  • Performance on image-text retrieval and multimodal prediction tasks was consistently higher with Brainish than with unimodal baselines, indicating effective cross-modal representation learning.
  • The model showed strong zero-shot generalization capabilities, suggesting that coordinated representations capture shared semantic structures across modalities.
  • While multimodal generation remains limited due to lack of high-fidelity generators and evaluation metrics, the framework provides a foundation for future development in this area.
  • The framework successfully operationalized Brainish within the CTM, suggesting that such a multimodal language is essential for modeling internal communication in conscious artificial systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.