Skip to main content
QUICK REVIEW

[Paper Review] Data Science: Challenges and Directions

Longbing Cao|arXiv (Cornell University)|Jun 28, 2020
Time Series Analysis and ForecastingComputer Science23 references81 citations
TL;DR

The paper surveys data science as a complex, interdisciplinary field, outlining X-complexities and X-intelligence, non-IID data challenges, and directions toward human-like machine intelligence. It argues for systematic, cross-disciplinary approaches to transform data into knowledge and actionable insights.

ABSTRACT

While data science has emerged as a contentious new scientific field, enormous debates and discussions have been made on it why we need data science and what makes it as a science. In reviewing hundreds of pieces of literature which include data science in their titles, we find that the majority of the discussions essentially concern statistics, data mining, machine learning, big data, or broadly data analytics, and only a limited number of new data-driven challenges and directions have been explored. In this paper, we explore the intrinsic challenges and directions inspired by comprehensively exploring the complexities and intelligence embedded in data science problems. We focus on the research and innovation challenges inspired by the nature of data science problems as complex systems, and the methodologies for handling such systems.

Motivation & Objective

  • Characterize data science as a complex system with embedded X-complexities across data, behavior, domain, society, environment, learning, and deliverables.
  • Identify limitations of current theories and methods in handling big data complexities and assumption violations.
  • Propose a framework for X-intelligence and data-to-decision transformation to guide disciplinary development.
  • Highlight non-IID data learning as a core research challenge and explore implications for theory and practice.
  • Discuss human-like machine intelligence prospects within data science and their potential impact on problem solving.

Proposed method

  • Comprehensive literature review to identify intrinsic complexities and intelligence in data science problems.
  • Conceptual framing of data science as complex systems with X-complexities and X-intelligence across multiple facets.
  • Propose a knowledge-to-delivery progression from known to unknown CKI (knowledge, intelligence) states and map problem spaces (Spaces A-D).
  • Introduce a structured landscape with three layers (data input, data-driven discovery, data output) and five research challenges across understanding, foundations, engineering, social issues, and value.
  • Discuss assumption violations (notably non-IID data) and their implications for theory, metrics, and learning.

Experimental results

Research questions

  • RQ1What constitutes data science as a trans-disciplinary field integrating statistics, informatics, computing, and social sciences?
  • RQ2What are the core X-complexities and X-intelligence embedded in data science problems, and how do they affect problem solving?
  • RQ3How do assumption violations, especially non-IID data, challenge current theories and methods in data science?
  • RQ4What strategic directions (data science landscape, learning for non-IID, and human-like intelligence) can advance data science as a discipline?
  • RQ5How can data-to-decision and action-taking processes be designed to effectively transform analytics into decision actions?

Key findings

  • Big data problems are complex systems with embedded X-complexities across data, behavior, domain, society, environment, learning, and deliverables.
  • Non-IID data learning and the need for new theories, algorithms, and metrics are central to advancing data science beyond IID-based methods.
  • A three-layer data science landscape (data input, data-driven discovery, data output) hosts several challenging research areas across understanding, foundations, engineering, social issues, and value.
  • Human-like machine intelligence—driven by curiosity and broader cognitive processes—could transform machine thinking within data science.
  • Assumption violations in big data require rethinking mathematical foundations, modeling, evaluation, and governance to ensure trustworthy, actionable insights.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.