Preprint
Machine Learning

A survey of machine learning for big data processing

Junfei Qiu(PLA Army Engineering University), Qihui Wu(PLA Army Engineering University), Guoru Ding(PLA Army Engineering University), Yuhua Xu(PLA Army Engineering University), Shuo Feng(PLA Army Engineering University)
May 28, 2016EURASIP Journal on Advances in Signal Processing896 citations

896

Citations

16

Influential Citations

EURASIP Journal on Advances in Signal Processing

Venue

2016

Year

Abstract

There is no doubt that big data are now rapidly expanding in all science and engineering domains. While the potential of these massive data is undoubtedly significant, fully making sense of them requires new ways of thinking and novel learning techniques to address the various challenges. In this paper, we present a literature survey of the latest advances in researches on machine learning for big data processing. First, we review the machine learning techniques and highlight some promising learning methods in recent studies, such as representation learning, deep learning, distributed and parallel learning, transfer learning, active learning, and kernel-based learning. Next, we focus on the analysis and discussions about the challenges and possible solutions of machine learning for big data. Following that, we investigate the close connections of machine learning with signal processing techniques for big data processing. Finally, we outline several open issues and research trends.

Analysis

Why This Paper Matters

This survey, published in 2016, arrives at a critical juncture when big data was becoming ubiquitous across science and engineering. The authors recognize that traditional data analysis methods are insufficient for the scale and complexity of big data, and they systematically review machine learning techniques that can address these challenges. By consolidating a wide range of methods—from deep learning to transfer learning—the paper provides a roadmap for researchers and practitioners navigating the rapidly evolving landscape.

The paper's emphasis on the synergy between machine learning and signal processing is particularly notable. Signal processing has long dealt with large-scale data, and the authors highlight how techniques from both fields can complement each other. This interdisciplinary perspective broadens the appeal of the survey beyond the machine learning community, making it a valuable resource for signal processing engineers and data scientists alike.

Technical Contributions

The paper's primary contribution is its structured taxonomy of machine learning techniques for big data. Key innovations highlighted include:

  • Representation learning: Automatically discovering features from raw data, reducing the need for manual feature engineering.
  • Deep learning: Leveraging hierarchical neural networks to model complex patterns in large datasets.
  • Distributed and parallel learning: Scaling algorithms across multiple machines to handle data that exceeds memory or computational limits.
  • Transfer learning: Reusing knowledge from related domains to improve learning efficiency when labeled data is scarce.
  • Active learning: Selecting the most informative data points for labeling, reducing annotation costs.
  • Kernel-based learning: Using kernel methods to capture non-linear relationships in high-dimensional spaces.

The paper also discusses the challenges of big data, such as data heterogeneity, noise, and scalability, and proposes potential solutions, including online learning and feature selection. Furthermore, it explicitly maps connections between machine learning and signal processing, such as using compressed sensing for dimensionality reduction and Bayesian methods for uncertainty quantification.

Results

As a survey, the paper does not present new experimental results or quantitative comparisons. Instead, its 'results' are the synthesized insights and identified trends. The authors successfully categorize a large body of research and highlight the most promising directions. They also outline open issues, such as the need for interpretable models and privacy-preserving techniques, which have since become major research areas. The paper's impact is evidenced by its 896 citations, indicating its widespread use as a reference.

Significance

The broader impact of this survey lies in its role as a catalyst for interdisciplinary research. By bridging machine learning and signal processing, it encourages cross-pollination of ideas. It also provides a clear framework for understanding the state of the art, which is essential for both newcomers and seasoned researchers. The open issues identified—such as scalability, interpretability, and privacy—have guided subsequent research agendas. In the years since publication, many of the highlighted techniques, especially deep learning, have become standard tools for big data processing, validating the survey's foresight. This paper remains a relevant and frequently cited resource, underscoring its lasting influence on the field.