Preprint
Machine Learning

Machine Learning and Deep Learning frameworks and libraries for large-scale data mining: a survey

Giang Nguyen(Institute of Chemistry of the Slovak Academy of Sciences), Štefan Dlugolinský(Institute of Chemistry of the Slovak Academy of Sciences), Martin Bobák(Institute of Chemistry of the Slovak Academy of Sciences), Viet Tran(Institute of Chemistry of the Slovak Academy of Sciences), Álvaro López García(Instituto de Física de Cantabria), Ignacio Heredia(Instituto de Física de Cantabria), Peter Malík(Institute of Chemistry of the Slovak Academy of Sciences), Ladislav Hluchý(Institute of Chemistry of the Slovak Academy of Sciences)
January 19, 2019Artificial Intelligence Review867 citations

867

Citations

29

Influential Citations

Artificial Intelligence Review

Venue

2019

Year

Abstract

The combined impact of new computing resources and techniques with an increasing avalanche of large datasets, is transforming many research areas and may lead to technological breakthroughs that can be used by billions of people. In the recent years, Machine Learning and especially its subfield Deep Learning have seen impressive advances. Techniques developed within these two fields are now able to analyze and learn from huge amounts of real world examples in a disparate formats. While the number of Machine Learning algorithms is extensive and growing, their implementations through frameworks and libraries is also extensive and growing too. The software development in this field is fast paced with a large number of open-source software coming from the academy, industry, start-ups or wider open-source communities. This survey presents a recent time-slide comprehensive overview with comparisons as well as trends in development and usage of cutting-edge Artificial Intelligence software. It also provides an overview of massive parallelism support that is capable of scaling computation effectively and efficiently in the era of Big Data.

Analysis

Why This Paper Matters

This survey addresses a critical need in the AI community: navigating the rapidly expanding ecosystem of Machine Learning and Deep Learning frameworks. With the explosion of big data and the increasing complexity of models, choosing the right software stack is a major challenge for both researchers and practitioners. The paper provides a comprehensive, time-slide overview that captures the state of the art at the time of publication, making it an essential reference for anyone involved in large-scale data mining.

The significance of this work lies in its systematic comparison of frameworks across multiple dimensions, including scalability, ease of use, and support for parallel computing. By highlighting trends in development and usage, the authors help readers understand not only what tools are available but also how the field is evolving. This is particularly valuable in a domain where new libraries and updates appear frequently, and where making an informed choice can significantly impact project success.

Moreover, the survey emphasizes the importance of massive parallelism, a key enabler for processing the vast amounts of data generated today. By focusing on this aspect, the paper connects software selection with the underlying hardware capabilities, offering guidance on how to achieve efficient and effective computation in the era of Big Data.

Technical Contributions

The paper's main technical contributions include:

  • Comprehensive Framework Overview: It catalogs a wide range of Machine Learning and Deep Learning frameworks, from established ones like TensorFlow and PyTorch to emerging libraries, providing a structured comparison.
  • Feature Comparison: It compares frameworks based on criteria such as programming language support, model flexibility, training speed, and deployment capabilities.
  • Scalability Analysis: It examines how different frameworks handle large-scale data and distributed computing, highlighting their support for massive parallelism.
  • Trend Identification: It analyzes development and usage trends, noting shifts in community adoption and industry preferences.
  • Open-Source Focus: It emphasizes the role of open-source software from academia, industry, and start-ups, which drives innovation and accessibility.

Results

As a survey, the paper does not present new experimental results but synthesizes existing information. It provides qualitative comparisons and categorizations of frameworks, offering insights into their relative strengths and weaknesses. The authors identify that frameworks like TensorFlow and PyTorch have gained significant traction, while also noting the emergence of specialized libraries for specific tasks. The survey also underscores the growing importance of GPU and distributed computing support, which is crucial for scaling deep learning models to big data. While no quantitative metrics are provided, the structured analysis serves as a practical guide for selecting appropriate tools.

Significance

The broader impact of this survey is substantial. It serves as a foundational reference for researchers entering the field, helping them navigate the complex landscape of AI software. For practitioners, it offers a decision-making framework that can save time and resources when building large-scale data mining systems. By highlighting trends and the importance of parallelism, the paper also informs future development of frameworks, encouraging the creation of more scalable and user-friendly tools. In an era where AI is increasingly applied to real-world problems, such guidance is invaluable. The survey's comprehensive nature and its focus on open-source solutions contribute to the democratization of AI, enabling a wider community to leverage cutting-edge techniques. Overall, this paper helps bridge the gap between algorithmic advances and practical implementation, fostering progress in both research and industry.