Prototype technique for anomaly detection in industrial big data.
problem Real-time condition monitoring and failure analysis of industrial systems.
method Independent ordinations of repeated bootstrapped partitions and inspection of ordinal distances.
result Prototype technique operates at industrial scale for anomaly detection.
Neural model learns company embeddings from data and news.
problem Subjective industry classification schemes in finance.
method Multimodal neural model training company embeddings.
result Objective company representations capture nuanced relationships.
Big data transforms accounting and auditing, enhancing insights but posing challenges.
problem Challenges in data privacy and security with increased data sources.
method Utilizing AI and machine learning for efficient data analysis and anomaly detection.
result Enhanced analytics tools and continuous learning are key to overcoming challenges.
PGPs handle big data uncertainty with parametric Gaussian processes.
problem Uncertainty quantification in big data.
method Parametric Gaussian processes designed for big data, avoiding stochastic variational inference.
result Demonstrated effectiveness in handling large datasets with uncertainty quantification.
Survey examines big data's role in network design.
problem Designing robust networks with refined performance and intelligent features.
method Integrates big data analytics with network control/traffic layers.
result Big data analytics can improve network design.
Big data analytics improves healthcare through early detection and quality life.
problem Limited access to healthcare data hinders evidence-based decision-making.
method Analysis of healthcare data using various tools and techniques.
result Big data analytics enhances healthcare quality and patient outcomes.
A new framework combines Spark and deep learning for big data analysis.
problem Efficient big data analysis for AI problems.
method Combines Apache Spark's distributed computing with deep learning's MLP architecture.
result Empirical analysis shows the new framework outperforms traditional methods.
Introduces Quantum Data Center for quantum era benefits.
problem No specific problem stated; focuses on future potential.
method Combines QRAM and quantum networks.
result QDC offers efficiency, security, and precision.
Proposes a methodology to improve data science ROI by addressing key business questions.
problem Companies often fail to maximize data science value, focusing on basic analysis.
method Categorizes and answers 'The Big Three' questions using data science methods.
result Shows how to apply the methodology to real business use cases.
Deep learning text embeddings improve fraud detection in healthcare insurance.
problem Improving fraud detection in healthcare insurance claims.
method Proposed deep learning architectures for text embeddings.
result Our approach outperforms other methods in detecting fraudulent claims.
Paper presents a data preprocessing method for PHM models.
problem Lack of consistent data preprocessing for PHM applications.
method Comprehensive pipeline for sensor data preprocessing.
result Creation of clean data sets for training machinery health state classifiers.
Robinhood users react strongly to overnight price changes and big losers, trading quickly after extreme losses.
problem Understanding trading behavior of Robinhood users, especially in high-frequency trading scenarios.
method Analyzed intraday and overnight price changes, focusing on big losers and gainers.
result Robinhood users react more to overnight price changes and big losers, trading quickly after extreme losses.
In real world industrial applications of topic modeling, the ability to capture gigantic conceptual space by learning an ultra-high dimensional topical representation, i.e., the so-called "big model", is becoming the next desideratum after enthusiasms on "big data", especially for fine-grained downstream tasks such as …
This research tackles unsupervised topic extraction in noisy social media data.
problem Capturing customer insights from social media data is challenging due to noise and heterogeneity.
method The research presents three nonparametric approaches based on the Variational Autoencoder framework: Embedded Dirichlet Process, Embedded Hierarchical Dirichlet Process, and time-aware Dynamic Embedded Dirichlet Process.
result The models achieve equal to better performance than state-of-the-art methods in topic extraction from noisy social media data.
New challenges in causal inference with big data.
problem Extending causal inference to incrementally available observational data.
method Formal definition of continual treatment effect estimation, solutions to challenges.
result New methods for handling challenges in causal inference with observational data.
What is a systematic way to efficiently apply a wide spectrum of advanced ML programs to industrial scale problems, using Big Models (up to 100s of billions of parameters) on Big Data (up to terabytes or petabytes)? Modern parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-…
Interpretable ML helps discover insights from big data.
problem Validating data-driven discoveries from complex datasets.
method Statistical and machine learning techniques for interpretable models.
result Challenges in validating data-driven discoveries remain.
This paper reviews scalable Gaussian processes for big data.
problem Scalability issues in Gaussian process regression with big data.
method Review of scalable Gaussian processes, categorizing them into global and local approximations.
result Comprehensive understanding of scalable Gaussian processes for both academia and industry.
Study tackles imbalanced data in car insurance claims prediction.
problem Predicting rare events (claims) in car insurance with imbalanced data.
method Various machine learning techniques (logistic-regression, decision tree, random forest, xgBoost, feed-forward network) applied to imbalanced dataset.
result Comparison of machine learning algorithms' performance in claim occurrence prediction.
This paper reviews integrating blockchain and machine learning.
problem Secure and efficient data sharing and analysis in blockchain.
method Review of existing research and demonstration of collaboration.
result Blockchain and machine learning can collaborate efficiently and effectively.
Model predicts telecom customer churn with high accuracy.
problem Predicting customers at risk of leaving telecom companies.
method Machine learning and social network analysis on big data platform.
result Model achieved 93.3% AUC, significantly improving churn prediction.
New method classifies nonlinear time series using deep CNNs and bispectra.
problem Classifying nonlinear time series data effectively.
method Combines HOSA with deep CNNs.
result Effective classification of nonlinear time series data.
Paper proposes a method to evaluate SME credit risk using meta paths.
problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.
The paper proposes an AI and IIoT framework for improved maintenance.
problem Current maintenance practices need improvement with AI and IIoT.
method Review of reliability modeling, introduction of Intelligent Maintenance framework, and novel probabilistic deep learning approach.
result Demonstrated novel probabilistic deep learning reliability modelling in Turbofan Engine Degradation Dataset.
Mobile game developers use a scalable churn prediction model to predict player abandonment.
problem Predicting player abandonment in mobile games.
method Survival ensembles approach for accurate churn prediction.
result Accurate predictions on player abandonment and playtime.
Paper tackles deep learning's data needs with rule-based augmentation for cartoon coloring.
problem Deep learning's need for large labeled datasets.
method Rule-based augmentation for small datasets, applied to image translation.
result Automated cartoon coloring with limited data achieved.
DSCOVR improves distributed optimization for big data with less communication and synchronization.
problem Efficiently optimizing large linear models with convex loss functions over distributed systems.
method Randomized primal-dual block coordinate algorithms with doubly stochastic coordinate optimization and variance reduction.
result DSCOVR algorithms require less overall computation and communication compared to other first-order distributed algorithms.
Transfer learning improves machine learning models for equipment diagnostics.
problem Limited model performance due to training data mismatch.
method Transfer learning to reuse knowledge from similar tasks.
result Transfer learning enhances model applicability in PHM.
Method estimates causal effects from incremental data, overcoming missing data challenges.
problem Estimating causal effects from non-stationary, incrementally available observational data.
method Continual Causal Effect Representation Learning
result Method achieves continual causal effect estimation without compromising original data.
NetDP predicts loan defaults using network data, addressing cold-start issues.
problem Cold-start problem in default prediction for new users.
method Combines unsupervised and supervised network representations, using parameter-server for scalability.
result Effectiveness in cold-start problem, especially for new users.
The rise of Big Data has led to new demands for Machine Learning (ML) systems to learn complex models with millions to billions of parameters, that promise adequate capacity to digest massive datasets and offer powerful predictive analytics thereupon. In order to run ML algorithms at such scales, on a distributed clust…
Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
AI enhances bank credit risk management through deep learning and data analysis.
problem Inaccurate credit decisions and potential risks in bank credit risk management.
method Innovative application of AI technology, including deep learning and big data analysis.
result AI provides more accurate and comprehensive credit decision support, reducing risks and losses.
This paper categorizes and analyzes mobile big data from wireless networks.
problem Understanding social characteristics of mobile big data.
method Categorization and analysis of real wireless cellular network data.
result Highlight several research directions in social computing for mobile big data.
LightGCNet simplifies AI for soft sensors, reducing complexity and training time.
problem Complex and resource-intensive deep learning models for soft sensors.
method LightGCNet uses compact angle constraints and node pool strategy for efficient learning.
result LightGCNet achieves small network size, fast learning, and good generalization.
This research improves debt collection strategies using advanced machine learning.
problem Accurate estimation of propensity to pay and cashflow for optimal debt collection.
method Developed a machine learning framework with pre-processing and model selection.
result The proposed model outperforms current industry strategies.
Paper analyzes electricity price and demand data to detect cyber-attacks using time series methods.
problem Detecting cyber-attacks in electricity price and demand data.
method Time series analysis, including moving average, moving standard deviation, and augmented Dickey-Fuller test.
result Identified anomalies in the data using time-series stationary criteria.
Deep learning faces adoption challenges in business analytics.
problem Adoption of deep learning in business analytics is hindered by various factors.
method Empirical study based on three industry use cases.
result Gradient boosting is recommended for structured datasets in business analytics.
DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
New problems on NNSC fill-ins for Bartnik data in high dimensions.
problem Conditions for (n−1)-dimensional Bartnik data to be NNSC-cobordant. method Formulating three problems related to nonnegative scalar curvature fill-ins.
result Conditions for (n−1)-dimensional Bartnik data to be NNSC-cobordant. InfDetect detects e-commerce insurance fraud using graph analysis.
problem Detecting fraudulent claims in e-commerce insurance with multiple parties involved.
method Developed a large-scale fraud detection system InfDetect using graph-based approaches.
result InfDetect successfully detected thousands of fraudulent claims and saved money daily.
Sabrina integrates financial data and domain knowledge for better visualization.
problem Scattered financial data across various sources makes it hard for analysts to understand the economy.
method Sabrina uses a pipeline to fuse firm-specific and macroeconomic data, visualizing it in a unified interface.
result Sabrina aids financial analysts in their analysis process, as shown in a user study.
Study assesses 'big data' in materials science, highlighting challenges.
problem Understanding what constitutes 'big data' in materials science.
method Selected examples of machine learning models, data quality, and infrastructure requirements.
result Big data presents unique challenges in materials science.
Big Data classifiers perform similarly to Small Data classifiers, suggesting scalability tradeoffs.
problem Comparing Big Data classifiers to Small Data classifiers for performance and scalability.
method Empirical study comparing Big Data classifiers to Small Data classifiers.
result Big Data classifiers are slightly inferior but catching up with Small Data classifiers.
Online learning improves big data accuracy quickly.
problem Heterogeneity in big data analysis.
method Online machine learning for big data.
result Online learning converges quickly to batch accuracy.
The paper explores machine learning in mobile big data analysis.
problem Challenges in mobile big data analysis.
method Discussion and review of existing methods.
result Identification of main challenges and future directions.
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…