Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
The paper explores machine learning in mobile big data analysis.
problem Challenges in mobile big data analysis.
method Discussion and review of existing methods.
result Identification of main challenges and future directions.
Study uses big data to analyze quantum invariants.
problem Investigate structural properties of Jones polynomial.
method Exploratory and topological data analysis, including coloring, rank increase, categorification.
result Contrasts behavior of Jones polynomial under various enhancements.
This study designs a financial risk control platform using big data and machine learning.
problem Traditional risk management models are inadequate for modern financial complexities.
method Big data mining, real-time streaming data processing, statistical analysis, and precise customer behavior mining.
result The platform effectively identifies and responds to potential risks in real-time.
The paper proposes a new model using financial big data to improve portfolio risk analysis.
problem Addressing potential information loss in portfolio risk measurement.
method Uses financial big data to incorporate out-of-target-portfolio information and overcomes the curse of dimensionality.
result The use of financial big data improves small portfolio risk analysis.
Astronomy needs efficient machine learning and image analysis for big data.
problem Efficient machine learning and image analysis for big astronomical data.
method Exemplary results, challenges, and methodological advancements in machine learning and image analysis.
result Astronomy pushes the boundaries of data analysis in machine learning.
Deep neural networks automate statistical analysis for big data.
problem Challenges in applying statistical analysis to big data.
method Constructing CNNs for automatic model selection and parameter estimation.
result CNNs demonstrate excellent performance in automatic model selection and estimation.
New method clusters mixed data types using homogeneity analysis.
problem Clustering datasets with a mix of numerical and categorical attributes.
method Homogeneity analysis to determine a Euclidean representation.
result The method is useful for analyzing big datasets with mixed data types.
This paper categorizes and analyzes mobile big data from wireless networks.
problem Understanding social characteristics of mobile big data.
method Categorization and analysis of real wireless cellular network data.
result Highlight several research directions in social computing for mobile big data.
New method improves efficiency analysis with big data.
problem Challenges in detecting inefficiency with big data.
method Post Double LASSO method using Neyman orthogonal moment conditions.
result Improved estimation of efficiency and inefficiency.
Online learning improves big data accuracy quickly.
problem Heterogeneity in big data analysis.
method Online machine learning for big data.
result Online learning converges quickly to batch accuracy.
A new framework combines Spark and deep learning for big data analysis.
problem Efficient big data analysis for AI problems.
method Combines Apache Spark's distributed computing with deep learning's MLP architecture.
result Empirical analysis shows the new framework outperforms traditional methods.
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…
Big data transforms accounting and auditing, enhancing insights but posing challenges.
problem Challenges in data privacy and security with increased data sources.
method Utilizing AI and machine learning for efficient data analysis and anomaly detection.
result Enhanced analytics tools and continuous learning are key to overcoming challenges.
New method classifies p>>n data using sparse nPCA.
problem Classifying p>>n data in medical imaging and genetics.
method Linear discriminant analysis with sparse nPCA for variable selection.
result The new method outperforms traditional methods in experiments.
LSAR efficiently estimates AR models for big time series data.
problem Efficiently analyzing large-scale time series data with high accuracy.
method Developed a fast algorithm to estimate leverage scores and an efficient LSAR algorithm for fitting AR models.
result LSAR algorithm finds maximum likelihood estimates with high probability and improved worst-case running time.
Neuroscience faces new challenges in data analysis as datasets grow richer.
problem How to analyze large, complex neuroscientific datasets effectively.
method Development of non-parametric, generative models combining frequentist and Bayesian approaches.
result New statistical methods will be essential for extracting meaningful insights from neuroscientific data.
Divide-and-conquer method splits large data sets for efficient analysis.
problem Handling large data sets that exceed computational limits.
method Split data into smaller sets, analyze each separately, then combine results.
result Combined results provide statistical inference similar to analyzing entire data set.
Enhances network monitoring with interpretable data analysis.
problem Making data analysis models understandable for network operators.
method Extended MBDA methodology for automatic feature derivation.
result Detects and diagnoses network anomalies with interpretable and interactive models.
Two algorithms improve fitting autoregressive models for big data.
problem Efficiently solving Toeplitz least squares problems for large time series data.
method Applied randomized numerical linear algebra (RandNLA) techniques.
result LSAR algorithm is more robust for real-world time series data.
Ethereum trends analyzed through blockchain transactions and Google searches.
problem Identifying market manipulation in crypto prices.
method Big data analysis of Ethereum transactions, smart contracts, and search volumes.
result Big players manipulate crypto markets after price drops.
PFBP algorithm speeds up feature selection in big data.
problem Feature selection in high-dimensional and/or large sample size data.
method PFBP algorithm partitions data and uses local computations with early decisions.
result Asymptotic optimality for causal networks, super-linear speedup, linear scalability.
Cloud-based BCI predicts epileptic seizures from big EEG data.
problem Nonstationary EEG signals and big data storage/computational challenges.
method Dimensionality reduction, stacked autoencoder, cloud computing.
result Proposes a patient-specific BCI system for real-time seizure prediction.
This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.
problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.
Gradient-HA aligns fMRI data faster and more accurately.
problem Aligning fMRI data from multiple subjects efficiently and accurately.
method Gradient-HA uses ICA and SGA to solve alignment issues.
result Gradient-HA outperforms other methods in big data fMRI analysis.
Big data analytics improves healthcare through early detection and quality life.
problem Limited access to healthcare data hinders evidence-based decision-making.
method Analysis of healthcare data using various tools and techniques.
result Big data analytics enhances healthcare quality and patient outcomes.
This paper analyzes interactive model analysis for machine learning.
problem Understanding, diagnosing, and refining machine learning models.
method Classification of relevant work into understanding, diagnosis, and refinement categories.
result Exploration of future research opportunities in interactive model analysis.
Neural network predicts stock prices using technical indicators.
problem Predicting stock prices for optimal trading.
method Converted financial data into buy-sell-hold signals, trained MLP ANN model with Apache Spark.
result Neural network model performs comparably to Buy and Hold strategy.
Prototype technique for anomaly detection in industrial big data.
problem Real-time condition monitoring and failure analysis of industrial systems.
method Independent ordinations of repeated bootstrapped partitions and inspection of ordinal distances.
result Prototype technique operates at industrial scale for anomaly detection.
A new method for handling imbalanced big data using ensembles and smart data.
problem Imbalanced data distribution in big data scenarios.
method Smart Data driven Decision Trees Ensemble (SD_DeTE) methodology.
result SD_DeTE outperforms Random Forest in handling imbalanced binary classification problems in big data.
New method for tensor completion using nonconvex dual total variation.
problem Tensor completion from partial measurements with exponential-family noise.
method Proposed dual-TV (DTV) regularizers for tensor completion under exponential-family noise.
result Theoretical upper bounds on recovery error for tensor completion.
Develops an efficient online robust PCA method for big data.
problem Efficiency and robustness in processing big data with changing subspaces.
method Online moving window robust principal component analysis (OMWRPCA) with change point detection.
result Successfully tracks both slowly and abruptly changing subspaces and detects change points.
New algorithm speeds up complex model analysis for big data.
problem Scalable analysis of large astronomical datasets with complex models.
method Collaborative Nested Sampling
result Parameter probability distributions derived for each observation with reduced model evaluations.
Improves GP expert combination for big data analysis.
problem Scaling Gaussian process experts for big data.
method Transductive combination of independently trained GP experts, providing theoretical justification for gPoE-GP.
result Empirical validation of an improved combination method over gPoE-GP.
Efficiently augments triplet data for better data analytics.
problem Lack of direct pairwise distance information for data analysis.
method Triplets augmentation to infer hidden information from existing data.
result Improves quality of kernel-based and kernel-free data analytics.
New SDA models for big data analysis using aggregated symbols.
problem Handling large and complex datasets efficiently.
method Developing likelihood functions for symbolic data based on underlying measurement-level data.
result Efficient analysis of big data through reduced distributional summaries.
Quantum SVM clustering speeds up big data analysis.
problem Performance degradation of classical SVM clustering on big data.
method Developed a quantum version of SVM clustering using quantum support vector machine and kernels.
result Significant speed-up gain on run-time complexity.
The paper proposes a simple algorithm for finding low-dimensional manifolds in big data.
problem Finding low-dimensional manifolds in large datasets.
method A novel algorithm for identifying low-dimensional manifolds in high-dimensional data.
result The algorithm effectively finds low-dimensional manifolds in data, outperforming classical methods like PCA and Isomap.
Survey of big data in cyber-physical systems, including data security and green challenges.
problem Managing vast amounts of data in cyber-physical systems.
method Taxonomy and overview of data collection, storage, access, processing, and analysis.
result First panoramic survey on big data for CPS, addressing cybersecurity and green challenges.
Parallelizes Bayesian MCMC for big data, improving efficiency and speed.
problem Efficiently analyzing large Bayesian hierarchical models with big data.
method Two-stage approach: first stage estimates group-specific parameters in parallel, second stage uses stage 1 posteriors as proposals.
result Agrees with full data analysis but with increased efficiency and reduced computation times.
CCA helps find hidden connections in complex biomedical data.
problem Analyzing large, multi-variable datasets in biology and medicine.
method Canonical correlation analysis (CCA) for exploring relationships between two sets of variables.
result CCA uncovers essential hidden associations between diverse data types.
Framework optimizes portfolios using big data from financial markets.
problem Optimizing investment decisions with structured and unstructured financial data.
method 5-stage methodology including DEA, text mining, clustering, ranking, and heuristics for portfolio optimization.
result Helps investors select, weight, and manage assets for informed investment decisions.
Divide data into subsets, analyze each, and recombine results for likelihood function computation.
problem Computing likelihood functions for large and complex data.
method Divide & Recombine (D&R) procedure to estimate density parameters of likelihood model (LM) from MCMC draws.
result The method successfully computes likelihood functions for logistic regression data model.
Proposes a Big Data framework for SC forecasting, including data preprocessing and machine learning.
problem Improving SC forecasting accuracy and efficiency.
method Data collection, preprocessing, machine learning model training, hyperparameter tuning, performance evaluation.
result Optimized SC forecasting models enhance workforce, inventory, and overall SC performance.
The paper improves smoothed analysis for online problems with adaptive adversaries.
problem Online prediction, discrepancy minimization, and online optimization with adaptive adversaries.
method General technique to prove smoothed guarantees against adaptive adversaries, reducing to simpler oblivious adversaries.
result Strong smoothed guarantees for three online problems, matching or improving previous results.
DeepDeath models predict underlying causes of death using big data.
problem Predicting health trajectories in large populations.
method Design and apply two model classes: Hadoop-based ensemble of random forests and DeepDeath (RNN with LSTMs).
result DeepDeath model outperforms N-gram models, learning temporal data aspects.
Unsupervised learning identifies phases and transitions in complex systems.
problem Discovering hidden patterns and phases in large datasets.
method Raw spin configurations were analyzed using principal component analysis and clustering.
result Unsupervised learning successfully identifies physical concepts like order parameter and structure factor.
Bayesian SVARs improve model construction and policy analysis in big data.
problem Manual selection of variables in SVAR models limits their applicability in big data.
method Develops a Bayesian methodology for constructing information sets and retaining the largest system.
result Output increases with housing production over household credit in SVAR models.