The paper improves Lasso inference methods for survey data.
problem Improving inference methods for survey data.
method Extends Lasso inferential methods to survey data.
result Establishes asymptotic validity of inference procedures in survey environments.
Study shows non-systematic bias in customer satisfaction surveys limits data value.
problem Non-systematic bias in customer satisfaction surveys limits data value.
method Used real customer satisfaction survey data of a large retail bank to show the irreducible error and suggest thoughtful survey design methods.
result A thoughtful survey design can reduce non-systematic error in customer satisfaction surveys.
The paper tests the credibility of public and private surveys using linear regression and differential privacy.
problem Ensuring the validity of data analysis results from sample surveys using linear regression.
method Designing an algorithm to test the credibility of surveys and extending it to handle LDP.
result The algorithm achieves optimal estimation error bound for ℓ1 linear regression and reduces sample complexity. Machine learning detects survey validity from user behavior.
problem Detecting valid responses in web surveys.
method Uses mouse activity and machine learning models (LSTM, HMM).
result Predicts survey validity without analyzing specific answers.
New ensemble models classify mouse movement trajectories to assess survey question difficulty.
problem Assessing survey question difficulty based on respondents' interaction data.
method Ensemble models combining semi-metric-based weak learners to classify multivariate functional data.
result Improved survey data quality through better identification of respondent difficulty.
We present a Bayesian framework for estimating the customer lifetime value (CLV) and the customer equity (CE) based on the purchasing behavior deducible from the market surveys on customer purchasing behavior. The proposed framework systematically addresses the challenges faced when the future value of customers is est…
Study shows k-NN regressor consistency in complex survey designs.
problem Lack of consistency results for k-NN regressor in complex survey data. method Analysis of regularity conditions on sampling design and data distribution.
result Consistency of k-NN regressor under complex survey designs. Paper extends conformal prediction to complex survey data.
problem Applying distribution-free prediction intervals to complex survey data.
method Design-based conformal prediction for non-exchangeable data.
result Empirical guarantees of finite-sample coverage for complex survey data.
Surveying machine learning methods for economic forecasting.
problem Improving accuracy of economic forecasts using machine learning.
method Nowcasting, textual data, panel and tensor data, high-dimensional Granger causality tests, time series cross-validation, classification with economic losses.
result Recent advances in machine learning methods enhance economic forecasting accuracy.
The paper proposes a method to assess survey data credibility without needing many samples, regardless of data dimension.
problem Assessing the credibility of survey data across different dimensions.
method Task-based approach and model-specific distance metric for verifying survey data credibility in regression models.
result The sample complexity of the proposed algorithm is independent of the data dimension, making it more efficient.
When considering answering important questions with data, unsupervised data offers extensive insight opportunity and unique challenges. This study considers student survey data with a specific goal of clustering students into like groups with underlying concept of identifying different poverty levels. Fuzzy logic is co…
SaML guides ML models to avoid survey biases.
problem ML models trained on survey data often ignore survey design metadata.
method Nine-step guideline for incorporating survey design metadata in ML lifecycle.
result SaML provides valid population inference from survey data.
The rapid development of computing power and efficient Markov Chain Monte Carlo (MCMC) simulation algorithms have revolutionized Bayesian statistics, making it a highly practical inference method in applied work. However, MCMC algorithms tend to be computationally demanding, and are particularly slow for large datasets…
The aim of this survey is to present some aspects of the Bérard-Besson-Gallot spectral embeddings of a closed Riemannian manifold from their origins in Riemannian geometry to more recent applications in data analysis.
Survey on learning with graph-dependent data, deriving new generalization bounds.
problem Traditional i.i.d. data assumption fails in many real-life applications.
method Collect and analyze graph-dependent concentration bounds, derive generalization bounds.
result New generalization bounds for graph-dependent data.
This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.
problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.
Survey on discovering causal relationships from data.
problem Discover causal relationships from data.
method Modern, continuous optimization methods for structure learning.
result Survey of methods and resources for structure discovery.
Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…
Survey of integrating physics knowledge into machine learning models.
problem Mitigating data shortage and ensuring physical plausibility.
method Combining physics knowledge with machine learning models.
result Summarizes recent works in physics-informed machine learning.
Survey on geometric foundations of data reduction methods.
problem High-dimensional data with intrinsic nonlinear structure.
method Spectral manifold learning methods.
result Derivation and convergence analysis of spectral manifold learning.
Survey on preserving curvature bounds for non-smooth Ricci flow.
problem Preserving curvature bounds for non-smooth initial data in Ricci flow.
method Survey of various weak initial data and preservation of curvature bounds.
result Various curvature lower bounds preserved up to a constant for non-smooth initial data.
Survey of IoT recommendation systems and their limitations.
problem Traditional recommender systems fail to handle IoT data.
method Comprehensive review of IoT recommender systems and techniques.
result Proposes a reference framework for future research.
Survey of data augmentation techniques for time series classification with neural networks.
problem Small datasets in time series recognition.
method Four families of data augmentation: transformation-based, pattern mixing, generative models, and decomposition methods.
result Empirical evaluation of 12 data augmentation methods on 128 datasets.
Survey data imputation methods impact feature selection and importance assessment.
problem Impact of different imputation methods on feature selection and importance assessment in survey data.
method Investigated eight imputation methods (listwise deletion, MICE, missRanger, mixGBoost) and three learners (Random Forest, XGBoost, linear model) in a simulation study.
result Different imputation methods yield varying feature selection and importance assessments.
Survey of deep RL in intelligent transportation systems.
problem Optimizing traffic signals and autonomous driving using deep RL.
method Comprehensive review of deep RL applications in traffic control and autonomous driving.
result Summarizes existing works in deep RL-based transportation applications.
AA extracts archetypes from data for clear feature extraction.
problem Non-convex optimization problem in AA.
method Computational procedure extracting archetypes as convex combinations of data.
result AA offers interpretable representations for high-dimensional data.
PICZL improves photometric redshifts for AGN in all-sky surveys.
problem Challenges in accurately computing photo-z for AGN due to interplay of SMBH and host galaxy emissions.
method PICZL uses an ensemble of CNNs with cross-channel integration of image and catalog data, leveraging Gaussian mixture models.
result PICZL achieves a photo-z variance of 4.5% and outlier fraction of 5.6% on a validation sample of 8098 AGN, outperforming previous methods.
Deep learning models outperform MICE in large survey imputation but with hyperparameter tuning.
problem Comparing deep learning and MICE for missing data imputation in large surveys.
method Extensive simulation studies comparing four machine learning-based MI methods: MICE with classification trees, MICE with random forests, generative adversarial imputation networks, and multiple imputation using denoising autoencoders.
result MICE with classification trees consistently outperforms deep learning methods in terms of bias, mean squared error, and coverage.
Paper develops NN models for diabetes screening using NHANES data.
problem Developing accurate predictive models for diabetes in diverse populations.
method Proposes a neural network framework with survey weights, uncertainty quantification.
result Robust risk score models for diabetes in US population.
Bayesian approach learns causal concepts from diverse social surveys.
problem Inferring causal concepts from heterogeneous data with sparse changes.
method Hierarchical Bayesian model with sequential Monte Carlo sampling.
result Model infers meaningful causal concepts and plausible relations.
Survey combines FL and control for better adaptability and privacy.
problem Combining FL and control for better adaptability and privacy.
method Combining Federated Learning (FL) and control methods.
result Combining FL and control enhances adaptability, scalability, generalization, and privacy.
This paper surveys techniques to personalize federated learning models.
problem Personalized models outperform shared models for some clients, reducing participation.
method Surveys recent research on personalizing federated learning models.
result Personalization techniques improve model performance for individual clients.
Balance corrects biased survey data for more accurate insights.
problem Bias in survey data leads to inaccurate insights and underperforming models.
method Three steps: bias understanding, weight adjustment, and evaluation.
result Corrected data leads to more accurate ML model training and insights.
Method completes mixed matrix from complex surveys with heterogeneous missingness.
problem Recovering a mixed dataframe matrix from complex survey sampling with different missingness patterns.
method Two-stage procedure: logistic regression for missingness modeling, and weighted log-likelihood maximization with low-rank constraint.
result The proposed method achieves sublinear convergence and shows superior performance compared to existing methods.
Survey on learning models for irregularly sampled time series data.
problem Challenges in learning from non-uniformly sampled time series data.
method Survey of recent models and architectures based on temporal discretization, interpolation, recurrence, attention, and structural invariance.
result Significant progress in machine learning for irregularly sampled time series data.
Survey of Gaussian process constraints for modeling expensive data.
problem Modeling expensive data with physical constraints.
method Overview of various Gaussian process constraints and their implementation.
result Discussion of computational challenges introduced by constraints.
Survey on privacy issues in deep learning and proposed solutions.
problem Privacy concerns in deep learning models due to sensitive data.
method Review of existing privacy techniques and gaps in research.
result Identification of test-time inference privacy as a research gap.
Survey of machine learning methods for Windows malware classification.
problem Difficulties in malware classification through data collection, labeling, feature creation, and selection.
method Review of current methods and challenges in malware classification.
result Discussion of constraints and unaddressed problems for machine learning in cybersecurity.
Surveying how to use unlabeled data in federated learning.
problem Costly labeling of data limits FL applications.
method Survey and analyze existing research.
result Potential for using unlabeled data in FL.
Survey of extreme value modeling techniques for insurance.
problem Modeling of insurance industry's extreme events.
method Truncation, tempering, censoring, regression techniques.
result Adapted techniques for insurance applications.
Surveying joint Gaussian graphical models to identify shared structures across domains.
problem Estimating shared structures across different data sources.
method Statistical inference of joint Gaussian graphical models.
result Improved estimation power for high-dimensional data.
Survey updates knowledge on homogeneous Einstein spaces.
problem Understanding homogeneous Einstein spaces.
method Building on previous surveys by Wang and Lauret.
result Current state of homogeneous Einstein spaces.
The paper provides a survey of results related to the "κ-generalized distribution", a statistical model for the size distribution of income and wealth. Topics include, among others, discussion of basic analytical properties, interrelations with other statistical distributions as well as aspects that are of special in…
Deep Neural Networks have shown tremendous success in the area of object recognition, image classification and natural language processing. However, designing optimal Neural Network architectures that can learn and output arbitrary graphs is an ongoing research problem. The objective of this survey is to summarize and …
Survey on submanifolds in nearly Kähler spaces.
problem Understanding submanifolds in nearly Kähler spaces.
method No specific method mentioned, just a survey of existing results.
result Summarizes existing results on submanifolds in nearly Kähler spaces.
As one of the most important types of (weaker) supervised information in machine learning and pattern recognition, pairwise constraint, which specifies whether a pair of data points occur together, has recently received significant attention, especially the problem of pairwise constraint propagation. At least two reaso…
Survey of EEG market and machine learning applications.
problem Improving neurology through data-driven research.
method Comprehensive survey of EEG applications and market.
result Machine learning enhances EEG applications and market growth.
Survey of minimal genus problem progress.
problem Finding the smallest genus surface in 3-manifolds.
method Survey and review of existing research.
result Summarizes recent advancements in the minimal genus problem.