New fuzzy clustering method for distribution-valued data using adaptive Wasserstein distances.
problem Clustering distribution-valued data with adaptive weights.
method Fuzzy c-means algorithms using adaptive L2 Wasserstein distances. result Adaptive distances improve clustering of distribution-valued data.
New framework models graph signals as distribution-valued signals in Wasserstein space.
problem Limitations of classical vector-based GSP, including synchronous observations and uncertainty.
method Introduces graph distribution-valued signals (GDSs) in the Wasserstein space.
result GDSs naturally encode uncertainty and stochasticity, generalizing traditional graph signals.
Wasserstein gradient boosting predicts probability distributions for supervised learning.
problem Distribution-valued supervised learning where outputs are probability distributions.
method Fits a new weak learner to Wasserstein gradients of loss functionals of probability distributions.
result Superior performance in probabilistic prediction compared to existing methods.
CAST predicts distribution-valued time series by stabilizing and transporting simplex-supported successors.
problem Forecasting distribution-valued time series with structural failure modes.
method CAST (Causal Anchored Simplex Transport) uses successors retrieved from causal context, stabilized with a persistence anchor, and locally transported on ordered supports.
result CAST outperforms baselines on eleven public and simulated benchmarks, achieving best average rank on both one-step KL and autoregressive rollout JSD.
Study new Ricci bounds for metric measure spaces, preserving properties under time changes.
problem Extend Ricci bounds to non-synthetic spaces and understand their behavior under time changes.
method Introduce distribution-valued lower Ricci bounds BE1(κ,∞), prove equivalence with gradient estimates, and show preservation under time changes. result Distribution-valued Ricci bounds BE1(κ,∞) are preserved under arbitrary time changes and imply sharp gradient estimates. Enhances financial market valuation and trading algorithms using distributional value functions.
problem Accurate valuation and optimal trading decisions in financial markets.
method Combines predictive knowledge and deep reinforcement learning to introduce CDG-Model, a flexible framework for financial market valuation and trading.
result Improved market valuation and trading performance through enhanced feature creation and integration of real-world market factors.
Price movements of stock market are not totally random. In fact, what drives the financial market and what pattern financial time series follows have long been the interest that attracts economists, mathematicians and most recently computer scientists [17]. This paper gives an idea about the trend analysis of stock mar…
This paper improves credit line impact analysis by considering spending as a distribution.
problem Previous studies on credit lines' impact on spending have overlooked the distributional nature of spending.
method Developed a distribution-valued estimator framework to extend existing real-valued estimators.
result Credit lines positively influence spending across all quantiles, but more towards luxuries as they increase.
We show injectivity of the X-ray transform and the d-plane Radon transform for distributions on the n-torus, lowering the regularity assumption in the recent work by Abouelaz and Rouvière. We also show solenoidal injectivity of the X-ray transform on the n-torus for tensor fields of any order, allowing the tensor…
Novel framework for quantifying data distribution values.
problem Quantifying the value of data distributions from samples.
method Generalized Bayesian Inference with loss from transferability measures.
result Unified solution for various practical problems.
New method explains anomalies in multivariate time series data.
problem Understanding and explaining anomalies in multivariate time series data.
method Counterfactual reasoning applied to MDI-detected anomalous intervals.
result Our method accurately identifies and explains anomalies in various extreme events.
Bayesian Additive Distribution Regression (DistBART) predicts distributions from grouped data.
problem Predicting distributions from grouped data with varying characteristics.
method Bayesian nonparametric approach using BART for modeling the regression function.
result Empirical and theoretical evidence supports DistBART's effectiveness in learning from low-dimensional marginals.
Paper develops privacy-preserving mechanisms for machine learning using wavelet transforms.
problem Improper data privacy methods compromise user data even with small preliminary knowledge.
method Three privacy-preserving mechanisms with discrete M-band wavelet transform.
result Successfully retains both differential privacy and learnability in various machine learning environments.
New method predicts spatial events like hurricanes and earthquakes with uncertainty.
problem Quantifying uncertainty in natural hazard predictions.
method Representing spatial point clouds as empirical measures, constraining prediction sets to spatial data manifold, using Wasserstein distance.
result Achieves near-nominal coverage and lower energy/manifold distances compared to baselines.
Estimates low-rank distributional matrices from incomplete samples.
problem Matrix completion for distributional entries with limited observed data.
method Kernel mean embeddings, Tucker rank, functional unfolding operators.
result Effective estimator for distributional matrix completion established.
The paper analyzes the risk of investing in a basket of 27 cryptocurrencies using statistical distributions.
problem Risk assessment of capital allocation in a basket of cryptocurrencies.
method Used statistical tests to determine the most appropriate distribution (SDI) for modeling returns, and adapted the generalized Pareto distribution for tail risk assessment.
result Found that a combination of stable and generalized Pareto distributions provides a more accurate risk assessment for the basket of cryptocurrencies.
New method for multivariate distribution regression using NPT metric.
problem Regression with multivariate distributional responses and Euclidean predictors.
method Fréchet regression with nonparanormal transport (NPT) metric.
result Efficient estimation and granular interpretation of predictor effects.
A new point process for clustering distributions with repulsion.
problem Clustering distributions with repulsion.
method Distributional Determinantal Point Process (dDPP) with sliced Wasserstein kernel.
result Validated dDPP as a well-defined point process and applied to gene expression and epilepsy data.
New SDA models for big data analysis using aggregated symbols.
problem Handling large and complex datasets efficiently.
method Developing likelihood functions for symbolic data based on underlying measurement-level data.
result Efficient analysis of big data through reduced distributional summaries.
DFR models dynamic distributional data with weighted Fréchet means.
problem Regression of distribution-valued responses over time.
method Dynamic Fréchet Regression (DFR) with index-aware weighting and feature selection.
result Improved predictive accuracy and feature recovery over existing methods.
The paper evaluates income credibility using a hierarchical correlation reconstruction technique.
problem Automatic evaluation of credibility of exogenous variables like income based on endogenous variables.
method Adapted hierarchical correlation reconstruction technique for credibility evaluation, combining statistics with machine learning.
result The method allows for the automatic evaluation of credibility of income data, with high density values considered credible.
Optimizes option portfolios for skewed-t returns using VaR and variance measures.
problem Optimizing portfolios for skewed-t returns with heavy tails and skewness.
method Uses variance and VaR measures, departing from normal returns, and provides explicit portfolio weights.
result Optimal portfolio weights differ significantly from variance optimal weights due to skewness.
Deep neural networks forecast financial return distributions accurately.
problem Forecasting probability distributions of financial returns.
method Used 1D CNN and LSTM architectures with custom loss functions to optimize distribution parameters.
result LSTM with skewed Student's t distribution outperformed classical models in multiple evaluation metrics.
The paper constructs random concave functions on the unit simplex.
problem Understanding probability measures on spaces of concave functions.
method Constructing random concave functions via a scaled minimum of random hyperplanes.
result There is a transition from deterministic to non-trivial limiting distributions as the number of hyperplanes increases.
Schrödinger bridge solved with Weyl calculus for quadratic state cost.
problem Optimal control policy to steer joint state statistics.
method Weyl calculus in quantum mechanics for reaction-diffusion PDEs.
result Explicit Markov kernel for quadratic state cost found.
Wasserstein Policy Learning for Distributional Outcomes
problem Offline policy learning with distribution-valued outcomes
method Establishing statistical guarantees for policy learning framework
result Proven leading dependence on N and N-dim(Π) for finite-sample regret
ARL bridges non-Markovian decision processes with reinforcement learning, improving foresight and stability.
problem Inaccurate foresight in non-Markovian environments due to state-based methods' limitations.
method Lifted state space into a signature-augmented manifold, using a self-consistent field approach to anticipate future path-law.
result ARL achieves deterministic evaluation of expected returns with reduced computational complexity and variance.
QR-MIX models joint state-action values as a distribution to handle randomness in MARL.
problem Randomness in rewards and observations leads to randomness in long-term returns in MARL.
method QR-MIX uses quantile regression and combines it with QMIX and IQN to model joint state-action values as a distribution.
result QR-MIX outperforms QMIX in the StarCraft Multi-Agent Challenge (SMAC) environment.
Proposes a new algorithm for click feedback in search results.
problem Learning to predict user clicks based on relevance and position.
method Developed a Bernoulli rank-1 bandit learning problem and proposed Rank1ElimKL to improve performance. result Rank1ElimKL outperforms Rank1Elim in various scenarios, including real-world data.
New algorithms for collaborative reinforcement learning with limited communication.
problem Efficiently learning value functions in multi-agent systems with strict information constraints.
method Distributed gradient-based temporal difference algorithms with consensus schemes.
result Parameter estimates converge to ODEs with defined invariant sets under general assumptions.
Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
Data preprocessing improves data quality for robust data mining.
problem Noisy and incomplete data hinders data mining models.
method Overview of data cleaning, transformation, and preprocessing methods.
result Preprocessing significantly affects data mining model performance.
A new method for handling imbalanced big data using ensembles and smart data.
problem Imbalanced data distribution in big data scenarios.
method Smart Data driven Decision Trees Ensemble (SD_DeTE) methodology.
result SD_DeTE outperforms Random Forest in handling imbalanced binary classification problems in big data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.
Survey on data collection challenges in machine learning.
problem Data scarcity and need for labeled data in machine learning.
method Comprehensive study of data acquisition, labeling, and improvement techniques.
result Identification of research challenges in data collection.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.
Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.
problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.
This paper evaluates how dirty data affects data mining and machine learning results.
problem Negative impacts of dirty data on data mining and machine learning results.
method Experimental comparison of missing, inconsistent, and conflicting data on classification and clustering algorithms.
result Guidelines for algorithm selection and data cleaning based on experimental findings.
DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
This paper introduces C-DSL to improve data mining outcomes by considering context.
problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.
Proposes using probabilistic models for privacy-preserving synthetic data.
problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.
Unlabeled data helps stop active learning better than labeled data.
problem Reducing the need for manual annotation in text classification.
method Compared stopping methods based on labeled, unlabeled, and training data.
result Stopping methods using unlabeled data are more effective.
New test ensures quality of shared data in machine learning.
problem Ensuring quality of external data in machine learning tasks.
method Distribution-free two-sample testing procedures grounded in conformal outlier detection.
result Identifies valuable external data agents for model personalization.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.
A new method classifies multiple correlated data streams simultaneously.
problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.