This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
problem Statistical issues with standard financial ratios at industry level.
method Compositional data analysis techniques for financial ratios.
result Improved analysis of financial ratios using compositional data methods.
Develops methods for causal inference in compositional data using instrumental variables.
problem Interpreting summary statistics like diversity indices as causal effects in compositional data.
method Statistical data transformations and regression techniques tailored for compositional data.
result Advantages and limitations of the proposed methods demonstrated on synthetic and real microbiome data.
This paper extends compositional data analysis using graph signal processing.
problem Traditional log-ratios between all variables are not suitable for specific variable relationships.
method Linking compositional data analysis with graph signal processing, it considers only selected log-ratios.
result The approach retains desirable properties of scale invariance and compositional coherence.
New ANN method for imputing rounded zeros in compositional data.
problem Imputing missing values in compositional data with rounded zeros.
method Artificial Neural Networks (ANNs) for imputation of compositional data.
result ANNs are competitive or better than conventional methods for imputing rounded zeros.
New geometric approach for analyzing compositional data like gut microbiomes.
problem Analyzing non-negative compositional data with relative values only.
method Reinterpret compositional data as quotient topology of a sphere, using spherical harmonics and reflection group actions.
result Construction of Reproducing Kernel Hilbert Space (RKHS) for compositional data.
Paper combines geometry and time-series analysis for spatiotemporal data.
problem Multivariate time-series data from multiple sensors.
method Combines manifold learning, Riemannian geometry, and spectral analysis.
result Proposes Riemannian multi-resolution analysis (RMRA) for dynamic mode extraction.
The paper explores using historical data to improve clinical trial analysis by optimizing covariate weights.
problem Limited covariates in small clinical trials reduce the effectiveness of analysis.
method Leverage historical data to pre-specify covariate weights as a composite covariate.
result A composite covariate improves the cost/benefit ratio and reduces overfitting in small clinical trials.
Sharp privacy bounds for sequential analysis of sensitive data.
problem Privacy degradation under sequential analysis of sensitive data.
method Edgeworth expansion in f-differential privacy framework.
result Improved privacy bounds under composition with refined approximation accuracy.
Gaussian process (GP) models provide a powerful tool for prediction but are computationally prohibitive using large data sets. In such scenarios, one has to resort to approximate methods. We derive an approximation based on a composite likelihood approach using a general belief updating framework, which leads to a recu…
We formalize notions of robustness for composite estimators via the notion of a breakdown point. A composite estimator successively applies two (or more) estimators: on data decomposed into disjoint parts, it applies the first estimator on each part, then the second estimator on the outputs of the first estimator. And …
Compositional data have two unique characteristics compared to typical multivariate data: the observed values are nonnegative and their summand is exactly one. To reflect these characteristics, a specific regularized regression model with linear constraints is commonly used. However, linear constraints incur additional…
Efficient logistic regression for aggregated data reduces computation time.
problem Inference for logistic regression models is computationally expensive for large datasets.
method Adapted symbolic data analysis to summarise predictor variables into histograms and use composite likelihoods.
result The method achieves comparable classification rates to full data analysis but at lower computational cost.
New composite indicators reveal hidden relationships between indicators.
problem Subjective aggregation of indicators leads to missed information.
method Used dimensionality reduction techniques (PCA, filtering, clustering) to reveal hidden relationships.
result Cluster-driven composite indicators outperform traditional ones in data reconstruction.
DeepCoDA provides personalized interpretability for complex health data.
problem Interpreting complex health data, especially compositional data, is challenging.
method DeepCoDA framework for high-dimensional compositional data, personalized interpretability through patient-specific weights.
result DeepCoDA maintains state-of-the-art performance and provides coherent, personalized interpretations.
New models for analyzing microbiome data with interactions.
problem Analyzing compositional data with interactions.
method Exponential family models with generalized score matching.
result Effective estimation methods for compositional data with interactions.
CARE method estimates precision matrix for compositional data, achieving optimality in high dimensions.
problem Challenges in inferring conditional dependence relationships in high-dimensional compositional data.
method Composition adaptive regularized estimation (CARE) method for sparse basis precision matrix.
result CARE estimator achieves minimax optimality in high dimensions, performing as well as if the basis were observed.
A folded type model is developed for analyzing compositional data. The proposed model involves an extension of the α-transformation for compositional data and provides a new and flexible class of distributions for modeling data defined on the simplex sample space. Despite its rather seemingly complex structure, emplo…
New financial ratios using compositional data improve analysis of firm health.
problem Statistical issues with standard financial ratios, especially skewness and outliers.
method Compositional data (CoDa) methodology to analyze financial statements.
result Outliers and skewness reduced, results invariant to numerator and denominator permutation.
Proposes a new model for clustering multiplex networks with compositional data.
problem Clustering multiplex networks with multiple types of relations and compositional data.
method Multiplex Dirichlet stochastic block model for compositional networks.
result Validated through simulation and applied to international export data.
Bayesian models predict evolving guest origin markets in tourism.
problem Forecasting the changing composition of guest origin markets in tourism.
method Developed and applied Bayesian Dirichlet autoregressive moving average (BDARMA) models to Airbnb booking data.
result BDARMA models outperform standard benchmarks in forecasting guest origin market shares.
We propose a novel approach for analysis of the composition of an equity mutual fund based on the time series decomposition of the price movements of the individual stocks of the fund. The proposed scheme can be applied to check whether the style proclaimed for a mutual fund actually matches with the fund composition. …
Bayesian model predicts evolving guest origin markets in tourism.
problem Forecasting the changing composition of guest origin markets in tourism.
method Developed and applied Bayesian Dirichlet autoregressive moving average (BDARMA) models to Airbnb booking data.
result BDARMA models achieve lower forecast error and competitive performance in guest origin market shares.
New method tackles composite optimization with error feedback.
problem Challenges in distributed machine learning training and message compression.
method Combines Dual Averaging with EControl for composite optimization.
result First strong convergence analysis for composite optimization with error feedback.
Adapts Altman's model to compositional data for bankruptcy prediction.
problem Predicting business default using standard financial ratios has issues.
method Uses compositional data methodology with log-ratios and machine learning.
result Compositional methods improve predictive performance, especially random forests.
Three supervised learning methods for selecting logratios in compositional data analysis.
problem Selecting logratios for predicting a dependent variable in compositional data.
method Three supervised learning methods: unrestricted search, parts restriction, and additive logratios.
result The first method excels in predictive power, while the other two are more interpretable.
The study uses CoDa to analyze family business financial ratios, highlighting methodological issues.
problem Asymmetry, non-normality, and non-linearity in financial ratios of family businesses.
method Compositional data analysis (CoDa) and classical analysis strategies.
result Results are sensitive to the methodology used, emphasizing the need for appropriate methodologies.
Paper extends FFT-based differential privacy method to heterogeneous compositions.
problem Computing accurate differential privacy guarantees for mixed mechanisms.
method Uses Fast Fourier Transform (FFT) for error analysis and parameter selection.
result Provides tighter bounds for heterogeneous compositions compared to homogeneous cases.
KernelBiome tackles microbiome research by improving predictive performance and interpretability.
problem Challenges in analyzing high-throughput sequencing data, especially in microbiome research.
method KernelBiome is a kernel-based nonparametric regression and classification framework for compositional data, incorporating prior knowledge and capturing complex signals.
result Improved predictive performance compared to state-of-the-art machine learning methods, with two novel quantities for interpretability.
Improved accuracy in dynamic response variation analysis using multi-fidelity data fusion.
problem Inefficient characterization of dynamic response variation due to limited high-fidelity data.
method Composite Neural Network fusion approach for multi-level, heterogeneous datasets.
result Improved accuracy in frequency response variation characterization.
New method recovers relative rates in spatial compositional data from IMS.
problem Challenges in analyzing spatial data from IMS due to competitive sampling.
method Hierarchical Variational Graph Fused Lasso using heavy-tailed graphical lasso prior and automatic differentiation variational inference.
result Our method outperforms state-of-the-practice point estimate methodologies in IMS and has superior posterior coverage.
Data augmentation improves microbiome disease prediction.
problem Improving predictive models for microbiome data.
method Defined novel data augmentation strategies for simplex-valued data.
result Set new state-of-the-art for disease prediction tasks.
Study introduces a variational approach for efficient KL divergence estimation in Dirichlet mixture models.
problem Efficient estimation of KL divergence in Dirichlet mixture models.
method Variational approach for a closed-form solution.
result Superior efficiency and accuracy compared to Monte Carlo methods.
FeDualEx tackles saddle point optimization in federated learning with composite objectives.
problem Saddle point optimization with constraints and non-smooth regularization in federated learning.
method Federated Dual Extrapolation (FeDualEx) algorithm for saddle point optimization and composite objectives.
result FeDualEx effectively solves saddle point optimization problems with composite objectives in federated learning.
Paper analyzes stability and generalization of SCO algorithms.
problem Understanding how SCO algorithms perform on unseen data.
method Algorithmic stability analysis in statistical learning theory.
result Derives dimension-independent excess risk bounds for SCGD and SCSC.
How can neural networks perform so well on compositional tasks even though they lack explicit compositional representations? We use a novel analysis technique called ROLE to show that recurrent neural networks perform well on such tasks by converging to solutions which implicitly represent symbolic structure. This meth…
Improved neural network model for predicting latent budgets in compositional data.
problem Predicting response variables in compositional data with non-negativity constraints.
method LBA-NN, a feed forward neural network model that incorporates K-means clustering for interpretation.
result LBA-NN outperforms traditional LBA in prediction accuracy, specificity, recall, and mean square error.
Develops consistent approximations for composite optimization problems.
problem Significant errors in solutions due to approximations in optimization problems.
method Specifies conditions for well-behaved approximations in minimizers, stationary points, and level-sets for a broad class of composite problems.
result Framework of consistent approximations for composite problems, including stochastic, neural-network, and multi-objective optimization.
Model estimates foreign exchange reserve compositions of undisclosed central banks.
problem Limited information on central bank reserve compositions hinders analysis.
method Hidden Markov Model relating portfolio valuation to exchange rates.
result China's reserve composition likely matches global average, while Singapore holds fewer US dollars.
Unified analysis of multi-task functional linear regression with manifold and composite penalties.
problem Estimating slope functions from functional data with multi-task learning.
method Penalized splines with manifold constraint and composite quadratic penalty.
result Unified convergence upper bound and phase transition behaviors for estimators.
Study reveals clusters of resilient and vulnerable Spanish agri-food firms post-Ukraine-Russia war.
problem Financial resilience of agri-food companies in Spain during the Ukraine-Russia conflict.
method Cluster analysis using centred log-ratios for compositional data of financial ratios.
result Increase in resilient firms by 2023, highlighting sectoral adaptation to economic challenges.
Here we study non-convex composite optimization: first, a finite-sum of smooth but non-convex functions, and second, a general function that admits a simple proximal mapping. Most research on stochastic methods for composite optimization assumes convexity or strong convexity of each function. In this paper, we extend t…
Model predicts composite structures assembly quality with input uncertainty.
problem Accurate prediction of dimensional deviations and residual stress in composite structures assembly.
method Neural Network Gaussian Process considering input uncertainty.
result NNGPIU model outperforms other methods for nonsmooth, nonlinear responses.
We consider in this work a system of two stochastic differential equations named the perturbed compositional gradient flow. By introducing a separation of fast and slow scales of the two equations, we show that the limit of the slow motion is given by an averaged ordinary differential equation. We then demonstrate that…
Edgeworth Accountant calculates privacy loss under differential privacy compositions efficiently.
problem Efficiently computing overall privacy loss under composition of private algorithms.
method Analytical approach using f-differential privacy framework and Edgeworth expansion. result Non-asymptotic (ε,δ)-differential privacy bounds with reduced computational cost. AutoScale improves LLM pre-training by adjusting data mixtures at different scales.
problem Data mixtures that work well at small scales may not perform as well at larger scales.
method AutoScale uses a two-stage approach: fitting a model to predict loss under different compositions and extrapolating optimal compositions to larger scales.
result AutoScale accelerates convergence and improves downstream performance.
A new geometry-preserving method for interpreting compositional data.
problem Statistical challenges in high-dimensional compositional data.
method Geometry-preserving framework for dimension reduction of compositional data.
result Identification of a central compositional subspace for compositional predictors.
New DP-CD method outperforms DP-SGD in solving composite DP-ERM problems.
problem Privacy-preserving machine learning with differential privacy.
method Differentially Private proximal Coordinate Descent (DP-CD) for composite Empirical Risk Minimization (ERM).
result DP-CD outperforms DP-SGD due to larger step sizes and better gradient exploitation.
The study visualizes Spanish fish and meat processing companies using financial, environmental, and social ratios.
problem Mapping financial, environmental, and social performance of Spanish processing companies.
method Used compositional data and principal-component analysis biplot for statistical analysis.
result Identified clusters of companies with similar financial, environmental, and social performance.