A new model integrates LSTM and copulas for high-dimensional financial data.
problem Modeling high-dimensional dependencies across financial markets.
method Variational LSTM with regular vine copulas.
result Outperforms benchmarks in cross-market portfolio forecasting.
New graph learning framework outperforms existing methods.
problem Learning effective representations of large graphs with high degree variability.
method Deep hierarchical decompositions and neural network template unrolling over the hierarchy.
result Empirically outperforms state-of-the-art graph classification methods on large social network datasets.
Detects dependencies between high-dimensional data and outcomes.
problem Analyzing educational data with high-dimensional student skills.
method n-TARP clustering to quantify and validate dependencies.
result Valid dependencies between student skills and course grades observed.
The problem of learning the structure of a high dimensional graphical model from data has received considerable attention in recent years. In many applications such as sensor networks and proteomics it is often expensive to obtain samples from all the variables involved simultaneously. For instance, this might involve …
The cohomology of complex irreducible polynomials stabilizes with degree or variables.
problem Understanding the cohomology of complex irreducible polynomials.
method Proving homological stability in cohomology as degree or variables increase.
result Cohomology stabilizes with respect to both degree and number of variables.
New algorithm learns ferromagnetic RBMs efficiently.
problem Learning RBMs with latent variables is hard.
method Greedy algorithm based on influence maximization.
result Ferromagnetic RBMs can be learned efficiently.
Solar algorithm selects variables faster and more accurately in high-dimensional data.
problem Variable selection in high-dimensional data with high accuracy and stability.
method Subsample-ordered least-angle regression (solar) and its coordinate descent generalization (solar-cd) using L0 norm solution path averaging. result Solar selects variables with high accuracy and stability, reducing redundant variable selection.
sgboost reduces variable selection bias in boosting with balanced group selection.
problem Reduces variable selection bias in boosting algorithms.
method Simulation-based approach to balance selection frequencies of base-learners.
result Demonstrates efficacy through simulations and flexible group variable selection.
The paper decouples shrinkage and selection in Bayesian Quantile Regression.
problem Improving prediction accuracy in high-dimensional Bayesian Quantile Regression.
method Two-step procedure: shrinkage through continuous priors, sparsification through SAVS.
result The method reduces bias and provides interpretable variable selection.
Polynomial distribution can be applied to dynamical systems in certain situations. Macroeconomic systems characterized by economic variables such as income and wealth can be modelled similarly using polynomials. We extend our previous work to data regarding income from a more diversified pool of countries, which contai…
In this paper we use wavelet concepts to show that correlation coefficient between two financial data's is not constant but varies with scale from high correlation value to strongly anti-correlation value This studies is important because correlation coefficient is used to quantify degree of independence between two va…
ClimAlign uses deep learning for unsupervised climate downscaling.
problem Downscaling climate variables from coarse to fine scales.
method Unsupervised statistical downscaling using normalizing flows.
result ClimAlign achieves comparable predictive performance to supervised methods.
A new co-clustering model for high-dimensional data reduces parameter complexity.
problem High-dimensional data challenges traditional co-clustering methods.
method Parameter-wise co-clustering model with SEM and Gibbs sampler for estimation.
result The model maintains parsimony while offering more flexibility.
The derivation of statistical properties for Partial Least Squares regression can be a challenging task. The reason is that the construction of latent components from the predictor variables also depends on the response variable. While this typically leads to good performance and interpretable models in practice, it ma…
Maximum likelihood estimation applied to high-frequency data allows us to quantify intermittency in the fluctu- ations of asset prices. From time records as short as one month these methods permit extraction of a meaningful intermittency parameter λ characterising the degree of volatility clustering of asset prices. We…
New objective improves disentanglement of discrete and continuous variables.
problem Learning disentangled representations of high-dimensional data.
method Two-level hierarchical objective to control statistical independence.
result Our objective improves disentanglement of discrete and continuous variables.
A novel GPDA method for high-dimensional functional data.
problem Classification and feature selection challenges in high-dimensional, non-stationary functional data.
method Unified two-layer non-stationary Gaussian process with Ising prior for variable selection and classification.
result Demonstrated superior performance on simulated and proteomics datasets.
ControlSHAP stabilizes Shapley value approximations using control variates.
problem High computational cost of exact Shapley values in blackbox models.
method ControlSHAP uses Monte Carlo control variates to stabilize Shapley value approximations.
result Significant reduction in Monte Carlo variability of Shapley estimates.
We propose nested sequential Monte Carlo (NSMC), a methodology to sample from sequences of probability distributions, even where the random variables are high-dimensional. NSMC generalises the SMC framework by requiring only approximate, properly weighted, samples from the SMC proposal distribution, while still resulti…
The proliferation of models for networks raises challenging problems of model selection: the data are sparse and globally dependent, and models are typically high-dimensional and have large numbers of latent variables. Together, these issues mean that the usual model-selection criteria do not work properly for networks…
Deep learning models perform variably across continents/seasons in land cover mapping.
problem Variability in deep learning model performance across different continents/seasons.
method Clustering techniques on satellite imagery from different continents.
result Model performance varies significantly between different continents/seasons.
Drawing a sample from a discrete distribution is one of the building components for Monte Carlo methods. Like other sampling algorithms, discrete sampling suffers from the high computational burden in large-scale inference problems. We study the problem of sampling a discrete random variable with a high degree of depen…
Study integrable geodesic flows on 2-surfaces with high-degree polynomial first integrals.
problem Integrable geodesic flows on 2-surfaces with high-degree polynomial first integrals.
method Semi-Hamiltonian systems of PDEs and generalized hodograph method.
result Construction of many local explicit and implicit integrable examples with polynomial first integrals of degrees 3, 4, 5.
Paper connects free-energy and low-degree hardness in high-dimensional statistics.
problem High-dimensional statistical inference problems are computationally hard.
method Defines a free-energy criterion and connects it to low-degree hardness.
result Establishes connection between free-energy and low-degree hardness for Gaussian models.
Bayesian inference models power-law graphs with efficient algorithms.
problem Modeling networks with heavy-tailed degree distributions.
method Constructs graphs using BFRY random variables and applies variational Bayesian inference.
result Automatic selection of power law behavior from data.
A new iterative algorithm improves RFDA for high-dimensional data.
problem High-dimensional data challenges conventional FDA and RFDA.
method Iterative sketching-based algorithm with accuracy guarantees.
result Accurate approximations can be achieved with smaller sample sizes.
This study uses Tsallis entropy to analyze diversification and integration in Italian stock market companies.
problem Examining the industrial structure and market reactions of cross-shareholding networks.
method Developed Tsallis entropy approach to model diversification and integration using copulas.
result Entropy analysis reveals insights into market polarisation and fairness.
In this paper we prove the following result: if two 2-dimensional 2-homogeneous rational vector fields commute, then either both vector fields can be explicitly integrated to produce rational flows with orbits being lines through the origin, or both flows can be explicitly integrated in terms of algebraic functions. In…
Generative adversarial networks sample unknown high-dimensional conditional distributions.
problem Sampling from unknown high-dimensional conditional distributions with limited data.
method Generative adversarial networks (GAN) for both sampling and distribution inference.
result GAN effectively samples target conditional distribution with minimal impact on sample quality.
The study constructs balanced and rigid curves on specific types of hypersurfaces and complete intersections.
problem Constructing balanced and rigid curves on Calabi-Yau and general-type complete intersections.
method Balanced and rigid curves are constructed using specific hypersurfaces and complete intersections.
result Rigid curves of various genera and balanced rational curves of high degrees are constructed.
Among the proposed network models, the hidden variable (or good get richer) one is particularly interesting, even if an explicit empirical test of its hypotheses has not yet been performed on a real network. Here we provide the first empirical test of this mechanism on the world trade web, the network defined by the tr…
New model measures changing strength of currency relationships over time.
problem Understanding how currency markets have become more or less synchronized over time.
method Presented a time-varying cointegration model for foreign exchange rates, allowing the loading matrix to change over time.
result Market comovement has strengthened over the past quarter century, but the rate of strengthening has slowed.
New study shows low-degree polynomial algorithms struggle at clause densities close to Fix's.
problem Finding satisfying assignments in random k-SAT formulas at high clause densities.
method Analysis of low-degree polynomial algorithms and a new many-way overlap gap property.
result No efficient algorithms can find satisfying assignments at clause densities close to Fix's.
Directly simulates squared Bessel processes efficiently.
problem Simulating squared Bessel processes accurately and efficiently.
method Two-dimensional Chebyshev expansion for non-central chi-square distribution inverse.
result Accurate and efficient simulation for various degrees of freedom.
Low-degree method fails to predict robust subspace recovery problem.
problem Predicting computational tractability of robust subspace recovery problem.
method Low-degree polynomial framework, anti-concentration properties.
result Low-degree method fails to predict computational tractability of robust subspace recovery problem even up to high degree.
Expander graphs have been a focus of attention in computer science in the last four decades. In recent years a high dimensional theory of expanders is emerging. There are several possible generalizations of the theory of expansion to simplicial complexes, among them stand out coboundary expansion and topological expand…
New method explains computational barriers in high-dimensional statistical models.
problem Understanding detection-recovery gaps in high-dimensional inference.
method Combining algorithmic contiguity and cross-validation reduction to obtain conditional computational lower bounds.
result Mild control of low-degree advantage is sufficient to explain computational barriers for recovery.
We consider unsupervised estimation of mixtures of discrete graphical models, where the class variable corresponding to the mixture components is hidden and each mixture component over the observed variables can have a potentially different Markov graph structure and parameters. We propose a novel approach for estimati…
Fourier analysis improves REINFORCE for binary models.
problem Improving gradient estimation for binary latent variable models.
method Connecting Fourier spectrum of Boolean functions to REINFORCE and developing low-variance unbiased gradient estimators.
result REINFORCE estimates degree-1 Fourier coefficients of a Boolean function.
We consider the problem of predicting an outcome variable using p covariates that are measured on n independent observations, in the setting in which flexible and interpretable fits are desirable. We propose the fused lasso additive model (FLAM), in which each additive function is estimated to be piecewise constant…
A new model corrects SBM's bias for power-law degree networks.
problem SBM's incapability to handle power-law degree distributions.
method Introducing degree decay variables to encode varying degree distributions.
result PLD-SBM approximately preserves the scale-free feature in real networks and corrects SBM's bias.
New approach reduces shape optimization anomalies and improves design quality.
problem Improving global optimization efficiency and avoiding geometrical anomalies in shape optimization.
method Reducing design variables, modeling generative process via probabilistic models, penalizing anomalous designs.
result Abnormal designs are penalized, leading to high-quality designs and improved convergence.
A new fairness metric for decision-making algorithms, conditioning on known fair variables.
problem Fairness issues in decision-making systems.
method Conditional fairness metric, Derivable Conditional Fairness Regularizer (DCFR), adversarial representation.
result Traditional fairness notations are special cases of the new conditional fairness notation.
This paper proposes a parsimoniously time varying parameter vector autoregressive model (with exogenous variables, VARX) and studies the properties of the Lasso and adaptive Lasso as estimators of this model. The parameters of the model are assumed to follow parsimonious random walks, where parsimony stems from the ass…
Proposes a boosting framework for sparsity in grouped covariates.
problem Sparsity and selection bias in grouped covariates.
method Component-wise and group-wise gradient boosting with adjusted degrees of freedom.
result Reduces bias and improves predictability in variable selection.
Statistical query algorithms and low-degree tests are nearly equivalent in high-dimensional hypothesis testing.
problem High-dimensional hypothesis testing and information-computation gaps.
method Analysis of statistical query framework and low-degree polynomials.
result Statistical query algorithms and low-degree polynomials are almost equivalent in power under mild conditions.
EBMs become opaque in high dimensions; LASSO sparsifies them.
problem Reducing complexity and improving interpretability of EBMs in high-dimensional settings.
method Applying LASSO to reweight and remove less relevant terms from EBMs.
result EBMs maintain transparency and fast scoring times with reduced complexity.
The paper introduces a new method to select high-quality clustering solutions in k-means.
problem Selecting the optimal number of clusters in k-means clustering.
method The paper introduces a new method to estimate the degrees of freedom in k-means clustering, which is used for model selection.
result The proposed method for selecting high-quality clustering solutions is competitive and reliable.