CSD learns a common component for domain generalization, outperforming existing methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
MCCA extracts shared structure from multiple tensor datasets.
A mixture of common skew-t factor analyzers model is introduced for model-based clustering of high-dimensional data. By assuming common component factor loadings, this model allows clustering to be performed in the presence of a large number of mixture components or when the number of dimensions is too large to be well…
New method improves deep CCA by modeling private components conditionally independent of common factors.
Study pairs of subspaces with or without a common complement in Hilbert spaces.
FMM fails to accurately determine the number of components even with consistent posterior.
A C_k-move is a local move that involves (k+1) strands of a link. A C_k-move is called a C_k^d-move if these (k+1) strands belong to mutually distinct components of a link. Since a C_k^d-move preserves all k-component sublinks of a link, we consider the converse implication: are two links with common k-component sublin…
Multitask learning, i.e. taking advantage of the relatedness of individual tasks in order to improve performance on all of them, is a core challenge in the field of machine learning. We focus on matrix regression tasks where the rank of the weight matrix is constrained to reduce sample complexity. We introduce the comm…
Methods for analysis of principal components in discrete data have existed for some time under various names such as grade of membership modelling, probabilistic latent semantic analysis, and genotype inference with admixture. In this paper we explore a number of extensions to the common theory, and present some applic…
We propose a penalized orthogonal-components regression (POCRE) for large p small n data. Orthogonal components are sequentially constructed to maximize, upon standardization, their correlation to the response residuals. A new penalization framework, implemented via empirical Bayes thresholding, is presented to effecti…
We present a large-scale study of commonality in liquidity and resilience across assets in an ultra high-frequency (millisecond-timestamped) Limit Order Book (LOB) dataset from a pan-European electronic equity trading facility. We first show that extant work in quantifying liquidity commonality through the degree of ex…
In this work we propose a method for reducing the dimensionality of tensor objects in a binary classification framework. The proposed Common Mode Patterns method takes into consideration the labels' information, and ensures that tensor objects that belong to different classes do not share common features after the redu…
Modern biomedical studies often collect multi-view data, that is, multiple types of data measured on the same set of objects. A popular model in high-dimensional multi-view data analysis is to decompose each view's data matrix into a low-rank common-source matrix generated by latent factors common across all data views…
FCPCA fuzzy clusters high-dimensional time series data efficiently.
In this letter, we propose an algorithm for recovery of sparse and low rank components of matrices using an iterative method with adaptive thresholding. In each iteration, the low rank and sparse components are obtained using a thresholding operator. This algorithm is fast and can be implemented easily. We compare it w…
Study on identifying AMP chain graph models under known and unknown component decompositions.
Let and be properly immersed closed locally convex subsets of a Riemannian manifold with pinched negative sectional curvature. Using mixing properties of the geodesic flow, we give an asymptotic formula as for the number of common perpendiculars of length at most from to , count…
The paper uses spectral flow on SPD matrices to analyze multimodal data.
New method calculates number of components in twisted torus links.
Two groups have a common model geometry if they act properly and cocompactly by isometries on the same proper geodesic metric space. The Milnor-Schwarz lemma implies that groups with a common model geometry are quasi-isometric; however, the converse is false in general. We consider free products of uniform lattices in …
The paper analyzes PLS-SVD in high-dimensional data integration, revealing its strengths and limitations.
Proposes joint LCA for multiview data to identify shared and view-specific components.
Study reveals structure of local minima in GMMs, identifying key cluster centers.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
Cardea automates machine learning for EHRs, improving model building efficiency.
In this paper, in following of the first part (which ADF tests using ACI evaluation) has conducted, Time Series (TSs) are analyzed using decomposition analysis. In fact, TSs are composed of four components including trend (long term behavior or progression of series), cyclic component (non-periodic fluctuation behavior…
Principal Components Regression (PCR) is a traditional tool for dimension reduction in linear regression that has been both criticized and defended. One concern about PCR is that obtaining the leading principal components tends to be computationally demanding for large data sets. While random projections do not possess…
The paper improves classification accuracy by leveraging a shared signal across domains in high-dimensional classification.
Differentiable pipeline replaces non-differentiable CAE components for shape optimization.
When estimating finite mixture models, it is common to make assumptions on the mixture components, such as parametric assumptions. In this work, we make no distributional assumptions on the mixture components and instead assume that observations from the mixture model are grouped, such that observations in the same gro…
Many deep reinforcement learning algorithms contain inductive biases that sculpt the agent's objective and its interface to the environment. These inductive biases can take many forms, including domain knowledge and pretuned hyper-parameters. In general, there is a trade-off between generality and performance when algo…
Finite mixture models are statistical models which appear in many problems in statistics and machine learning. In such models it is assumed that data are drawn from random probability measures, called mixture components, which are themselves drawn from a probability measure P over probability measures. When estimating …
HeMPPCAT improves PCA for data with varying noise.
We study large deviations and rare default clustering events in a dynamic large heterogeneous portfolio of interconnected components. Defaults come as Poisson events and the default intensities of the different components in the system interact through the empirical default rate and via systematic effects that are comm…
In many settings, we have multiple data sets (also called views) that capture different and overlapping aspects of the same phenomenon. We are often interested in finding patterns that are unique to one or to a subset of the views. For example, we might have one set of molecular observations and one set of physiologica…
Using recent advances in the econometrics literature, we disentangle from high frequency observations on the transaction prices of a large sample of NYSE stocks a fundamental component and a microstructure noise component. We then relate these statistical measurements of market microstructure noise to observable charac…
We present an approach to model-based hierarchical clustering by formulating an objective function based on a Bayesian analysis. This model organizes the data into a cluster hierarchy while specifying a complex feature-set partitioning that is a key component of our model. Features can have either a unique distribution…
We introduce Gluon Time Series (GluonTS, available at https://gluon-ts.mxnet.io), a library for deep-learning-based time series modeling. GluonTS simplifies the development of and experimentation with time series models for common tasks such as forecasting or anomaly detection. It provides all necessary components and …
Inpatient care is a large share of total health care spending, making analysis of inpatient utilization patterns an important part of understanding what drives health care spending growth. Common features of inpatient utilization measures include zero inflation, over-dispersion, and skewness, all of which complicate st…
ProJIVE integrates multiple data types to explain joint and individual variation.
Develops FGL for better portfolio allocation under common factor influence.
Optimal SGD rates achieved with shuffling, covering non-convex and convex cases.
Eigen component analysis combines quantum mechanics with machine learning for efficient data analysis.
CbMLC improves multi-label classification with noisy labels.
The paper tackles noisy functional data by exploring a multivariate perspective.
A new method reduces data movement in neural network training.
High-dimensional shrinkage risk depends on the default prior for the common scale.
We consider the problem of recovering a common latent source with independent components from multiple views. This applies to settings in which a variable is measured with multiple experimental modalities, and where the goal is to synthesize the disparate measurements into a single unified representation. We consider t…