Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · May 199319922001200920172026
48 results for common component

CSD learns a common component for domain generalization, outperforming existing methods.

problem Training models to generalize across unseen domains.
method CSD decomposes the model into a common and specific component, discarding the latter.
result CSD outperforms state-of-the-art domain generalization methods.

A mixture of common skew-t factor analyzers model is introduced for model-based clustering of high-dimensional data. By assuming common component factor loadings, this model allows clustering to be performed in the presence of a large number of mixture components or when the number of dimensions is too large to be well…

2013-07-21abs ↗pdf ↗

New method improves deep CCA by modeling private components conditionally independent of common factors.

problem Discovering latent co-variation in multiview datasets with weak common factors.
method Proposes a novel formulation that models private components conditionally independent of common factors.
result Validates the approach with synthetic and real datasets, showing improved identification of common factors.

Study pairs of subspaces with or without a common complement in Hilbert spaces.

problem Characterize pairs of subspaces with or without a common complement in Hilbert spaces.
method Analyze pairs of subspaces (S, T) in the Grassmann manifold Gr(H) of a Hilbert space H, identifying Delta and Gamma based on the existence of a common complement.
result Delta is open and its connected components are parametrized by dimension and codimension. Gamma is a C^\infty submanifold characterized by dimensions and semi-Fredholm indices.

FMM fails to accurately determine the number of components even with consistent posterior.

problem Determining the number of subpopulations in a data set using FMM.
method Analysis of FMM component-count posterior under model misspecification.
result FMM component-count posterior diverges under model misspecification, contrary to intuition.

A C_k-move is a local move that involves (k+1) strands of a link. A C_k-move is called a C_k^d-move if these (k+1) strands belong to mutually distinct components of a link. Since a C_k^d-move preserves all k-component sublinks of a link, we consider the converse implication: are two links with common k-component sublin…

2012-02-13abs ↗pdf ↗

Multitask learning, i.e. taking advantage of the relatedness of individual tasks in order to improve performance on all of them, is a core challenge in the field of machine learning. We focus on matrix regression tasks where the rank of the weight matrix is constrained to reduce sample complexity. We introduce the comm…

2019-10-27abs ↗pdf ↗

Methods for analysis of principal components in discrete data have existed for some time under various names such as grade of membership modelling, probabilistic latent semantic analysis, and genotype inference with admixture. In this paper we explore a number of extensions to the common theory, and present some applic…

2012-07-11abs ↗pdf ↗

We propose a penalized orthogonal-components regression (POCRE) for large p small n data. Orthogonal components are sequentially constructed to maximize, upon standardization, their correlation to the response residuals. A new penalization framework, implemented via empirical Bayes thresholding, is presented to effecti…

2008-11-25abs ↗pdf ↗

We present a large-scale study of commonality in liquidity and resilience across assets in an ultra high-frequency (millisecond-timestamped) Limit Order Book (LOB) dataset from a pan-European electronic equity trading facility. We first show that extant work in quantifying liquidity commonality through the degree of ex…

2014-06-20abs ↗pdf ↗

In this work we propose a method for reducing the dimensionality of tensor objects in a binary classification framework. The proposed Common Mode Patterns method takes into consideration the labels' information, and ensures that tensor objects that belong to different classes do not share common features after the redu…

2019-02-06abs ↗pdf ↗

FCPCA fuzzy clusters high-dimensional time series data efficiently.

problem Ambiguous clustering of multivariate time series data with overlapping distributions.
method FCPCA based on common principal component analysis.
result FCPCA outperforms existing methods in fuzzy clustering of multivariate time series.

Study on identifying AMP chain graph models under known and unknown component decompositions.

problem Identifying AMP chain graph models with known and unknown chain component decompositions.
method Analyzes conditions for identifiability of AMP models and proposes algorithms for structure recovery.
result Conditions for DAG identifiability in AMP models extend equal variance criteria for Bayes nets.

Let DD^- and D+D^+ be properly immersed closed locally convex subsets of a Riemannian manifold with pinched negative sectional curvature. Using mixing properties of the geodesic flow, we give an asymptotic formula as t+t\to+\infty for the number of common perpendiculars of length at most tt from DD^- to D+D^+, count…

2013-05-06abs ↗pdf ↗

Two groups have a common model geometry if they act properly and cocompactly by isometries on the same proper geodesic metric space. The Milnor-Schwarz lemma implies that groups with a common model geometry are quasi-isometric; however, the converse is false in general. We consider free products of uniform lattices in …

2019-10-21abs ↗pdf ↗

The paper analyzes PLS-SVD in high-dimensional data integration, revealing its strengths and limitations.

problem Understanding the behavior of PLS-SVD in high-dimensional data integration.
method Analysis using random matrix theory and singular value decomposition.
result PLS-SVD exhibits counter-intuitive or limiting behavior in certain regimes and outperforms PCA when detecting common latent subspace.

Proposes joint LCA for multiview data to identify shared and view-specific components.

problem Extracting shared components sequentially from multiview data.
method Formulates a matrix decomposition model with joint and individual structures, proposes a penalty term objective function, and employs a refitting procedure.
result Achieves simultaneous estimation and rank selection for cross covariance.

Study reveals structure of local minima in GMMs, identifying key cluster centers.

problem Identifying optimal cluster centers in non-convex GMM landscapes.
method Analyzing the negative log-likelihood function of GMMs in the population limit.
result Local minima share a common structure that partially identifies true cluster centers.

This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.

problem Identifying pollution sources in rivers for effective environmental management.
method Principal Component Analysis and its modifications, along with Independent Component Analysis and Factor Analysis, are applied to nitrogen pollution source identification.
result PCA and related techniques can be powerful tools for uncovering nitrogen pollution sources in rivers.

Cardea automates machine learning for EHRs, improving model building efficiency.

problem Lack of a trusted, open-source framework for automated machine learning in EHRs.
method Uses FHIR for data structure, AUTOML frameworks for feature engineering, model selection, and tuning, and an adaptive data assembler.
result Demonstrates framework's effectiveness on 5 prediction tasks, highlighting its flexibility and human competitiveness.

The paper improves classification accuracy by leveraging a shared signal across domains in high-dimensional classification.

problem Improving classification accuracy in high-dimensional data with shared signals across domains.
method Transfer learning for linear discriminant analysis, decomposing mean differences into common and domain-specific components.
result Deterministic limits for transfer performance, leading to optimal weights and corrections for bias.

Differentiable pipeline replaces non-differentiable CAE components for shape optimization.

problem Gradient-based optimization is limited by non-differentiable components in CAE workflows.
method Surrogate models replace non-differentiable pipeline components, enabling gradient-based optimization.
result Gradient-based shape optimization possible without differentiable solvers.

When estimating finite mixture models, it is common to make assumptions on the mixture components, such as parametric assumptions. In this work, we make no distributional assumptions on the mixture components and instead assume that observations from the mixture model are grouped, such that observations in the same gro…

2016-06-30abs ↗pdf ↗

Many deep reinforcement learning algorithms contain inductive biases that sculpt the agent's objective and its interface to the environment. These inductive biases can take many forms, including domain knowledge and pretuned hyper-parameters. In general, there is a trade-off between generality and performance when algo…

2019-07-05abs ↗pdf ↗

Finite mixture models are statistical models which appear in many problems in statistics and machine learning. In such models it is assumed that data are drawn from random probability measures, called mixture components, which are themselves drawn from a probability measure P over probability measures. When estimating …

2015-02-23abs ↗pdf ↗

We study large deviations and rare default clustering events in a dynamic large heterogeneous portfolio of interconnected components. Defaults come as Poisson events and the default intensities of the different components in the system interact through the empirical default rate and via systematic effects that are comm…

2013-11-03abs ↗pdf ↗

In many settings, we have multiple data sets (also called views) that capture different and overlapping aspects of the same phenomenon. We are often interested in finding patterns that are unique to one or to a subset of the views. For example, we might have one set of molecular observations and one set of physiologica…

2015-07-14abs ↗pdf ↗

We present an approach to model-based hierarchical clustering by formulating an objective function based on a Bayesian analysis. This model organizes the data into a cluster hierarchy while specifying a complex feature-set partitioning that is a key component of our model. Features can have either a unique distribution…

2013-01-16abs ↗pdf ↗

We introduce Gluon Time Series (GluonTS, available at https://gluon-ts.mxnet.io), a library for deep-learning-based time series modeling. GluonTS simplifies the development of and experimentation with time series models for common tasks such as forecasting or anomaly detection. It provides all necessary components and …

2019-06-12abs ↗pdf ↗

Develops FGL for better portfolio allocation under common factor influence.

problem Sparsity assumption fails for stock returns driven by common factors.
method Integrates graphical models with factor structure to estimate portfolio weights and risk exposure robust to heavy-tailed distributions.
result FGL-based portfolios outperform equal-weighted and Index portfolios in empirical applications.

Eigen component analysis combines quantum mechanics with machine learning for efficient data analysis.

problem Efficiently extracting linearly separable components from complex data.
method Eigen component analysis (ECA) incorporates quantum mechanics principles into linear learning models.
result ECA outperforms classical linear models and can be integrated with deep neural networks.

High-dimensional shrinkage risk depends on the default prior for the common scale.

problem Choosing the default prior for the common scale in high-dimensional shrinkage.
method Using radial-power benchmark to compare variance-flat and standard deviation-flat priors.
result The standard deviation-flat prior has a one-unit asymptotic risk advantage near the origin.