Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

13263952 · Jun 202019922001200920172026
48 results for overlapping vs. disjoint

New study on time series anomaly detection shows overlapping inference improves performance.

problem Heterogeneous evaluation practices and inference procedures in time series anomaly detection.
method Unified training, tuning, and evaluation protocol on TSB-AD benchmark, analyzing overlapping vs. disjoint inference.
result Overlapping inference yields consistent improvements, with average relative gain up to +28%.

New model improves histopathology classification across magnifications.

problem Robust histopathology classification is difficult due to magnification shift.
method Domain-general model using stable sparse embedding signatures.
result Domain-general model outperformed baseline and GAN augmentation.

Community detection is a fundamental problem in machine learning. While deep learning has shown great promise in many graphrelated tasks, developing neural models for community detection has received surprisingly little attention. The few existing approaches focus on detecting disjoint communities, even though communit…

2019-09-26abs ↗pdf ↗

New systemic risk models for banks choosing their group memberships.

problem Analyzing systemic risk for banks in disjoint and overlapping groups.
method Proposed new models with realistic game features, introducing Nash equilibrium for optimal solution.
result Explicit solution for risk allocation and existence/uniqueness of Nash equilibrium.

Overlapping clusters are common in models of many practical data-segmentation applications. Suppose we are given nn elements to be clustered into kk possibly overlapping clusters, and an oracle that can interactively answer queries of the form "do elements uu and vv belong to the same cluster?" The goal is to recov…

2019-10-28abs ↗pdf ↗

Paper proposes a new co-clustering method for overlapping clusters and outliers.

problem Real-world datasets often contain overlaps and outliers in co-clusters.
method Formulated Non-Exhaustive, Overlapping Co-Clustering problem and developed NEO-CC algorithm.
result NEO-CC algorithm effectively captures underlying co-clustering structure of real-world data.

The paper addresses causal estimation for text data with apparent overlap violations.

problem Estimating causal effects from text data with unknown confounders and apparent overlap.
method Uses supervised representation learning to create a representation that preserves confounding information while eliminating predictive information, satisfying overlap assumptions.
result Shows how to obtain robust causal estimation in the presence of apparent overlap violations.

Latent variable models for network data extract a summary of the relational structure underlying an observed network. The simplest possible models subdivide nodes of the network into clusters; the probability of a link between any two nodes then depends only on their cluster assignment. Currently available models can b…

2012-06-27abs ↗pdf ↗

Clustering is one of the most universal approaches for understanding complex data. A pivotal aspect of clustering analysis is quantitatively comparing clusterings; clustering comparison is the basis for many tasks such as clustering evaluation, consensus clustering, and tracking the temporal evolution of clusters. In p…

2017-06-19abs ↗pdf ↗

Hypothesis testing in singular models is fundamentally about identifiable vs. non-identifiable parameters.

problem Testing in singular models is inherently problematic due to non-identifiability and degeneracy of Fisher information.
method Formalized the overlap obstruction and showed that hypotheses over non-identifiable parameters are untestable, while those over identifiable parameters reduce to classical testing.
result Hypotheses over non-identifiable parameters are untestable, while those over identifiable parameters reduce to classical testing.

Solves a triangulation problem by showing minimum tetrahedra equals minimum integral 3-chain.

problem Finding the minimum number of tetrahedra to extend a triangulation of a 2-sphere to a 3-ball.
method Relates the minimum number of tetrahedra to the minimum integral 3-chain norm, proving them equal and showing how to achieve the minimum.
result The minimum number of tetrahedra needed to extend a triangulation of a 2-sphere to a 3-ball equals the minimum integral 3-chain norm.

We give the first provably efficient algorithm for learning a one hidden layer convolutional network with respect to a general class of (potentially overlapping) patches. Additionally, our algorithm requires only mild conditions on the underlying distribution. We prove that our framework captures commonly used schemes …

2018-02-07abs ↗pdf ↗

We consider a class of learning problems regularized by a structured sparsity-inducing norm defined as the sum of l_2- or l_infinity-norms over groups of variables. Whereas much effort has been put in developing fast optimization techniques when the groups are disjoint or embedded in a hierarchy, we address here the ca…

2011-04-11abs ↗pdf ↗

Let S S be a hyperbolic surface. We investigate the topology of the space of all curves on S S which start and end at given points in given directions, and whose curvatures are constrained to lie in a given interval (κ1,κ2) (κ_1,κ_2) . Such a space falls into one of four qualitatively distinct classes, according to whet…

2016-11-28abs ↗pdf ↗

We prove some rigidity theorems for configurations of closed disks. First, fix two collections C\mathcal{C} and C~\tilde{\mathcal{C}} of closed disks in the Riemann sphere C^\hat{\mathbb{C}}, sharing a contact graph which (mostly-)triangulates C^\hat{\mathbb{C}}, so that for all corresponding pairs of intersecting dis…

2013-02-11abs ↗pdf ↗

Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. While naturally cast as a combinatorial optimization problem, variable or feature selection admits a convex relaxation through the regularization by the 1\ell_1-norm. In this paper, we consider situations where we…

2011-09-12abs ↗pdf ↗

U-statistics improve gradient estimation in importance-weighted variational inference.

problem High variance in gradient estimation for importance-weighted variational inference.
method Use U-statistics to average base gradient estimators on overlapping batches of size m, achieving lower variance.
result U-statistic variance reduction leads to modest to significant improvements in inference performance.

Classification with a sparsity constraint on the solution plays a central role in many high dimensional machine learning applications. In some cases, the features can be grouped together so that entire subsets of features can be selected or not selected. In many applications, however, this can be too restrictive. In th…

2014-02-18abs ↗pdf ↗

The family of f-divergences is ubiquitously applied to generative modeling in order to adapt the distribution of the model to that of the data. Well-definedness of f-divergences, however, requires the distributions of the data and model to overlap completely in every time step of training. As a result, as soon as the s…

2019-06-01abs ↗pdf ↗

Study shows how transformers classify symbols without naming them, proving a margin-versus-collision criterion.

problem How transformers classify symbols without naming them.
method Logistic classification analysis of transformer-kernel regime, colored collision graph.
result Decomposes learned predictor into ideal template-level classifier and finite-sample perturbation.

We consider a class of learning problems that involve a structured sparsity-inducing norm defined as the sum of \ell_\infty-norms over groups of variables. Whereas a lot of effort has been put in developing fast optimization methods when the groups are disjoint or embedded in a specific hierarchical structure, we add…

2010-08-31abs ↗pdf ↗

Develops methods for causal inference in longitudinal data.

problem Estimating Individual Treatment Effects (ITEs) in high-dimensional, time-varying data.
method Causal Dynamic Variational Autoencoder (CDVAE) and long-term counterfactual regression framework.
result CDVAE outperforms baselines and improves state-of-the-art models, approaching oracle performance.

A Heegaard splitting of a closed, orientable three-manifold satisfies the disjoint curve property if the splitting surface contains an essential simple closed curve and each handlebody contains an essential disk disjoint from this curve [Thompson, 1999]. A splitting is full if it does not have the disjoint curve proper…

2004-01-28abs ↗pdf ↗

Overlapping clustering problem is an important learning issue in which clusters are not mutually exclusive and each object may belongs simultaneously to several clusters. This paper presents a kernel based method that produces overlapping clusters on a high feature space using mercer kernel techniques to improve separa…

2012-11-29abs ↗pdf ↗

Deconfounding scores improve causal effect estimation with weak overlap.

problem Challenges in causal treatment effect estimation due to weak overlap in high-dimensional data.
method Propose deconfounding scores to preserve identification and target estimation while improving overlap.
result Prognostic scores are overlap-optimal under a broad family of generalized linear models with Gaussian features.

A new method speeds up overlapping group lasso computations.

problem Time-consuming optimization of overlapping group lasso on large-scale problems.
method Non-overlapping statistical approximation to overlapping group lasso.
result The proposed penalty is statistically equivalent to overlapping group lasso.

Proposes a sensitivity framework to handle limited overlap in causal inference.

problem Limited overlap between treated and control groups in observational studies.
method Sensitivity framework based on worst-case confidence bounds on bias introduced by trimming.
result Protects against spurious findings by quantifying uncertainty in regions with limited overlap.

Producing overlapping schemes is a major issue in clustering. Recent proposed overlapping methods relies on the search of an optimal covering and are based on different metrics, such as Euclidean distance and I-Divergence, used to measure closeness between observations. In this paper, we propose the use of another meas…

2012-11-29abs ↗pdf ↗

The study simplifies assessing overlap in logistic regression models using empirical likelihood.

problem Assessing overlap in multidimensional logistic regression models.
method Translation of Silvapulle's condition to empirical likelihood maximization, mechanized with R code.
result Minimal overlapping structures are cataloged in dimensions less than four, providing rules for higher dimensions.

Biomedical documents such as Electronic Health Records (EHRs) contain a large amount of information in an unstructured format. The data in EHRs is a hugely valuable resource documenting clinical narratives and decisions, but whilst the text can be easily understood by human doctors it is challenging to use in research …

2019-12-18abs ↗pdf ↗

We present a new property, the Disjoint Path Concordances Property, of an ENR homology manifold X which precisely characterizes when X times R has the Disjoint Disks Property. As a consequence, X times R is a manifold if and only if X is resolvable and it possesses this Disjoint Path Concordances Property.

2009-03-17abs ↗pdf ↗

Recently, to solve large-scale lasso and group lasso problems, screening rules have been developed, the goal of which is to reduce the problem size by efficiently discarding zero coefficients using simple rules independently of the others. However, screening for overlapping group lasso remains an open challenge because…

2014-10-25abs ↗pdf ↗

Twisting a knot KK in S3S^3 along a disjoint unknot cc produces a twist family of knots {Kn}\{K_n\} indexed by the integers. Comparing the behaviors of the Seifert genus g(Kn)g(K_n) and the slice genus g4(Kn)g_4(K_n) under twistings, we prove that if g(Kn)g4(Kn)<Cg(K_n) - g_4(K_n) < C for some constant CC for infinitely many integers $…

2017-05-29abs ↗pdf ↗

Community detection is a fundamental problem in network analysis which is made more challenging by overlaps between communities which often occur in practice. Here we propose a general, flexible, and interpretable generative model for overlapping communities, which can be thought of as a generalization of the degree-co…

2014-12-10abs ↗pdf ↗

RISA improves VFL by using imputed samples with low uncertainty.

problem Limited overlapping samples constrain VFL performance.
method Imputing non-overlapping samples and using evidence theory to select reliable imputed samples.
result Significant performance gains achieved, especially with limited overlapping samples.