Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

3026039051,206 · Jun 202019922001200920182026
48 results for level set analysis

Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…

2014-11-21abs ↗pdf ↗

A Python tool generates synthetic data for cluster analysis from high-level descriptions.

problem Creating synthetic data for cluster analysis is laborious and requires detailed geometric parameters.
method Proposes natural language-based synthetic data generation and implements it in a Python package.
result Makes it easy to set up interpretable and reproducible benchmarks for cluster analysis.

The level set tree approach of Hartigan (1975) provides a probabilistically based and highly interpretable encoding of the clustering behavior of a dataset. By representing the hierarchy of data modes as a dendrogram of the level sets of a density estimator, this approach offers many advantages for exploratory analysis…

2013-07-30abs ↗pdf ↗

The clusters of a distribution are often defined by the connected components of a density level set. However, this definition depends on the user-specified level. We address this issue by proposing a simple, generic algorithm, which uses an almost arbitrary level set estimator to estimate the smallest level at which th…

2014-09-30abs ↗pdf ↗

VarFA efficiently estimates student skill levels with uncertainty for adaptive testing.

problem Efficiently estimating student skill levels with uncertainty for adaptive testing.
method VarFA uses variational inference to extend factor analysis models for educational data.
result VarFA efficiently handles large datasets and produces uncertainty estimates.

LOO prediction method improves generalization guarantees for arbitrary datasets.

problem Understanding LOO error guarantees in fully transductive settings for arbitrary datasets.
method Median of Level-Set Aggregation (MLSA) for empirical-risk level sets.
result Multiplicative oracle inequality for LOO error with complexity scaling.

A new framework for discovering macro-level causes and effects from micro-level data.

problem Understanding macro-level causal relations from poorly understood micro-level data.
method Generalizes Chalupka et al. (2015) to discover macro-variable causes and effects from micro-level measurements.
result Identifies multiple levels of causal structure under specific conditions.

The so-called level crossing analysis has been used to investigate the empirical data set. But there is a lack of interpretation for what is reflected by the level crossing results. The fractional Gaussian noise as a well-defined stochastic series could be a suitable benchmark to make the level crossing findings more s…

2011-12-07abs ↗pdf ↗

New method for online learning IC models with node-level feedback.

problem Learning IC models with node-level feedback in social networks.
method Detailed analysis and online algorithm with O(T)\mathcal{O}( \sqrt{T}) cumulative regret.
result First confidence-region result and online algorithm for IC models with node-level feedback.

Develops methods to adjust prediction set coverage based on post-selection analysis.

problem Adjusting prediction set coverage after initial analysis to better fit specific needs.
method Post-selection conformal inference to adjust miscoverage levels.
result Allows for trade-off between coverage and prediction set quality.

The paper explores using set-level ratings for better user-item preference prediction in recommender systems.

problem Capturing user preferences on individual items using set-level ratings.
method Developed collaborative filtering-based methods to model user behaviors in set-level ratings.
result Collaborative filtering-based models can recover and predict user preferences on individual items using set-level ratings.

This paper tackles multilayer graph clustering via convex layer aggregation.

problem Challenges in clustering multilayer graphs and combining information from each layer.
method Theoretical framework for multilayer spectral graph clustering via convex layer aggregation.
result Establishes a critical value on the noise level for reliable cluster separation.

GANs generate DOOM levels similar to human-designed ones.

problem Generating levels similar to human-designed ones in first-person shooter games.
method Extracted features from human-designed levels, trained GANs on these features and level images, generated new levels, compared results.
result GANs can generate levels similar to human-designed ones.

Introduces Motion Programs for better video analysis of human motion.

problem Current video analysis focuses on raw pixels or keypoints, missing higher-level motion primitives.
method Introduces Motion Programs as a neuro-symbolic representation of motions as a composition of high-level primitives.
result Motion Programs accurately describe diverse human motions and improve downstream tasks.

Linear Discriminant Analysis (LDA) is a well-known method for dimensionality reduction and classification. Previous studies have also extended the binary-class case into multi-classes. However, many applications, such as object detection and keyframe extraction cannot provide consistent instance-label pairs, while LDA …

2013-09-21abs ↗pdf ↗

Study on spin random fields using chaos decomposition for cosmic microwave background modeling.

problem Modeling polarization of Cosmic Microwave Background using spin random fields.
method Explicit Wiener-Itô chaos decomposition of area measures of level sets.
result Reveals a clear difference between high frequency regime and zero spin case.

Study robust mean estimation under coordinate-level corruptions using Hamming distance.

problem Robust mean estimation under realistic coordinate-level corruptions.
method Introduce a novel Hamming distance-based measure and present information-theoretic analysis.
result Data cleaning-inspired approaches can match information theoretic bounds for robust mean estimation.

A new algorithm improves Wasserstein discriminant analysis for better data classification.

problem Improving data classification in machine learning.
method Bi-level nonlinear eigenvector algorithm (WDA-nepv) for optimal transport and trace ratio optimizations.
result WDA-nepv enhances classification accuracy and scalability.

Scalable psFA for fMRI data extracts sparse components.

problem Extracting neural representations from fMRI data with probabilistic formulation.
method Group level scalable probabilistic sparse factor analysis (psFA) with spatial sparsity, component pruning, and heteroscedastic noise modeling.
result Sparse components similar to group ICA and reduced noise in activated areas.

New framework analyzes pre-stock jump trading behaviors using multivariate time series analysis.

problem Understanding micro-trading behaviors before stock price jumps.
method Multivariate time series analysis considering temporal information.
result Identifies highly informative attributes for predicting price jumps.

Study uses machine learning to analyze state drug policies and reduce overdose deaths.

problem Epidemic opioid overdose rates and ineffective state-level policies.
method Hierarchical clustering of 138 binomial variables to generate policy bundles, then regression analysis.
result Balancing certain policies leads to reduced overdose deaths, but only after second year.

Constructs bivariate quantiles using vine copulas for multivariate analysis.

problem Need for research in multivariate quantiles, especially for bivariate responses.
method Constructs bivariate (conditional) quantiles using vine copula based bivariate regression model with a novel tree sequence graph structure.
result Avoids typical shortfalls of regression like transformations, interactions, collinearity, and quantile crossings.

The level crossing and inverse statistics analysis of DAX and oil price time series are given. We determine the average frequency of positive-slope crossings, να+ν_α^+, where Tα=1/να+T_α =1/ν_α^+ is the average waiting time for observing the level αα again. We estimate the probability P(K,α)P(K, α), which provides us the probab…

2010-01-25abs ↗pdf ↗

Paper analyzes churn behavior in mobile games at micro and macro levels.

problem Understanding churn behavior in mobile games, especially at micro and macro levels.
method Developed a semi-supervised and inductive embedding model for micro-level churn prediction and constructed a relationship graph for macro-level churn ranking.
result Accurate micro-level churn prediction and macro-level churn ranking were achieved using novel techniques.

Study active learning for multi-level user preferences in recommendation systems.

problem Efficiently learning user preferences through active querying in recommendation systems.
method Proposes a theoretically optimal active learning strategy based on Fisher information matrix for collective matrix factorization.
result Demonstrates strong improvements over active learning methods in personalized, cold-start, and noisy data settings.

Estimates population mean from user-level data with privacy, accounting for heterogeneity.

problem Heterogeneous user data with varying numbers of data points and distributions.
method Simple model of heterogeneous user data, differential privacy mechanism for estimation.
result Asymptotic optimality of the proposed estimator and general lower bounds on error.

New methods test correlation between network structure and node features.

problem Assessing correlation between network structure and node-level covariates.
method Four novel methods based on linear models and canonical correlation analysis.
result Theoretical guarantees and computational efficiency for testing network dependency.

A distributed SGD method for heterogeneous networks with hubs and workers.

problem Learning in heterogeneous multi-level networks with worker heterogeneity and varying communication.
method Multi-Level Local SGD: distributed SGD with hub-and-spoke paradigm and hub averaging.
result The method converges with error dependent on worker heterogeneity, hub network topology, and iterations.

New control on diameter and curvature for evolving surfaces.

problem Controlling the diameter and curvature of evolving surfaces under mean curvature flow.
method Detailed analysis of cylindrical regions under mean curvature flow.
result Intrinsic diameter stays uniformly controlled as surfaces approach first singular time.

The main object of study in the paper is the distance from a point to a line in the Riemannian manifold associated with the Heston model. We reduce the problem of computing such a distance to certain minimization problems for functions of one variable over finite intervals. One of the main ideas in this paper is to use…

2014-09-21abs ↗pdf ↗

In this paper we present a new approach to the study of asymptotically flat static metrics arising in general relativity. In the case where the static potential is bounded, we introduce new quantities which are proven to be monotone along the level set flow of the potential function. We then show how to use these prope…

2015-04-17abs ↗pdf ↗

Develops consistent approximations for composite optimization problems.

problem Significant errors in solutions due to approximations in optimization problems.
method Specifies conditions for well-behaved approximations in minimizers, stationary points, and level-sets for a broad class of composite problems.
result Framework of consistent approximations for composite problems, including stochastic, neural-network, and multi-objective optimization.

Unified view of monotonicity formulas for inverse mean curvature flow and pp-capacitary potentials.

problem Understanding monotonicity formulas for various geometric flows and potentials.
method Refined analysis of pp-capacitary potentials and their level sets.
result Strong convergence of pp-capacitary potentials to inverse mean curvature flow and curvature varifolds.

A key problem in statistical modeling is model selection, how to choose a model at an appropriate level of complexity. This problem appears in many settings, most prominently in choosing the number ofclusters in mixture models or the number of factors in factor analysis. In this tutorial we describe Bayesian nonparamet…

2011-06-14abs ↗pdf ↗