Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

105209314418 · Jun 202019922001200920172026
48 results for statistical identifiability

Paper establishes identifiability conditions for a model with two latent vectors and auxiliary data.

problem Identifying conditions for a statistical model with two latent vectors and auxiliary data.
method Proposes a statistical model with two latent vectors and auxiliary data, establishing various identifiability conditions.
result Identifiability conditions reveal a dimensionality relation and link model indeterminacies to maximum link weights.

We identify action representations from video data, proving their statistical benefits.

problem Identifying latent action policies from video data.
method Entropy-regularized LAPO objective, formalizing desiderata for action representations.
result Entropy-regularized LAPO identifies action representations satisfying desiderata under suitable conditions.

New research shows LLMs can't be explained by statistical generalization alone.

problem Understanding why large language models (LLMs) perform well despite statistical generalization limitations.
method Examined the non-identifiability of AR probabilistic models and their implications for LLMs.
result Non-identifiability of LLMs leads to different behaviors and requires a separate theoretical explanation.

Paper uses algebraic signatures to identify probabilistic structures in empirical data.

problem Identifying probabilistic structure from observed binomials in empirical probability tensors.
method Treating vanishing binomials as algebraic signatures, matching signatures to identify models without parameter estimation.
result The method successfully identified rank-one structures in real language data, revealing interpretable sets of words.

Develops a framework for identifying mispriced assets through attention factors for statistical arbitrage.

problem Identifying mispriced assets in statistical arbitrage trading.
method Uses conditional latent factors learned from firm characteristic embeddings to identify time-series signals and form a trading strategy.
result Achieves an out-of-sample Sharpe ratio above 4 on the largest U.S. equities over a 24-year period.

Unified framework for singular statistical models using observable charts.

problem Non-identifiability and breakdown of classical asymptotic theory in singular models.
method Invariant framework based on observable charts to define local coordinate systems in model space.
result Observable order provides a lower bound on KL divergence vanishing rate in singular models.

New method identifies latent relationships in deep models without additional constraints.

problem Latent representations in deep latent variable models are not statistically identifiable.
method Identifies relationships between latent variables (distances, angles, volumes) under mild model conditions.
result Empirically demonstrates more reliable latent distances without additional labeled data.

Unified framework for disentangled representations using mechanistic independence.

problem Identifiability of disentangled latent factors under statistical dependencies.
method Introduces mechanistic independence to characterize latent factors by their actions on observed variables, proposing various independence criteria.
result Establishes conditions for identifiability of latent subspaces without statistical assumptions.

Hypothesis testing in singular models is fundamentally about identifiable vs. non-identifiable parameters.

problem Testing in singular models is inherently problematic due to non-identifiability and degeneracy of Fisher information.
method Formalized the overlap obstruction and showed that hypotheses over non-identifiable parameters are untestable, while those over identifiable parameters reduce to classical testing.
result Hypotheses over non-identifiable parameters are untestable, while those over identifiable parameters reduce to classical testing.

ICCNLS models complex relationships as convex and concave components.

problem Complex input-output relationships with affine ambiguity.
method Sub-gradient constrained affine functions, global orthogonality constraints, L1, L2, and elastic net regularisation.
result Improved predictive accuracy and model simplicity compared to conventional methods.

New model identifies cell-specific genes for cancer prognosis.

problem No statistical model to integrate multiscale cancer data.
method Bayesian generalized promotion time cure models (GPTCMs).
result Improves cancer prognosis by identifying cell-specific genes.

Study identifies and analyzes three types of errors in learning Fourier operators.

problem Statistical, discretization, and truncation errors in learning Fourier operators.
method Analysis of a Discrete Fourier Transform (DFT) based least squares estimator.
result Established upper and lower bounds on statistical, discretization, and truncation errors.

Symmetry helps VI recover certain statistics.

problem Understanding how symmetry in variational inference affects the recovery of statistics.
method Developed a general theory of symmetry-induced statistic recovery in variational inference.
result Symmetry can force the recovery of certain statistics in VI, even under model misspecification.

New formulae identify discrete probability laws without needing normalization constants.

problem Characterizing non-normalized discrete probability distributions.
method Derive explicit formulae for mass functions using Stein's method.
result Developed tools for solving statistical problems without normalization constants.

Fitting models for non-Poisson point processes is complicated by the lack of tractable models for much of the data. By using large samples of independent and identically distributed realizations and statistical learning, it is possible to identify absence of fit through finding a classification rule that can efficientl…

2007-12-02abs ↗pdf ↗

This paper tackles CRL for multi-node interventions, achieving identifiability guarantees.

problem CRL under unknown multi-node interventions, focusing on single-node assumptions.
method Establishes identifiability results for general latent causal models under stochastic interventions.
result Identifiability up to ancestors using soft interventions, perfect identifiability using hard interventions.

Paper identifies and estimates CAPCEs in continuous treatment settings.

problem Estimating heterogeneous causal effects of continuous treatments.
method Instrumental variable approach to identify CAPCEs under weaker conditions.
result Developed three families of CAPCE estimators with statistical properties analyzed.

Cookbook transforms constrained statistical inference into unconstrained problems.

problem Transforming constrained statistical inference into unconstrained problems.
method Bijective and diffeomorphisms parametrizations.
result Maintains statistical inference properties like identifiability.

Identifies learning rules from neural network observables.

problem Determine the underlying plasticity rules governing learning in biological systems.
method Simulated idealized neuroscience experiments with artificial neural networks to generate a dataset of learning trajectories. Used linear and non-linear classifiers to identify learning rules from aggregate statistics of weights, activations, and activity changes.
result Different classes of learning rules can be separated solely on the basis of aggregate statistics of the weights, activations, or instantaneous layer-wise activity changes.

High throughput screening of compounds (chemicals) is an essential part of drug discovery [7], involving thousands to millions of compounds, with the purpose of identifying candidate hits. Most statistical tools, including the industry standard B-score method, work on individual compound plates and do not exploit cross…

2017-09-28abs ↗pdf ↗

New method uses statistical physics to detect financial market manipulation.

problem Detecting financial market manipulation activities like spoofing and layering.
method Modeling order book dynamics as particle motion and using momentum measure.
result Method outperforms conventional Z-score-based anomaly detection.

How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…

2013-09-30abs ↗pdf ↗

The paper investigates topic models, ensuring their statistical identifiability and accuracy.

problem Lack of formal theoretical investigation of topic model identifiability and estimation accuracy.
method Proposes a maximum likelihood estimator (MLE) based on integrated likelihood, introducing new geometric identifiability conditions.
result Introduces weaker conditions for topic model identifiability, allowing a broader investigation.

Paper extends SI method for detecting CPs in complex systems' frequency domain.

problem Identifying change points in complex systems' frequency domain.
method Extends SI framework to frequency domain using DFT properties and develops valid p-values.
result Reliable detection of genuine CPs with strong statistical guarantees.

AI needs causal inference to avoid being just a correlation machine.

problem AI's inability to distinguish correlation from causation.
method Develops a unified framework connecting various causal statistical estimators and proves a Statistical Necessity Theorem for causal generalization.
result AI systems without causal grounding are brittle and biased, highlighting the need for causal statistics.

Establishes statistical and computational bounds for influence diagnostics.

problem Identifying influential datapoints or subsets in machine learning models.
method Finite-sample statistical bounds and computational complexity for influence functions and approximate maximum influence perturbations.
result Established statistical and computational guarantees for influence diagnostics.

For analysis of a high-dimensional dataset, a common approach is to test a null hypothesis of statistical independence on all variable pairs using a non-parametric measure of dependence. However, because this approach attempts to identify any non-trivial relationship no matter how weak, it often identifies too many rel…

2015-05-09abs ↗pdf ↗

This work closes the gap between theory and practice for nICA identifiability.

problem Identifying latent components in nonlinearly mixed data.
method Finite-sample analysis of GCL-based nICA, combining GCL properties, statistical generalization, and numerical differentiation.
result Establishes a trade-off between function learner complexity and expressiveness.

AI helps forecasters understand TC convective evolution before intensification.

problem Challenges in extracting scientific insights from complex TC data.
method Combining AI prediction algorithms and classical statistical inference.
result Identifies patterns in TC convective structure leading to intensification.

New method identifies causal relationships without strong assumptions.

problem Causal Representation Learning (CRL) is ill-posed due to representation and causal discovery issues.
method Identifiability based on grouping of observational variables, self-supervised estimation framework.
result Practical identifiability conditions without temporal structure, interventions, or weak supervision.

Method provides statistical guarantees for identifying subgroups in ML studies.

problem Bias and noise in estimating conditional average treatment effects (CATE).
method Develops uniform confidence bands (GATES) for estimating group average treatment effects (GATEs).
result Identifies subgroups with statistical guarantees, regardless of effect size.

Finding statistically significant high-order interaction features in predictive modeling is important but challenging task. The difficulty lies in the fact that, for a recent applications with high-dimensional covariates, the number of possible high-order interaction features would be extremely large. Identifying stati…

2015-06-26abs ↗pdf ↗

Deep neural networks identify robust arbitrage strategies in financial markets.

problem Identifying profitable trading strategies under model ambiguity.
method Data-driven deep neural networks considering high-dimensional financial markets.
result Empirical investigations show profitable trading performances in various market conditions.

The Rashomon effect shows many models can perform similarly, explored in this paper.

problem Why do many models perform similarly in machine learning?
method Categorized causes into statistical, structural, and procedural sources.
result Structural multiplicity persists and cannot be resolved without additional assumptions.