Fast nonparametric conditional independence testing via two-stage regression
problem Fast nonparametric conditional independence testing
method BLITZ (Broad-to-Local Independence Testing via residualiZation)
result Better null calibration than fast kernel, random-feature, and regression-based competitors
DIET tests conditional independence using marginal dependence measures of residual information.
problem Computational intractability of conditional randomization tests (CRTs).
method DIET avoids fitting large models by leveraging marginal independence statistics of information residuals.
result DIET achieves higher power than other tractable CRTs on synthetic and real benchmarks.
In this paper we consider portmanteau tests for testing the adequacy of multiplicative seasonal autoregressive moving-average (SARMA) models under the assumption that the errors are uncorrelated but not necessarily independent.We relax the standard independence assumption on the error term in order to extend the range …
Develops efficient inference for noise heterogeneity in machine learning models.
problem Downstream procedures based on residuals can be biased in additive noise models.
method Semiparametrically efficient inference using a novel Hilbert-valued one-step estimator.
result Constructs tests and confidence intervals for residual independence and goodness of fit.
Extends VAEs to handle complex Bayesian network structures.
problem Handling complex dependency structures in Bayesian networks.
method Extends VAEs with graphical residual flows to model arbitrary dependency structures.
result Demonstrates improved performance on synthetic datasets.
GaussDetect-LiNGAM eliminates Gaussianity tests for causal discovery.
problem Causal direction identification without Gaussianity assumptions.
method Leverages the equivalence between noise Gaussianity and residual independence in reverse regression.
result Gaussianity tests replaced with robust kernel-based independence tests.
The article consist of two main parts: an analog of the Leray Theory for Singular Varieties and its application to the Theory of Parshin's Residues. The first part is independent from the second. It uses the theory of Whitney stratifications. The second part is an application of the first. In particular, a geometric an…
This work improves fair tensor decomposition using a kernel criterion.
problem Learning fair low-rank tensor decompositions with statistical parity.
method Regularizes Canonical Polyadic Decomposition with KHSIC to ensure approximate statistical parity.
result The proposed algorithm achieves better fairness and fit than state-of-the-art FATR.
Alternative to likelihood-based LSNM model selection, residual independence testing is more robust to noise misspecification.
problem Cause-effect inference in location-scale noise models with misspecified noise distributions.
method Residual independence testing as an alternative to likelihood-based model selection.
result Residual independence testing is more robust to noise misspecification.
Two methods are proposed to filter correlations in DCC-GARCH residuals for foreign exchange rates.
problem Filtering correlations in DCC-GARCH residuals for accurate foreign exchange rate prediction.
method Two approaches: estimating correlation matrix as a parameter and using eigenvalue decomposition.
result The DCC-GARCH residual can be almost independent using these methods.
Graph neural networks often assume vertex labels are independent, but we show this is rarely true and propose a method to improve predictions.
problem Graph neural networks often assume vertex labels are conditionally independent given their neighborhood features, which is rarely true.
method We model the joint distribution of residuals on vertices with a parameterized multivariate Gaussian and estimate parameters by maximizing the marginal likelihood of the observed labels.
result Our method achieves substantially higher accuracy than competing baselines and can be interpreted as the strength of correlation among connected vertices.
We introduce a new deep convolutional neural network, CrescendoNet, by stacking simple building blocks without residual connections. Each Crescendo block contains independent convolution paths with increased depths. The numbers of convolution layers and parameters are only increased linearly in Crescendo blocks. In exp…
The covariance matrix is formulated in the framework of a linear multivariate ARCH process with long memory, where the natural cross product structure of the covariance is generalized by adding two linear terms with their respective parameter. The residuals of the linear ARCH process are computed using historical data …
Residually finite groups found in manifold automorphisms.
problem Residual finiteness of automorphism groups of high-dimensional manifolds.
method Embedding calculus, Weiss fibre sequence, convergence of embedding calculus tower, smoothing theory.
result Topological mapping class group of high-dimensional manifolds is residually finite.
Probabilistic principal component analysis (PPCA) seeks a low dimensional representation of a data set in the presence of independent spherical Gaussian noise, Sigma = (sigma^2)*I. The maximum likelihood solution for the model is an eigenvalue problem on the sample covariance matrix. In this paper we consider the situa…
The paper studies residues of manifolds and their applications in geometry.
problem Understanding the residues of manifolds and their geometric implications.
method Analytic continuation and Möbius invariance of residues, introduction of relative and weighted residues.
result Scalar curvature, mean curvature, and Euler characteristic can be expressed in terms of residues.
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
We have developed an automatic sleep stage classification algorithm based on deep residual neural networks and raw polysomnogram signals. Briefly, the raw data is passed through 50 convolutional layers before subsequent classification into one of five sleep stages. Three model configurations were trained on 1850 polyso…
This paper begins to explore the determinants of the topological properties of the international - trade network (ITN). We fit bilateral-trade flows using a standard gravity equation to build a "residual" ITN where trade-link weights are depurated from geographical distance, size, border effects, trade agreements, and …
Paper studies M-estimators with derivatives and residual distribution for robust adaptive tuning.
problem Tackles robustness and adaptive tuning of M-estimators with heavy-tailed noise.
method Provides formulae for derivatives, characterizes residual distribution, proposes adaptive criterion.
result Characterizes distribution of residuals and proposes adaptive criterion as out-of-sample error proxy.
Robust forecast framework reduces distribution error by 63%.
problem Accurate distribution forecast for planning decisions.
method Backtest-based bootstrap and adaptive residual selection.
result Reduces Absolute Coverage Error by more than 63%.
CIRCE measures conditional independence for learning invariant features.
problem Learning invariant features while being conditionally independent of a distractor.
method CIRCE is a measure of conditional independence applied as a regularizer in feature learning.
result CIRCE provides a zero value if and only if features are conditionally independent of the distractor given the target.
Constructs currents representing Baum-Bott residues for foliations.
problem Calculating Baum-Bott residues for complex foliations.
method Explicit construction of currents with support on singular components.
result Currents represent Baum-Bott residues and are independent under certain conditions.
We add size factor to CAPM and normalize residuals by Volatility Index.
problem Capturing the size effect in CAPM and making residuals Gaussian.
method Insert size effect, normalize residuals by Volatility Index, and fit model to real-world data.
result The new model shows long-term stability and connects to Stochastic Portfolio Theory.
TRA detects causal direction from bivariate data using geometric shapes.
problem Inferring causal direction from observational data is challenging and unreliable.
method TRA compares rank-based copula-standardized residual clouds to detect causal direction.
result TRA is robust and superior in detecting causal direction across various scenarios.
A new method, Residual-Permuted Sums, improves confidence region construction for linear regression models.
problem Constructing reliable confidence regions for linear regression models with non-symmetric noise.
method Residual-Permuted Sums (RPS) method, which permutes residuals instead of perturbing their signs.
result RPS provides exact finite sample coverage probabilities and is uniformly strongly consistent.
Framework calculates positional influence in causal residual Transformers.
problem Understanding positional influence in causal residual Transformers.
method Adjoint-sensitivity framework for positional influence in causal residual Transformers.
result Exact evolution of adjoint-energy influence density and decomposition into residual transmission, nonlocal Volterra, and local channels.
Probabilistic principal component analysis (PPCA) seeks a low dimensional representation of a data set in the presence of independent spherical Gaussian noise. The maximum likelihood solution for the model is an eigenvalue problem on the sample covariance matrix. In this paper we consider the situation where the data v…
Causal inference uses observations to infer the causal structure of the data generating system. We study a class of functional models that we call Time Series Models with Independent Noise (TiMINo). These models require independent residual time series, whereas traditional methods like Granger causality exploit the var…
Variational auto-encoders (VAEs) are a popular and powerful deep generative model. Previous works on VAEs have assumed a factorized likelihood model, whereby the output uncertainty of each pixel is assumed to be independent. This approximation is clearly limited as demonstrated by observing a residual image from a VAE …
Develops a new multivariate regression model for complex outcomes.
problem Flexible, heterogeneous, and residual-dependent multivariate regression problems.
method MultiVCBART framework with Graphical Horseshoe priors.
result Empirically outperforms existing models on sparse, high-dimensional datasets.
We construct minimal laminations by hyperbolic surfaces whose generic leaf is a disk and contain any prescribed family of surfaces and with a precise control of the topologies of the surfaces that appear. The laminations are constructed via towers of finite coverings of surfaces for which we need to develop a relative …
Inverted file and asymmetric distance computation (IVFADC) have been successfully applied to approximate nearest neighbor search and subsequently maximum inner product search. In such a framework, vector quantization is used for coarse partitioning while product quantization is used for quantizing residuals. In the ori…
New method for learning causal relationships in PNL models.
problem Learning causal relationships from empirical observations in PNL models.
method Rank-based methods to estimate non-linear functions, disentangling from independence tests.
result Consistent method for PNL causal discovery, validated in experiments.
In recent years, the RFRS condition has been used to analyze virtual fibering in 3-manifold topology. Agol's work shows that any 3-manifold with zero Euler characteristic satisfying the RFRS condition on its fundamental group virtually fibers over the circle. In this note we will show that a finitely generated nilpoten…
New method recovers causal order from dependent data.
problem Causal discovery methods fail with shared volatility or common scale effects.
method Linear Mean-Independent Acyclic Model (LiMIAM) with mean-independence restrictions.
result Compatible causal order can be recovered from dependent disturbances.
REMAL: Residual Equilibrium Manifold Active Learning for Surrogate-Based Multidisciplinary Design Analysis
problem Multidisciplinary design analysis of coupled engineering systems requires solving equilibrium states where all disciplinary coupling variables are consistent.
method Residual manifold surrogate modeling framework for coupled systems.
result REMAL learns a surrogate model of the joint residual manifold via multitask Gaussian process models.
Polynomial-time algorithm learns causal graphs without parametric assumptions.
problem Learning causal graphs from data without assuming linearity or parametric forms.
method Model-free polynomial-time algorithm with finite-sample guarantees.
result Algorithm achieves linear cost in dimension and samples compared to optimal.
Gradient descent converges to global minima for ResNets with linearly scaled width.
problem Understanding the convergence of deep residual networks with varying network width and dataset size.
method Analyzing the Jacobian of ResNets and applying gradient descent for quadratic loss.
result Gradient descent converges to global minima for ResNets with linearly scaled width and independent of depth.
Sharp stability threshold found for deep residual architectures.
problem Ensuring stable training and inference in deep residual networks.
method Sublinear-growth principle and optimal-control analysis.
result Stable training condition: input-magnitude exponent q ≤ 1.
Recent results in the literature indicate that a residual network (ResNet) composed of a single residual block outperforms linear predictors, in the sense that all local minima in its optimization landscape are at least as good as the best linear predictor. However, these results are limited to a single residual block …
Constraint-based structure learning algorithms infer the causal structure of multivariate systems from observational data by determining an equivalent class of causal structures compatible with the conditional independencies in the data. Methods based on additive-noise (AN) models have been proposed to further discrimi…
Causal inference from observational data requires assumptions. These assumptions range from measuring confounders to identifying instruments. Traditionally, causal inference assumptions have focused on estimation of effects for a single treatment. In this work, we construct techniques for estimation with multiple treat…
The study analyzes numerical stability in large language models using mixed-precision arithmetic.
problem Numerical stability of large language models using low-precision arithmetic.
method Developed a mixed-precision analysis of transformer inference, deriving bounds for condition numbers and forward error.
result Established that numerical stability is determined by the interplay between weight magnitude and the growth of the residual stream.
Temporal aggregation reveals latent default correlation from monthly data.
problem Understanding effective default correlation from monthly default data.
method Temporal coarse-graining of latent default-probability paths.
result Temporal coarse-graining improves identifiability and reduces over-allocation of long-horizon fluctuations.
Study tests adequacy of FARIMA models with uncorrelated but non-independent errors.
problem Testing adequacy of FARIMA models with specific error characteristics.
method Derive asymptotic distributions of residual autocovariances and autocorrelations, propose self-normalization approach.
result Asymptotic distributions of modified portmanteau statistics for weak FARIMA models.
Temporal coarse-graining of latent default paths explains effective correlation in corporate defaults.
problem Understanding effective default correlation in corporate defaults.
method Temporal coarse-graining of latent default-probability paths, applied to corporate default-count data.
result Temporal coarse-graining provides a scale-consistent baseline that improves identifiability and reduces over-allocation of long-horizon fluctuations.
LANCA uses ANM to learn latent causal factors without supervision.
problem Learning latent causal factors without supervision.
method LANCA employs a deterministic Wasserstein Auto-Encoder coupled with a differentiable ANM Layer.
result LANCA outperforms baselines on physics and photorealistic environments.