High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
Novel approach ensures stability of compact schemes for variable PDEs.
problem Ensuring stability of compact schemes for variable coefficient PDEs.
method Difference equation approach to derive stability conditions.
result Derives sufficient condition for unconditional stability.
Improving the detection of relevant variables using a new bivariate measure could importantly impact variable selection and large network inference methods. In this paper, we propose a new statistical coefficient that we call the rank minrelation coefficient. We define a minrelation of X to Y (or equivalently a majrela…
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
problem The standard Pearson correlation coefficient is limited to two variables and doesn't meet the needs for multi-variable analysis.
method The authors use random matrix theory to extend Pearson's correlation coefficient to an arbitrary number of variables.
result The extended correlation coefficient is useful for gauging noise and selecting features, particularly in classification.
We consider the problem of predicting several response variables using the same set of explanatory variables. This setting naturally induces a group structure over the coefficient matrix, in which every explanatory variable corresponds to a set of related coefficients. Most of the existing methods that utilize this gro…
The study uses DCC for financial market analysis, revealing hidden correlations.
problem Identifying hidden nonlinear correlations in financial markets.
method Agglomerative hierarchical clustering with distance correlation coefficient.
result DCC reveals more information than Pearson correlation for financial data.
Abstract: Determines thermoelastic coefficients from boundary data.
problem Determining coefficients of thermoelastic system from boundary information.
method Explicit expression for thermoelastic Dirichlet-to-Neumann map with variable coefficients.
result Thermoelastic Dirichlet-to-Neumann map uniquely determines coefficients on the manifold.
Bayesian method discovers PDEs with variable coefficients robustly.
problem Discovering PDEs from noisy data is challenging.
method Bayesian sparse learning with tBGL-SS and Gibbs sampler.
result Method enhances robustness and model selection criteria.
The performance of Orthogonal Matching Pursuit (OMP) for variable selection is analyzed for random designs. When contrasted with the deterministic case, since the performance is here measured after averaging over the distribution of the design matrix, one can have far less stringent sparsity constraints on the coeffici…
The article proposes modified Gower's coefficients for handling mixed type variables in nearest neighbor methods.
problem Handling mixed type variables in nearest neighbor methods, especially imputation and statistical matching.
method Suggests modifications to the Gower's distance for interval and ratio scaled variables to address unbalanced contributions and outlier sensitivity.
result Improved distance calculations reduce the unbalanced contribution of different variable types and attenuate outlier effects.
This paper addresses parameter estimation for wave equations with Markovian switching.
problem Parameter estimation for wave equations with abrupt changes.
method Bayesian statistical framework using discrete sparse Bayesian learning.
result Strong performance in parameter estimation for variable coefficient PDEs.
We consider the problem of constructing a reduced-rank regression model whose coefficient parameter is represented as a singular value decomposition with sparse singular vectors. The traditional estimation procedure for the coefficient parameter often fails when the true rank of the parameter is high. To overcome this …
We study learning problems in which the conditional distribution of the output given the input varies as a function of additional task variables. In varying-coefficient models with Gaussian process priors, a Gaussian process generates the functional relationship between the task variables and the parameters of this con…
We introduce a general setting for multidimensional dispersionless integrable hierarchy in terms of differential m-form Ωm with the coefficients satisfying the Plücker relations, which is gauge-invariantly closed and its gauge-invariant coordinates (ratios of coefficients) are (locally) holomorphic with respect to…
New method for MTL with varying sparsity patterns across tasks.
problem Jointly training multiple linear models with differing sparsity patterns.
method Mixed-integer programming formulation and scalable algorithms.
result Our methods leverage shared support information to improve variable selection.
In this paper we use wavelet concepts to show that correlation coefficient between two financial data's is not constant but varies with scale from high correlation value to strongly anti-correlation value This studies is important because correlation coefficient is used to quantify degree of independence between two va…
The paper derives theoretical foundations for two common machine learning variable importance measures.
problem Understanding variable importance in machine learning problems.
method The paper derives closed-form expressions for Permute-and-Predict (PaP) and Leave-One-Covariate-Out (LOCO) methods.
result Theoretical derivations explain the behavior of PaP and LOCO under collinearity, linking them to coefficients and predictor variability.
Knoop enhances variable selection with over-parameterization and knockoffs.
problem Challenges of variable selection in high-dimensional datasets.
method Generates knockoff variables, integrates them into an over-parameterized model, and uses anomaly-based significance tests.
result Superior performance in variable selection compared to existing methods.
A novel non-supervised method detects anomalies in multivariate time series.
problem Detecting anomalies in multivariate time series data.
method Partitioning based on clustering of correlation coefficients.
result Significant improvement in anomaly detection performance.
We introduce the Randomized Dependence Coefficient (RDC), a measure of non-linear dependence between random variables of arbitrary dimension based on the Hirschfeld-Gebelein-Rényi Maximum Correlation Coefficient. RDC is defined in terms of correlation of random non-linear copula projections; it is invariant with respec…
MIC consistently estimates dependence in large datasets.
problem Estimating dependence between variable pairs in large datasets.
method Proving consistency of MIC as an estimator.
result MIC is a consistent estimator of population statistic MIC*.
Method identifies causal drivers from background features.
problem Distinguishing causal influence from hidden confounding.
method Stability of regression coefficients measured by statistic V.
result V converges to zero if and only if no causal drivers exist.
Abstract: Determines Lamé coefficients from boundary measurements.
problem Determining Lamé coefficients from elastic boundary measurements.
method Explicit symbol of elastic Dirichlet-to-Neumann map, partial derivatives determination.
result Elastic Dirichlet-to-Neumann map uniquely determines Lamé coefficients.
Generalizes underlap coefficient for multivariate group separation.
problem Quantifying distributional separation across groups in statistical learning.
method Generalizes underlap coefficient (UNL) to multivariate variables, establishes key properties, interprets as dependence measure, proposes efficient estimator.
result Highlights the UNL's utility in clustering for evaluating group structure dependence on covariates.
The paper discusses methods for interval estimation of coefficients in penalized regression models for insurance data.
problem Valid inference on coefficients after feature selection in GLM family for insurance data.
method Proposes methodologies for constructing confidence intervals of coefficients after feature selection in GLM family.
result Valid inference on coefficients after feature selection in GLM family for insurance data.
We reproduced the results of CheXNet with fixed hyperparameters and 50 different random seeds to identify 14 finding in chest radiographs (x-rays). Because CheXNet fine-tunes a pre-trained DenseNet, the random seed affects the ordering of the batches of training data but not the initialized model weights. We found subs…
SCOPE fuses categorical variable levels to estimate high-dimensional linear models.
problem Estimating high-dimensional linear models with nominal categorical data.
method SCOPE uses nonconvex concave penalties to fuse levels and achieve efficient computation.
result SCOPE achieves oracle least squares solution under certain conditions.
Extends compactness theory to variable-coefficient pseudo-differential operators on manifolds.
problem Compensated compactness for pseudodifferential operators on vector bundles.
method Establishes a theorem for weakly convergent sequences of sections under a pseudo-differential operator.
result Quadratic form converges in distributional sense under certain conditions.
An AI approach selects variables in linear models.
problem Selecting significant variables in linear regression models.
method Artificial Neural Network trained to determine variable significance based on OLS estimates.
result The AI approach outperforms traditional methods in accuracy and variable selection.
We proposed a new statistical dependency measure called Copula Dependency Coefficient(CDC) for two sets of variables based on copula. It is robust to outliers, easy to implement, powerful and appropriate to high-dimensional variables. These properties are important in many applications. Experimental results show that C…
New concentration inequalities for tensors with heavy-tailed coefficients.
problem Developing bounds for Euclidean functions of tensors with sub-Weibull distributions.
method Extending concentration inequalities to sub-Weibull random tensors, using new inequalities for heavy-tailed random variables and martingale analysis.
result Established a phase transition between sub-gaussian and heavy-tailed regimes for Euclidean functions of tensors.
We derive an explicit formula for likelihood function for Gaussian VARMA model conditioned on initial observables where the moving-average (MA) coefficients are scalar. For fixed MA coefficients the likelihood function is optimized in the autoregressive variables Φ's by a closed form formula generalizing regression c…
In this study, we have investigated factors of determination which can affect the connected structure of a stock network. The representative index for topological properties of a stock network is the number of links with other stocks. We used the multi-factor model, extensively acknowledged in financial literature. In …
A new method treats all variables equally in fitting data.
problem Fitting relationships to data with multiple variables, especially when dependent and independent variables are not clearly defined.
method A general method treating all variables impartially, using geometric mean functional relationships and correlation.
result The method provides coefficients that are easily calculated from covariances or correlations, making it scale-invariant and applicable to various units.
Proposes a Varying-Coefficient MoE model for analyzing dynamic data.
problem Inadequate constant coefficients in MoE models for dynamic settings.
method Varying-Coefficient Mixture of Experts (VCMoE) model with varying coefficients in gating and expert models.
result Established identifiability and consistency of the VCMoE model.
The SLOPE estimates regression coefficients by minimizing a regularized residual sum of squares using a sorted-ℓ1-norm penalty. The SLOPE combines testing and estimation in regression problems. It exhibits suitable variable selection and prediction properties, as well as minimax optimality. This paper introduces …
A new method combines machine learning with mixed-effects models for better repeated measurement analysis.
problem Inference of linear coefficients in partially linear mixed-effects models with complex interactions and high-dimensional variables.
method Double machine learning approach to estimate nonparametrically nonlinear variables, then use standard linear mixed-effects techniques to estimate the linear coefficient.
result The estimated fixed effects coefficient converges at the parametric rate and is semiparametrically efficient.
New theorem bounds link volume using surface coefficients.
problem Bounding hyperbolic volume of links on surfaces.
method Analogue of Dasbach-Lin theorem for surface links.
result Bounds on link volume from surface polynomial coefficients.
We show a connection between the Fourier spectrum of Boolean functions and the REINFORCE gradient estimator for binary latent variable models. We show that REINFORCE estimates (up to a factor) the degree-1 Fourier coefficients of a Boolean function. Using this connection we offer a new perspective on variance reduction…
In this article, we consider a 2 factors-model for pricing defaultable bond with discrete default intensity and barrier where the 2 factors are stochastic risk free short rate process and firm value process. We assume that the default event occurs in an expected manner when the firm value reaches a given default barrie…
New algorithm selects relevant variables in high-dimensional graphical models.
problem Automatic selection of relevant variables in high-dimensional graphical models.
method Extends Chow and Liu's algorithm using mutual information and entropy coefficient of determination.
result Outperforms existing methods in selecting variables with explanatory power.
A novel Bayesian method for dynamic sparsity in Gaussian dynamic linear regression.
problem Variable selection and shrinkage in time-varying regression models.
method Time-varying sparsity via Markov switching priors for coefficients' variances, extending spike-and-slab priors.
result Induces smoothness or shrinkage towards zero at each time point, leading to improved model performance.
We propose a method for estimating coefficients in multivariate regression when there is a clustering structure to the response variables. The proposed method includes a fusion penalty, to shrink the difference in fitted values from responses in the same cluster, and an L1 penalty for simultaneous variable selection an…
SIP framework discovers governing equations in uncertain systems.
problem Discovering governing equations in systems with input variability and noisy data.
method SIP framework treats unknown coefficients as random variables and infers their posterior distribution by minimizing Kullback-Leibler divergence.
result SIP consistently identifies correct equations and lowers coefficient error by 82% relative to SINDy.
Characterizes differential forms and vector fields with constant coefficients on manifolds.
problem Understanding constant coefficient differential forms and vector fields on manifolds.
method Analyzes differential forms and vector fields of specific degrees, proving obstructions and characterizing solutions to partial differential systems.
result Characterizes differential forms and vector fields with constant coefficients of various degrees on smooth manifolds.
New robust estimator improves variable selection and coefficient estimation in linear regression with heavy-tailed errors and outliers.
problem Heavy-tailed errors and anomalous predictors in high-dimensional regression.
method Adaptive PENSE estimator for robust variable selection and estimation.
result Adaptive PENSE estimator provides reliable results even under very heavy-tailed errors and aberrant predictors.
Reduces selection bias in estimating individual treatment effects.
problem Selection bias in counterfactual reasoning.
method Auto-encoder with regularized loss based on Pearson Correlation Coefficient.
result Improves performance in estimating individual treatment effects.
Improves Gower's similarity for mixed-type variables with automatic weighting.
problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.