Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

189379568757 · Jun 202019922001200920182026
48 results for two variables

The paper introduces methods to identify key variables discriminating between two datasets.

problem Identifying variables that distinguish between two datasets.
method Introduces a mathematical notion of discriminating variables and proposes two methods for their selection.
result Proposed methods improve upon existing techniques in two-sample variable selection.

Proposes a method to select variables for kernel two-sample tests.

problem Determining whether two samples have the same distribution using informative variables.
method A framework based on kernel maximum mean discrepancy (MMD) for selecting a subset of variables.
result The sample size requirements for the three kernels depend on the number of selected variables, not the data dimension.

Many inference problems involving questions of optimality ask for the maximum or the minimum of a finite set of unknown quantities. This technical report derives the first two posterior moments of the maximum of two correlated Gaussian variables and the first two posterior moments of the two generating variables (corre…

2009-10-01abs ↗pdf ↗

Proposes a two-stage method for selecting correlated predictors in high-dimensional data.

problem Selecting correlated predictors in high-dimensional data with unknown group structures.
method Two-stage approach: variable clustering followed by group selection.
result The two-stage method improves prediction accuracy and active predictor selection.

This paper categorifies Chebyshev polynomials using diagrammatic algebra.

problem Categorifying two-variable Chebyshev polynomials of the second kind.
method Using A2A_2 spider and Karoubi envelope of A2A_2 spider, the recursive formula is shown.
result A qq-deformation of the two-variable Chebyshev polynomials is defined.

The paper shows how neural networks with less decision boundary variability generalize better.

problem Improving neural network generalizability by reducing decision boundary variability.
method Introduces new measures (algorithm DB variability and (ε,η)(ε, η)-data DB variability) to quantify decision boundary variability and proves theoretical bounds on generalizability.
result Neural networks with lower decision boundary variability have better generalizability, as shown by extensive experiments and theoretical bounds.

A Gaussian restricted Boltzmann machine (GRBM) is a Boltzmann machine defined on a bipartite graph and is an extension of usual restricted Boltzmann machines. A GRBM consists of two different layers: a visible layer composed of continuous visible variables and a hidden layer composed of discrete hidden variables. In th…

2015-12-03abs ↗pdf ↗

New polynomials from virtual knot invariants detect cosmetic changes.

problem Detecting cosmetic crossing changes in virtual knots.
method Introducing two-variable polynomials based on crossing and dwrithe values, proving they are invariants.
result Polynomials LKnL^n_K and FKnF^n_K detect when virtual knots do not admit cosmetic crossing changes.

Symmetric game analysis shows Nash equilibria in three strategic states.

problem Analyzing Nash equilibria in a symmetric multi-player zero-sum game with two strategic variables.
method Using the minimax theorem by Sion to show equivalence of Nash equilibria.
result Nash equilibria are equivalent in three strategic states.

This paper explores what causal structures can be distinguished by observational and interventional probing schemes.

problem Identifying causal structures with latent variables using observational and interventional data.
method Investigates the power of different probing schemes (observation vs. intervention) to distinguish causal structures.
result Two causal structures are indistinguishable if they share the same mDAG structure.

Bayesian method models binary response and covariates for two groups, estimating causal relationships.

problem Estimating causal relationships between binary response and covariates in observational data.
method Gaussian DAG-probit model with MCMC sampling for posterior distribution estimation.
result Validated method on simulated and real datasets, showing value of grouping variable in causality.

Paper compares largest claim amounts from two interdependent portfolios.

problem Comparing claim amounts from two sets of interdependent portfolios.
method Stochastic comparisons using dependent non-negative random variables and Bernoulli variables.
result Stochastic order results for largest claim amounts.

The paper addresses statistical estimation in MDPs with confounders using instrumental variables.

problem Statistical estimation of value functions in MDPs with unobservable confounders.
method Two-stage estimator based on instrumental variables for confounded linear MDPs.
result Established statistical properties of the two-stage estimator, including error bounds and asymptotic normality.

Approach selects variables and time intervals for comparing high-dimensional time-series data.

problem Comparing high-dimensional time-series data for significant differences.
method Data is split into subintervals, and two-sample tests are performed on each to identify distinguishing variables.
result The approach effectively identifies variables and time intervals where data significantly differs.

In data science and machine learning, hierarchical parametric models, such as mixture models, are often used. They contain two kinds of variables: observable variables, which represent the parts of the data that can be directly measured, and latent variables, which represent the underlying processes that generate the d…

2014-08-25abs ↗pdf ↗

Two Bayesian optimization methods tackle dynamic design spaces with mixed variables.

problem Optimizing complex systems with varying numbers and types of variables and constraints.
method Two Bayesian optimization approaches: budget allocation and kernel function.
result Both methods converge faster and more consistently than standard approaches.

Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…

2014-11-21abs ↗pdf ↗

In this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate Gaussian d…

2017-05-30abs ↗pdf ↗

Standard probabilistic linear discriminant analysis (PLDA) for speaker recognition assumes that the sample's features (usually, i-vectors) are given by a sum of three terms: a term that depends on the speaker identity, a term that models the within-speaker variability and is assumed independent across samples, and a fi…

2017-04-07abs ↗pdf ↗

A mathematical paradox shows secant planes don't always form a tangent plane, but some analogies hold with a specific vector product.

problem Secant planes of a two-variable smooth function do not always form a tangent plane, even for simple polynomials.
method Analogies with the one-variable case are explored, using Clifford's geometric vector product.
result Some analogies with the one-variable case still hold in the multi-variable context with a specific vector product.

We extend common entropy concept and propose algorithms to distinguish causation from correlation.

problem Discovering the simplest latent variable for conditional independence of observed variables.
method Renyi common entropy, iterative algorithm, constraint-based methods modification.
result Improved constraint-based methods for causal inference in small samples.

New invariant CWRCWR for alternating links is stronger than existing invariants.

problem Developing a stronger invariant for alternating links.
method Introducing CWRCWR invariant as an array of two-variable polynomials.
result The CWRCWR invariant is stronger than classical invariants like HOMFLYPT and Kauffman polynomials.

New bounds on continuous random variables' right-tail probabilities.

problem Finding precise upper and lower limits for right-tail probabilities of continuous random variables.
method Developed new bounds based on PDF, first derivative, and two parameters.
result The new bounds are tight for various continuous random variables.

New methods rank variables for Gaussian processes better than automatic relevance determination.

problem Variable selection for Gaussian process models using inverse length-scale parameters has limitations.
method Two novel methods rank variables based on their predictive relevance using posterior predictive distribution predictions.
result Improved variable selection compared to automatic relevance determination in terms of variability and predictive performance.

By taking into account the nonlinear effect of the cause, the inner noise effect, and the measurement distortion effect in the observed variables, the post-nonlinear (PNL) causal model has demonstrated its excellent performance in distinguishing the cause from effect. However, its identifiability has not been properly …

2012-05-09abs ↗pdf ↗

Derives derivatives and geometric framework for functions with non-independent variables.

problem Characterizing functions with non-independent variables in probabilistic models.
method Derives actual and dependent partial derivatives, dependent Jacobian matrix, and tensor metric.
result Derives gradient, Hessian, and Taylor expansion for functions with non-independent variables.

The paper compares one-hot encoding to Naïve Bayes for categorical variables.

problem Incorrect one-hot encoding affects Naïve Bayes performance.
method Mathematical and experimental analysis of PoB vs. categorical Naïve Bayes.
result Posterior probabilities are usually greater in the PoB case, but agree on the maximum a posteriori class label.

Two statistical tasks are shown to have equivalent sample complexity.

problem Determining if a function depends on only a few variables and identifying those variables.
method Proved statistical equivalence of feature selection and junta testing through sample complexity analysis.
result Brute-force algorithm is sample-optimal for both tasks with optimal sample size.

Proposes a two-stage method for testing variable interactions with FDR control.

problem Testing pairwise interactions in high-dimensional data with dependence.
method Two-stage testing procedure with FDR control using Cramér type moderate deviation technique.
result The proposed method controls FDR and has comparable or improved statistical power.

In some speaker recognition scenarios we find conversations recorded simultaneously over multiple channels. That is the case of the interviews in the NIST SRE dataset. To take advantage of that, we propose a modification of the PLDA model that considers two different inter-session variability terms. The first term is t…

2015-11-20abs ↗pdf ↗

A serious problem in learning probabilistic models is the presence of hidden variables. These variables are not observed, yet interact with several of the observed variables. Detecting hidden variables poses two problems: determining the relations to other variables in the model and determining the number of states of …

2013-01-10abs ↗pdf ↗

DualIV simplifies non-linear IV regression via dual formulation.

problem Non-linear instrumental variable regression with potential first-stage regression bottleneck.
method Dual formulation of non-linear IV regression as a convex-concave saddle-point problem, leading to a kernel-based algorithm with analytic solution.
result Empirical results show competitive performance compared to existing algorithms.

This article considers the problem of multi-group classification in the setting where the number of variables pp is larger than the number of observations nn. Several methods have been proposed in the literature that address this problem, however their variable selection performance is either unknown or suboptimal to…

2014-11-23abs ↗pdf ↗