Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

151301452602 · Jun 202019922001200920182026
48 results for Binary distributions

Paper constructs unfaithful probability distributions in binary causal graphs.

problem Unfaithful probability distributions in binary causal graphs.
method Constructs unfaithful probability distributions in binary causal graphs.
result Examples of unfaithful probability distributions in binary causal graphs.

Derives pricing formulae for power binary and normal distribution standard options.

problem Developing pricing models for binary and standard options.
method Incorporates Buchen's formulae into power binary options and derives a formula for normal distribution standard options.
result Derives pricing formulae for power binary and normal distribution standard options.

Paper resolves open problems on sample complexity in binary hypothesis testing.

problem Open problems in distributed simple binary hypothesis testing under information constraints.
method One-shot lower bound on Bayes error, streamlined sample complexity formula, reverse data-processing inequality.
result Optimally tight sample complexity bounds for communication-constrained simple binary hypothesis testing.

BIND removes background noise from binary matrices, improving detection accuracy and fairness.

problem Real data often violates the i.i.d assumption for binary matrix entries, leading to inaccurate detection.
method BIND optimizes detection by estimating row- and column-wise mixture distributions and eliminating background noise.
result BIND effectively removes background noise and increases detection accuracy and fairness.

A new method combines simple binary classifiers to build complex multiclass classifiers, achieving performance limits in a Gaussian setting.

problem Building a sophisticated multiclass classifier from simple binary decisions.
method Combining O(logK)O(\log K) simple binary classifiers to form a KK-class classifier.
result Explicit performance bounds across various decoding and dimensional regimes for a stylized Gaussian setting.

Study binary hypothesis testing with privacy and communication constraints.

problem Binary hypothesis testing under local differential privacy and communication constraints.
method Qualifies results as minimax or instance optimal, develops instance-optimal algorithms.
result Achieves minimum possible sample complexity under both privacy and communication constraints.

Study three types of uncertainty quantification for binary classification without distributional assumptions.

problem Uncertainty quantification for binary classification in a distribution-free setting.
method Established theorems connecting calibration, confidence intervals, and prediction sets for score-based classifiers.
result Distribution-free calibration is only possible using scoring functions that partition feature space into countably many sets.

Modified Metropolis algorithm ensures convergence for multivariate binary distributions with fixed-order updates.

problem Infeasibility of standard Metropolis algorithm for multivariate binary distributions with fixed-order updates.
method Proposed a modified Metropolis transition operator ensuring irreducibility and convergence.
result Ensures convergence to the limiting distribution in multivariate binary case with fixed-order updates.

New tests for binary classification regression functions without distribution assumptions.

problem Testing regression functions in binary classification without distributional assumptions.
method Conditional kernel mean embeddings and resampling-based framework.
result Distribution-free hypothesis tests with exact type I error control.

Discovering causal relations among observed variables in a given data set is a major objective in studies of statistics and artificial intelligence. Recently, some techniques to discover a unique causal model have been explored based on non-Gaussianity of the observed data distribution. However, most of these are limit…

2014-01-22abs ↗pdf ↗

RBMs model binary interactions with hidden node activation effects.

problem Understanding how RBM hidden node activation affects binary variable distributions.
method Investigated RBM marginal distributions with different hidden node activation functions.
result Found exact expressions for RBM marginals as interacting binary variables.

The study examines the discrepancies between binary forecasts and real-world outcomes, revealing their often misleading nature.

problem The confusion between binary forecasts and real-world payoffs in decision-making and prediction.
method Comparative analysis of binary forecasts, bets, and real-world continuous payoffs under different tail conditions.
result Binary forecasting abilities do not translate to better real-world performance, and vice versa, especially under nonlinearities.

New algorithm separates audio sources better using alpha-stable distributions.

problem Improving audio source separation using complex distributions.
method Estimating mixtures of alpha-stable distributions using characteristic function matching.
result Better separation performance than Gaussian-based methods.

Alternative to convolutions using decision trees for neural networks.

problem Replacing complex convolutions with simpler decision-based layers.
method Binary decisions as indices to conditional distributions, trained using backpropagation.
result Performance similar to conventional neural networks, with runtime improvements.

A novel method for feature selection using a reparameterized logitNormal distribution.

problem Feature selection for reconstruction in high-dimensional data.
method Introducing a reparameterization of the logitNormal distribution to address differentiability and covariance issues.
result The method provides an effective exploration scheme and efficient feature selection for reconstruction.

Binary testing for softmax models requires many samples, similar to leverage score models.

problem Binary hypothesis testing for softmax models and leverage score models.
method Analyzing sample complexity and drawing analogies between models.
result Sample complexity is asymptotically \(O(ε^{-2})\), where \(ε\) is the distance between model parameters.

Quantum circuits represent binary classification trees with binary features.

problem Classifying data using binary classification trees with binary features.
method Quantum circuits and probabilistic approach for traversing decision trees.
result First realization of a decision tree classifier on a quantum device.

Paper analyzes impact of PRM on binary random variables and distribution shifts.

problem Impact of performative risk minimization on binary random variables and distribution shifts.
method Formulated two measures of impact, derived explicit formulas for full information, and provided estimators for partial information.
result PRM can have amplified side effects compared to methods that do not model data shift.

A new method for binary ICA using non-stationary sources.

problem Independent component analysis of binary data.
method Linear mixing model in latent space, followed by binary observation model with non-stationary sources.
result Proves non-identifiability with few observed variables but identifies with more variables.

The paper develops efficient estimators for semi-parametric binary models in distributed computing.

problem Estimation and inference challenges in large-scale data under non-smooth objective functions.
method Proposes one-shot and multi-round divide-and-conquer estimators with adaptive kernel smoothing to relax constraints and achieve superlinear optimization error.
result Establishes quadratic convergence up to optimal statistical error rate and handles dataset heterogeneity and high-dimensional sparse parameters.

Continuous Sweep improves binary quantifier performance.

problem Estimating class prevalence in datasets.
method Parametric binary quantifier inspired by Median Sweep, using parametric class distributions and mean of Adjusted Count estimates.
result Continuous Sweep outperforms other quantifiers in simulations and empirical data analysis.

New method recovers predictions from unobservable source subpopulation in binary classification.

problem Challenging binary classification with unobservable subpopulation in source domain.
method Distribution matching method to estimate subpopulation proportions, rigorous derivation of prediction models.
result Our method outperforms naive benchmarks in synthetic and real-world datasets.

ENTED efficiently decomposes binary and count tensors using nonparametric Gaussian processes.

problem Handling high-dimensional and sparse binary and count data with traditional tensor decompositions.
method ENTED uses nonparametric Gaussian processes and sparse orthogonal variational inference to handle binary and count tensors.
result ENTED outperforms traditional methods in binary and count tensor completion tasks.

This study suggests replacing Ising distribution with Cox distribution.

problem Handling correlated binary data efficiently.
method Exploring conditions for replacing Ising distribution with Cox distribution as a latent variable model.
result The Ising distribution can be treated as a latent variable model with a quasi-normal distribution.

Binary classification models get more efficient predictive probabilities.

problem Computing predictive probabilities in Bayesian probit models is computationally challenging.
method Use of expectation propagation (EP) to find a closed-form expression for predictive probabilities.
result Closed-form predictive probabilities improve over existing methods.

Sharp analysis of isotonic regression for binary data, improving calibration bounds.

problem Improving the calibration of probabilistic predictors using isotonic regression.
method Sharp finite-sample characterization of isotonic regression's degrees of freedom using analytic number theory.
result First nontrivial distribution-free guarantee on Expected Calibration Error (ECE) of isotonic regression.

The paper proposes using theoretical ROC curves to categorize classifier responses.

problem The lack of explicit probability distributions for classifier responses in machine learning.
method Fit beta distributions to classifier responses and use them to categorize responses into different classes.
result Established a categorization of classifier responses into classes with different ROC curve extremal behaviors.

New algorithm optimizes margin distribution in binary classifiers.

problem Optimizing margin distribution in binary classifiers.
method Proposes an algorithm that searches the hypothesis space to ensure a pre-set margin level is a robust estimator of the margin location.
result Empirical tests show the method is effective and promising for classification.

BEGIN network models binary data without parametric assumptions.

problem Conditional independence in non-parametric families of binary data.
method BEGIN network models binary data using sparse linear representations and block factorizations.
result BEGIN network captures conditional independence for arbitrary binary and multinomial variables.

Optimal gradient quantization reduces communication costs in distributed deep learning.

problem High communication costs in distributed training of deep neural networks.
method Deduced optimal gradient quantization conditions for binary and multi-level quantization, developed novel schemes for dynamic quantization levels.
result Demonstrated superior performance of proposed quantization schemes on CIFAR and ImageNet datasets.

A new neural network using chi-square test for binary classification.

problem Improving binary classification accuracy.
method Backpropagation neural network with chi-square test redefined cost and error functions.
result Significantly improved classification accuracy compared to related approaches.

Sparse binary compression reduces communication costs in distributed deep learning.

problem Limited communication bandwidth in distributed deep learning.
method Combines gradient sparsification, binarization, and optimal weight update encoding.
result Reduces upstream communication by more than four orders of magnitude.

Improved optimization methods for discrete distributions reduce bias in gradient estimation.

problem Estimating gradients for discrete distribution parameters is challenging.
method Analyzed and proposed methods to reduce bias in gradient estimation, including Gumbel-Softmax and piece-wise linear continuous relaxation.
result Reduced bias leads to better performance in variational inference and binary optimization tasks.