New method learns signals from binary measurements, surpassing existing techniques.
problem Learning signals from noisy, incomplete, and quantized binary measurements.
method Self-supervised learning approach (SSBM) for binary data.
result SSBM outperforms supervised learning and sparse reconstruction methods.
Study shows how adjusting for a binary proxy can bound causal effects.
problem Bounding causal effects with a binary confounder and proxy.
method Monotonicity assumption applied to a binary confounder and observed proxy.
result Adjusting for a proxy produces a measure of the effect between unadjusted and true measures.
Paper extends FOFC algorithm to work with mixed data types.
problem Designing causal discovery algorithms for mixed data types.
method Proves tetrad constraint can be entailed for mixed data types and applies FOFC algorithm.
result FOFC algorithm can work on mixed data types.
In this paper, we study the problem of compressed sensing using binary measurement matrices and ℓ 1 \ell_1 ℓ 1 -norm minimization (basis pursuit) as the recovery algorithm. We derive new upper and lower bounds on the number of measurements to achieve robust sparse recovery with binary matrices. We establish sufficient conditi…
A new weighted MCC measure improves classifier performance evaluation.
problem Lack of measures sensitive to observation weights in multiclass classification.
method Proposes weighted versions of Pearson-Matthews Correlation Coefficient (MCC) for binary and multiclass classification.
result Weighted MCC values are higher for classifiers that perform better on highly weighted observations.
Confidence intervals improve evaluation of binary prediction rules in data mining.
problem Uncertainty in performance measures estimation from finite datasets.
method Asymptotic normal approximations for confidence intervals, with a blurring correction.
result Improved finite sample coverage probabilities and general performance measures inference.
Networked sensing, where the goal is to perform complex inference using a large number of inexpensive and decentralized sensors, has become an increasingly attractive research topic due to its applications in wireless sensor networks and internet-of-things. To reduce the communication, sensing and storage complexity, t…
Study 1-bit compressive sensing with generative models, improving recovery accuracy.
problem Accurately recover sparse vectors from binary measurements with generative models.
method Analyzes noiseless and noisy 1-bit measurements with i.i.d.~Gaussian and Lipschitz continuous generative priors, proving sample complexity bounds and stability properties.
result Proves sample complexity bounds and stability properties for 1-bit compressive sensing with generative models.
The volume of a credal set correlates with epistemic uncertainty in binary classification but not in multi-class.
problem Representing and quantifying epistemic uncertainty in machine learning.
method Examined the geometric representation of credal sets as d d d -dimensional polytopes and their volume as a measure of uncertainty. result The volume of a credal set is a meaningful measure of epistemic uncertainty in binary classification but not in multi-class.
G-Net constructs binary neural networks with high accuracy using randomized binary embeddings.
problem Creating high-accuracy binary neural networks with theoretical guarantees.
method Proposes a novel floating-point G-Net family with randomized binary embeddings and theoretical accuracy guarantees.
result Empirically, G-Net achieves almost 30% higher accuracy on CIFAR-10 compared to prior HDC models.
New algorithm recovers sparse binary vectors from generalized linear measurements efficiently.
problem Recovering sparse binary vectors from generalized linear measurements.
method Linear estimation algorithm and information theoretic lower bounds.
result Optimal sample complexity of O ( ( k + σ 2 ) log n ) O((k+σ^2)\log{n}) O (( k + σ 2 ) log n ) for noisy one bit quantized linear measurements. Complex performance measures, beyond the popular measure of accuracy, are increasingly being used in the context of binary classification. These complex performance measures are typically not even decomposable, that is, the loss evaluated on a batch of samples cannot typically be expressed as a sum or average of losses…
It has been shown recently that deep convolutional generative adversarial networks (GANs) can learn to generate music in the form of piano-rolls, which represent music by binary-valued time-pitch matrices. However, existing models can only generate real-valued piano-rolls and require further post-processing, such as ha…
A new clustering algorithm for functional data using binary trees.
problem Clustering multivariate functional data with measurement errors.
method Recursive binary tree splitting for data clustering.
result Good performance in various complex settings, including vehicle trajectories.
New algorithms optimize metrics for binary classification with class imbalance.
problem Optimizing metrics like Fβ, AM, Jaccard for imbalanced classes.
method Reformulates metric optimization as cost-sensitive learning, using surrogate loss functions.
result METRO algorithms provide strong theoretical guarantees and outperform baselines.
Binary linear classifiers are the most explainable up to negligible sets.
problem Measuring the explainability of machine learning classifiers.
method Introducing pointwise coverage to measure explainability and proving the binary linear classifier is the most explainable up to negligible sets.
result The binary linear classifier is uniquely the most explainable classifier up to negligible sets.
Quantum circuits represent binary classification trees with binary features.
problem Classifying data using binary classification trees with binary features.
method Quantum circuits and probabilistic approach for traversing decision trees.
result First realization of a decision tree classifier on a quantum device.
Study proposes a differentiable surrogate loss function for optimizing F β F_β F β score in binary classification with imbalanced data.
problem Non-differentiability of F β F_β F β score makes it unsuitable for optimization by gradient-based learning. method Investigated relationship between F β F_β F β score and loss functions, proposed a differentiable surrogate loss function. result Gradient paths of the proposed surrogate F β F_β F β loss function approximate the gradient paths of the F β F_β F β score. An attractive approach for fast search in image databases is binary hashing, where each high-dimensional, real-valued image is mapped onto a low-dimensional, binary vector and the search is done in this binary space. Finding the optimal hash function is difficult because it involves binary constraints, and most approac…
The F-measure, which has originally been introduced in information retrieval, is nowadays routinely used as a performance metric for problems such as binary classification, multi-label classification, and structured output prediction. Optimizing this measure is a statistically and computationally challenging problem, s…
BO algorithms improve binary and preferential optimization by distinguishing between types of uncertainty.
problem Optimizing expensive functions with binary or pairwise comparisons.
method Proposed new acquisition functions distinguishing between epistemic and aleatoric uncertainty.
result New acquisition functions outperform state-of-the-art heuristics in binary and preferential BO.
The paper proposes a least squares method for binary compressive sampling with low intrinsic dimension signals.
problem Recovering signals from binary measurements with noise and sign flips.
method Least squares decoder for signals with low generative intrinsic dimension.
result The least squares decoder achieves a sharp estimation error of O ( k log ( L n ) m ) O(\sqrt{\frac{k\log (Ln)}{m}}) O ( m k l o g ( L n ) ) under certain conditions. The paper tackles binary classification with measure data using topological descriptors.
problem Binary classification with measure data.
method Develops classifiers for measure data using topological descriptors (persistence diagrams).
result Upper and lower bounds on the Rademacher complexity of classifiers on measures.
We propose a method for maximizing a partial area under a receiver operating characteristic (ROC) curve (pAUC) for binary classification tasks. In binary classification tasks, accuracy is the most commonly used as a measure of classifier performance. In some applications such as anomaly detection and diagnostic testing…
New metric reduces arbitrariness in fair binary classification predictions.
problem Variance in predictions leads to arbitrary decisions in fair classification.
method Developed a self-consistency metric and an abstention algorithm.
result Fair binary classification is often close to fair due to variance, not interventions.
New binary loss functions improve density ratio estimation accuracy.
problem Improving accuracy of density ratio estimators using binary classifiers.
method Characterized loss functions based on prescribed error measures in Bregman divergences.
result Novel loss functions prioritize accurate estimation of large density ratio values.
Compressive Sensing (CS) theory asserts that sparse signal reconstruction is possible from a small number of linear measurements. Although CS enables low-cost linear sampling, it requires non-linear and costly reconstruction. Recent literature works show that compressive image classification is possible in CS domain wi…
Develops RF-GLS for binary geospatial data.
problem Challenges in extending RF to binary geospatial data.
method Proposes RF-GLS for binary data, embedding it in generalized mixed effects models.
result Establishes consistency of RF-GP for mean function and covariate effect estimation.
This paper develops convex surrogates for optimizing the multi-label F-measure.
problem Optimizing the F-measure for multi-label classification is computationally hard.
method Designing convex surrogate losses calibrated for the F-measure.
result The F-measure for multi-label problems has a rank of at most s 2 + 1 s^2+1 s 2 + 1 . Develops a continuous compliance index for Islamic equity screening.
problem Binary rulebooks lead to inconsistent compliance assessment of firms.
method Integrates six leading financial and business activity standards into a single continuous index.
result Firms with the same pass/fail label can differ significantly in compliance strength.
This work extends score-based methods to binary data on the Boolean hypercube.
problem Learning and sampling binary data on the Boolean hypercube.
method Adopting Bernoulli noise as a smoothing device, deriving a TMF-like expression for the optimal denoiser, and using a Langevin-like sampler.
result The method successfully samples noisy binary data and reduces effective noise through multiple measurements.
Deep belief networks can approximate any multivariate density with binary hidden units.
problem Approximating multivariate probability densities with binary hidden units.
method Sharp quantitative bounds on approximation error in terms of hidden units.
result Deep belief networks can approximate any multivariate density with binary hidden units under mild integrability requirements.
BinaryDuo improves BNNs by coupling binary activations, outperforming state-of-the-art models.
problem Gradient mismatch in BNNs due to binarizing activations.
method Using gradient of smoothed loss function to estimate gradient mismatch, proposing BinaryDuo scheme with coupled ternary activations.
result BinaryDuo outperforms state-of-the-art BNNs on various benchmarks.
Paper introduces ILD algorithm to determine Bayes error for binary classification.
problem Determining the best possible performance in binary classification problems.
method Model-agnostic ILD algorithm to calculate Bayes error.
result Provides intrinsic limits of any binary classification algorithm.
Binary Iterative Hard Thresholding converges with optimal number of 1-bit measurements.
problem Recovering sparse signals from 1-bit compressed measurements.
method Binary Iterative Hard Thresholding (BIHT) algorithm.
result BIHT converges with only O(k/ε) measurements, optimal for recovery.
The paper derives a new theorem for predicting batches of data.
problem Finding lower bounds on minimal batch regret.
method Derives a conditional version of the regret-capacity theorem.
result Reveals a connection between conditional Rényi divergence and conditional Sibson's mutual information.
Paper analyzes impact of PRM on binary random variables and distribution shifts.
problem Impact of performative risk minimization on binary random variables and distribution shifts.
method Formulated two measures of impact, derived explicit formulas for full information, and provided estimators for partial information.
result PRM can have amplified side effects compared to methods that do not model data shift.
This paper deals with two related problems, namely distance-preserving binary embeddings and quantization for compressed sensing . First, we propose fast methods to replace points from a subset X ⊂ R n \mathcal{X} \subset \mathbb{R}^n X ⊂ R n , associated with the Euclidean metric, with points in the cube { ± 1 } m \{\pm 1\}^m { ± 1 } m and we associa…
Combines BO with context to optimize binary feedback.
problem Optimizing expensive binary functions with context.
method Bayesian active learning and optimization.
result Efficiently chooses best context and parameters.
New method predicts activity coefficients for binary mixtures without using physical descriptors.
problem Predicting activity coefficients for unexplored binary mixtures.
method Probabilistic matrix factorization model.
result Method outperforms state-of-the-art models requiring less training effort.
New ROC tools assess predictive abilities for any linearly ordered outcomes.
problem Fundamental restriction in ROC analysis for non-dichotomous outcomes.
method ROC movies and UROC curves for linearly ordered outcomes.
result CPA equals AUC for binary outcomes and relates to Spearman's coefficient for pairwise distinct outcomes.
New clustering methods for binary data using combinatorial optimization.
problem Clustering binary data efficiently and effectively.
method Five new combinatorial optimization heuristics (SA, TA, TS, GA, ACO) applied to binary data.
result Simulated annealing performs exceptionally well compared to classical methods.
Probit regression was first proposed by Bliss in 1934 to study mortality rates of insects. Since then, an extensive body of work has analyzed and used probit or related binary regression methods (such as logistic regression) in numerous applications and fields. This paper provides a fresh angle to such well-established…
A new method for measuring prediction uncertainty in classifiers.
problem Measuring uncertainty of predictions from machine learning methods.
method Density Based Calibration (DBCal) technique.
result Expected calibration error of less than 0.2% on binary classifiers and less than 3% on semantic segmentation networks.
Quantum algorithm improves sparse vector recovery from noisy measurements.
problem Accurately recover sparse vectors from noisy linear measurements.
method Formulated as a QUBO task, solved using quantum technology.
result Quantum approach outperforms classical methods in sparse coding.
This work examines uncertainty sampling in binary classification using equivalent loss.
problem Lack of consensus on proper uncertainty definition and theoretical guarantees for active learning.
method Systematically examines uncertainty sampling via equivalent loss, proving its optimality.
result Established that uncertainty sampling optimizes against equivalent loss, providing theoretical guarantees.
TCE measures calibration error with a test-based approach.
problem Measuring calibration error of probabilistic binary classifiers.
method TCE uses a novel loss function based on a statistical test.
result TCE offers clear interpretation, consistent scale, and enhanced visual representation.
The study evaluates AI model performance measures for medical use.
problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.