Distribution networks model novel classes in open set learning.
problem Modeling novel classes in open set learning.
method Distribution networks map samples to a latent space where known and novel classes' distributions are jointly learned.
result Distribution networks accurately detect and model novel classes for subsequent classification.
A simple framework predicts unseen classes using exponential family distributions.
problem Learning to predict previously unseen classes.
method Estimating class-attribute-gated class-conditional distributions modeled as exponential families.
result Natural representation of classes as probability distributions, leveraging unlabeled data.
New approach uses class domains for classification when distributions are unknown.
problem Traditional classification rules are inadequate when class distributions are ill-defined or unknown.
method Use class domains instead of class distributions for constructing a reliable decision function.
result Illustrated examples show the effectiveness of the new approach.
Friedman's method performs well for estimating class distributions.
problem Estimating prior class probabilities without label observations.
method Friedman's method and DeBias method for designing linear equation systems.
result Friedman's method performs well for binary and multi-class quantification.
New class of heavy-tailed distributions shows weighted averages dominate individual variables.
problem Understanding and comparing risks in heavy-tailed distributions.
method Introducing a new class of heavy-tailed distributions and proving stochastic dominance relations.
result Weighted averages of random variables in this class are stochastically larger than individual variables.
Generative framework tackles zero-shot learning with adversarial domain adaptation.
problem Domain shift between seen and unseen class distributions in zero-shot learning.
method End-to-end learning of seen and unseen class distributions, adversarial domain adaptation.
result Superior accuracies compared to state-of-the-art models on various benchmark datasets.
Investigates a conjugate prior for Dirichlet distribution.
problem No specific problem stated; focuses on mathematical investigation.
method Investigates a conjugate class for the Dirichlet distribution within the exponential family.
result Identifies a conjugate prior for the Dirichlet distribution.
New bounds on learning from multiple distributions for VC classes.
problem Understanding the sample complexity of learning from multiple data distributions.
method Analyzing the gap between known upper and lower bounds for PAC-learnable classes.
result Recent progress on sample complexity for VC dimension d classes on k distributions.
We provide a complete characterization of the class of one-dimensional time-homogeneous diffusions consistent with a given law at an exponentially distributed time using classical results in diffusion theory. To illustrate we characterize the class of diffusions with the same distribution as Brownian motion at an expon…
New framework for understanding adversarial and stochastic learning.
problem Understanding the continuum from adversarial to stochastic settings in online learning.
method Distributionally constrained adversaries framework.
result Characterization of learnable distribution classes for various function classes.
The paper studies Atiyah and Todd classes for DG manifolds derived from integrable distributions.
problem Understanding Atiyah and Todd classes for DG manifolds.
method Analyzing DG manifolds (F[1],dF) corresponding to integrable distributions F. result Atiyah and Todd classes of DG manifolds are identical to those of Lie pairs $(T_{\mathbb{K}} M, F).
Cloud-based filter blocks privacy-sensitive images using distributed one-class learning.
problem Preventing third parties from uploading privacy-sensitive images to social media.
method Distributed One-Class Learning, autoencoders, edge devices, multi-class filter.
result The proposed filter can cope with imbalanced and complex image distributions and is robust to attacks.
This Colloquium reviews statistical models for money, wealth, and income distributions developed in the econophysics literature since the late 1990s. By analogy with the Boltzmann-Gibbs distribution of energy in physics, it is shown that the probability distribution of money is exponential for certain classes of models…
New method improves ensemble inference for high-class tasks.
problem High inference costs for ensemble models.
method Proxy-Dirichlet target to minimize reverse KL-divergence.
result Resolves gradient issues for large-scale classification tasks.
Algorithm selects private hypothesis from unknown distribution.
problem Private selection of hypothesis from unknown distribution.
method Differentially private algorithm for hypothesis selection.
result Sample complexity of O(α2logm+αεlogm). Study shows realizable learnability doesn't imply agnostic learnability for distributions.
problem Learnability and robustness of distribution classes.
method Analyzes the relationship between learnability and robustness for distribution learning.
result Realizable learnability does not imply agnostic learnability for distributions.
Undirected graphical models, or Markov networks, are a popular class of statistical models, used in a wide variety of applications. Popular instances of this class include Gaussian graphical models and Ising models. In many settings, however, it might not be clear which subclass of graphical models to use, particularly…
Survey of Gauss map value distribution for various surface classes.
problem Understanding the value distribution of Gauss maps for different surface types.
method Analyzing recent results for minimal surfaces, improper affine spheres, and flat surfaces.
result Elucidation of geometric background for the results.
The paper extends logistic regression for unbounded majority classes and derives asymptotic properties.
problem Infinitely imbalanced logistic regression inference.
method Derive a second order expansion for slope parameter under unbounded majority class.
result The second order term converges to a normal distribution with a variance depending only on the minority class's mean.
Paper extends stochastic dominance for compound binomial distributions.
problem Stochastic dominance for infinite-mean random variables.
method Investigates properties and inclusion relationships of distribution classes, extends results to compound binomial distributions.
result Establishes necessary and sufficient conditions for first-order stochastic dominance preservation.
Research examines the distribution of curve components in random multicurves.
problem Distribution of curve components in random multicurves.
method Action of the mapping class group on random multicurves.
result Distribution of curve components analyzed.
Unified approach to non-standard classification tasks.
problem Non-standard classification tasks like semi-supervised, positive-unlabelled, multi-positive-unlabelled and noisy-label learning.
method Probabilistic, unified approach training a classifier to predict label-distributions, then inferring class-distributions.
result Unified model for various non-standard classification tasks.
A Fourier-based learning algorithm for multiclass classification.
problem Highly nonlinear multiclass classification problems.
method Smoothing technique with low-pass filters to calculate probability distributions.
result Probabilistic explanation for classification without kernel functions.
Quantum interference improves clustering accuracy.
problem Improving clustering accuracy in Gaussian mixture models.
method Modeling classes as wave functions, then mixing them.
result Quantum method outperforms Gaussian mixture in all aspects.
Improves SSL with doubly robust estimation of unlabeled class distribution.
problem Limited labeled data and long-tailed class distributions in unlabeled data.
method Explicitly estimate unlabeled class distribution using doubly robust estimator.
result Improves performance of SSL methods on unlabeled data.
We present a framework for online inference in the presence of a nonexhaustively defined set of classes that incorporates supervised classification with class discovery and modeling. A Dirichlet process prior (DPP) model defined over class distributions ensures that both known and unknown class distributions originate …
New approach calibrates predictions for better decision-making.
problem Achieving reliable predictions for multi-class problems is hard.
method Introduces decision calibration, a new approach to calibrate predictions.
result Designs a recalibration algorithm that makes predictions reliable for decision-making.
No single parameter characterizes the learnability of probability distributions.
problem Finding a parameter to characterize the learnability of probability distributions.
method Analyzing various notions of learnability and showing impossibility results.
result No such parameter exists for characterizing learnability of probability distributions.
New algorithms for learning under s-concave distributions, including Pareto and t-distributions.
problem Learning under broad and natural generalizations of log-concave distributions, including fat-tailed ones.
method Introduce new convex geometry tools to study s-concave distributions and use these properties to provide bounds on learning quantities. result Significantly generalize prior results for margin-based, disagreement-based, and passive learning of intersections of halfspaces.
Solves biased pseudo-labels in imbalanced SSL by refining them.
problem Imbalanced class distributions in semi-supervised learning lead to biased pseudo-labels.
method Formulates a convex optimization problem to refine pseudo-labels and develops an efficient algorithm, DARP.
result Demonstrates the effectiveness of DARP in various imbalanced semi-supervised scenarios.
This paper tackles domain generalization by learning invariant class conditional distributions.
problem Learning invariant representations across different domains with varying distributions.
method Proposes a conditional invariant representation to ensure invariance of class conditional distributions.
result Guarantees invariance of the joint distribution P(h(X),Y) if class prior P(Y) remains invariant. This paper explores using KDE for balanced sampling in imbalanced datasets.
problem Imbalanced class distribution in data science.
method Kernel density estimation (KDE) for resampling the minority class.
result KDE-based resampling outperforms other techniques in F1-score and G-mean.
Optimal income crossover found using particle swarm optimization.
problem Determining the crossover point between two income distributions.
method Particle swarm optimization for finding crossover income, temperature, and Pareto index.
result Optimization method finds boundaries of two income distributions.
Review of multivariate Poisson-based distributions for count data.
problem Dependencies in high-dimensional count data.
method Categorization and empirical comparison of multivariate Poisson-based distributions.
result Empirical comparison of multivariate Poisson-based distributions on real-world datasets.
AGGAN uses genetic algorithm with simulated annealing to generate minority class data.
problem Overcoming class imbalance in minority class data.
method AGGAN combines genetic algorithm and simulated annealing to train GANs on scarce minority class data.
result AGGAN effectively generates minority class data distributions from limited samples.
New method accelerates diffusion models for broader target distributions.
problem Current diffusion models have limited acceleration for certain target distributions.
method Developed a novel accelerated stochastic DDPM sampler.
result Achieved accelerated performance for three broad distribution classes.
Example shows learnable distributions not privately learnable.
problem Learnable distributions under non-private conditions not transferable to differential privacy.
method Example of a distribution class learnable up to constant error in total variation distance but not under differential privacy.
result Contradicts conjecture of Ashtiani on learnability under differential privacy.
Study on distributed nonparametric function estimation with optimal rate and cost of adaptation.
problem Optimal rate of convergence and cost of adaptation in distributed nonparametric function estimation.
method Distributed minimax estimation and adaptive estimation under communication constraints for Gaussian sequence model and white noise model.
result Established minimax rate of convergence and exact communication cost for adaptation.
Proposes a multimodal deep generative model for semi-supervised learning with class imbalance.
problem Class imbalance in semi-supervised learning with partial supervision.
method Separate encoders for each modality, sharing latent variables, and using Student's t-distributions for prior, encoder, and decoder.
result Outperforms baseline methods in generalization and classification performance for partially labeled multimodal data.
Paper optimizes classification of distributions using Wasserstein metric.
problem Classifying instances represented by distributions on a vector space.
method Maximizing Fisher's ratio in the Wasserstein metric space through iterative algorithm.
result The method enhances classification performance and is robust to variations in distribution summaries.
Characterizes distribution-free rates in unbalanced classification problems.
problem Minimizing error under two different distributions in unbalanced settings.
method Characterizes minimax rates over all pairs of distributions using a geometric condition.
result Identifies a dichotomy between hard and easy classes based on a three-points-separation condition.
New bounds for contrastive learning handle domain shifts and generalization.
problem Domain shifts and generalization challenges in downstream tasks.
method Novel generalization bounds accounting for both domain shift and generalization.
result Performance of contrastively learned representations depends on statistical discrepancy between pretraining and downstream distributions.
A new sampling method balances imbalanced data using gamma distribution.
problem Imbalanced class distribution in data causes bias in classification algorithms.
method Intelligent resampling of minority class instances via gamma distribution.
result The proposed method outperforms existing techniques on 12 out of 24 datasets.
New CPS model tackles conditional probability shift in machine learning.
problem Discrepancy between source and target distributions in machine learning.
method Conditional Probability Shift Model (CPSM) using multinomial regression and EM algorithm.
result Superior balanced classification accuracy on target data compared to existing methods.
Efficiently estimates densities of multidimensional shift-invariant distributions.
problem Density estimation for shift-invariant multidimensional distributions.
method Efficient algorithms for learning any distribution in the class from samples, using total variation distance.
result Shift-invariant distributions can be learned efficiently with a number of samples and time proportional to 1/εd+2 and 1/ε2d+2 respectively. In many real-world classification problems, the labels of training examples are randomly corrupted. Most previous theoretical work on classification with label noise assumes that the two classes are separable, that the label noise is independent of the true class label, or that the noise proportions for each class are …
Paper establishes sufficient condition for comparing linear combinations of infinite-mean risks.
problem Comparing linear combinations of infinite-mean risks under stochastic dominance.
method Introduced a new class of distributions and used majorization order to compare weights.
result Linear combinations of random variables are stochastically larger when their weight vectors are smaller in majorization order.
PROTOCOL tackles imbalanced multi-view clustering by enhancing contrastive learning.
problem Class imbalance in real-world multi-view data.
method PROTOCOL uses partial optimal transport to perceive and mitigate imbalance, enhancing contrastive learning.
result PROTOCOL significantly improves clustering performance on imbalanced multi-view data.