A new method for weakly supervised learning that improves model accuracy.
problem Training machine learning models with precise labels is expensive; weak supervision provides a low-cost alternative.
method Data consistent weak supervision algorithm that searches over classifiers to find plausible labelings, considering features of the training data and estimating labels for low/no coverage data.
result Empirically, the method significantly outperforms state-of-the-art weak supervision methods on text and image classification tasks.
Study on tradeoffs between mistakes and ERM oracle calls in online and transductive learning.
problem Analyzing online and transductive learning with limited ERM and weak consistency oracle access.
method Proves lower bounds and upper bounds on mistakes and oracle calls, considering realizable and agnostic cases.
result Achieves optimal mistake bounds with weak consistency queries for certain concept classes.
The paper defines a new equivalence relation for knot projections and finds an infinite number of distinct classes.
problem Classifying knot projections based on weak homotopy equivalence.
method Defining weak (1, 2, 3) homotopy and using it to find an invariant.
result There are an infinite number of weak (1, 2, 3) homotopy equivalence classes of knot projections.
Tensor completion requires fewer samples with weak side information.
problem Tensor completion with limited samples and side information.
method Algorithm utilizing weak side information to reduce sample complexity.
result Consistent estimator with O(n1+κ) samples for any small constant κ>0. The paper establishes conditions for Bayesian consistency in supremum metric.
problem Ensuring Bayesian consistency in the supremum metric.
method Using a triangle inequality and weak convergence, the paper establishes conditions for Bayesian consistency.
result Demonstrates supremum consistency with weaker conditions than previously used.
Weak diffusion priors can still perform well in inverse problems.
problem Using mismatched or low-fidelity diffusion priors in inverse problems.
method Extensive experiments and theoretical analysis combining Bayesian-consistency theory and local-correlation analysis.
result Weak priors succeed when measurements are highly informative, and they fail in other regimes.
Study on consistency of ML methods for moving objects in non-stationary environments.
problem Consistency of machine learning methods for moving objects in non-stationary environments.
method Least squares, ridge regression, and ℓs-penalized least squares methods under non-stationary spatial-temporal sampling. result Consistency and asymptotic normality of the estimates under weak conditions.
Study shows financial value of weak information converges in discrete vs continuous markets.
problem Analyzing financial value of weak information in discrete vs continuous markets.
method Defined minimal probability measure and financial value of weak information, then showed convergence.
result Financial value of weak information converges in discrete vs continuous markets.
New methods lift weak supervision to structured prediction, providing robustness guarantees.
problem Applying weak supervision techniques to structured prediction problems.
method Introducing pseudo-Euclidean embeddings, tensor decompositions, and invariants for consistent noise rate estimation.
result Generalization guarantees nearly identical to those for models trained on clean data.
New method improves IV estimation with many weak and invalid instruments.
problem Identification in linear IV models with unknown validity.
method Non-convex penalized approaches, surrogate sparsest penalty.
result Advantages over other IV estimators in selection consistency and weak IV strength conditions.
The problem of learning from label proportions (LLP) involves training classifiers with weak labels on bags of instances, rather than strong labels on individual instances. The weak labels only contain the label proportion of each bag. The LLP problem is important for many practical applications that only allow label p…
Introduces a new framework for Riemannian diffeology.
problem No specific problem stated; focuses on a new framework.
method Uses tangent functor and metric from Iglesias-Zemmour to establish weak Riemannian diffeological spaces.
result Establishes a category of weak Riemannian diffeological spaces and shows induced pseudodistance is a distance under technical conditions.
This manuscript studies statistical properties of linear classifiers obtained through minimization of an unregularized convex risk over a finite sample. Although the results are explicitly finite-dimensional, inputs may be passed through feature maps; in this way, in addition to treating the consistency of logistic reg…
Formalizes weak and strong verification for LLMs, controlling errors without assumptions.
problem Balancing cost and reliability in reasoning with LLMs.
method Formalizes weak-strong verification policies, introduces metrics, develops online algorithm.
result Optimal policies admit a two-threshold structure, and calibration and sharpness govern value of weak verifiers.
WSINDy algorithm proves robust to noise in identifying differential equations.
problem Identifying differential equations from noisy data.
method Weak-form sparse identification of nonlinear dynamics (WSINDy) algorithm.
result WSINDy is asymptotically consistent for a wide class of models, including Navier-Stokes and Kuramoto-Sivashinsky equations.
Boosting weak learners to strong ones from aggregate labels is possible for LLP but not for MIL.
problem Boosting weak learners to strong ones from aggregate labels in learning from label proportions (LLP).
method Using a weak learner on large enough bags to obtain a strong learner for small bags in polynomial time.
result Boosting is possible for LLP but not for MIL.
Spectral algorithm recovers community structure in sparse hypergraphs.
problem Community detection in sparse random hypergraphs with community structure and higher-order interactions.
method Spectral algorithm with three steps: hyperedge selection, spectral partition, and correction/merging.
result Weak consistency achieved for weak signal-to-noise ratio.
We consider stationary autoregressive processes with coefficients restricted to an ellipsoid, which includes autoregressive processes with absolutely summable coefficients. We provide consistency results under different norms for the estimation of such processes using constrained and penalized estimators. As an applica…
Large learning rates cause oscillations in NN weights that improve generalization.
problem Improving generalization of neural networks trained with large learning rates.
method Theoretical analysis and feature-noise data generation model.
result Oscillating SGD with large learning rates benefits NN generalization by effectively learning weak features.
The problem of prescribing conformally the scalar curvature of a closed Riemannian manifold as a given Morse function reduces to solving an elliptic partial differential equation with critical Sobolev exponent. Two ways of attacking this problem consist in subcritical approximations or negative pseudo gradient flows. W…
We consider a jump-type Cox--Ingersoll--Ross (CIR) process driven by a standard Wiener process and a subordinator, and we study asymptotic properties of the maximum likelihood estimator (MLE) for its growth rate. We distinguish three cases: subcritical, critical and supercritical. In the subcritical case we prove weak …
WRENCH benchmarks weak supervision datasets for machine learning.
problem Lack of standardized evaluation for weak supervision datasets.
method Developed a comprehensive benchmark platform WRENCH.
result Demonstrated the efficacy of WRENCH as a benchmark platform.
Unified approach for learning with weak labels across various tasks.
problem Learning with noisy or incomplete labels in diverse machine learning settings.
method Implicit posterior models for joint label inference.
result Unified training objective for various machine learning tasks.
Framework for quantifying uncertainty in dynamic processes.
problem Quantifying uncertainty in dynamic stochastic processes.
method Define dynamic uncertainty sets and dynamic robust risk measures.
result Dynamic robust risk measures are time-consistent under specific uncertainty sets.
An active learner is given a hypothesis class, a large set of unlabeled examples and the ability to interactively query labels to an oracle of a subset of these examples; the goal of the learner is to learn a hypothesis in the class that fits the data well by making as few label queries as possible. This work addresses…
New heat flow for harmonic maps avoids singularities but not bubbles.
problem Finite time singularities in harmonic maps.
method Introduces a conformal heat flow for harmonic maps defined by an evolution equation.
result Global weak solution exists, smooth except at most finitely many points.
A family of Markov blankets in a faithful Bayesian network satisfies the symmetry and consistency properties. In this paper, we draw a bijection between families of consistent Markov blankets and moral graphs. We define the new concepts of weak recursive simpliciality and perfect elimination kits. We prove that they ar…
Proposes a new estimator for weak instrumental variables in panel data models.
problem Weak instrumental variables due to ignored nonlinearities in panel data.
method Triangular simultaneous equation model with a nonlinear reduced form equation and a control function approach using Super Learner.
result The proposed SLCF estimator is consistent and asymptotically normal, achieving a parametric rate of convergence.
Novel weak solutions for volume-preserving mean curvature flow established.
problem Existence and uniqueness of solutions to volume-preserving mean curvature flow.
method Introducing varifold solutions coupled with phase volumes and new calibrations.
result Uniqueness of classical solutions among varifold solutions.
Paper uses ensemblers to predict sepsis early from patient records.
problem Early detection of sepsis in patients.
method Imputation and weak ensembler technique applied to 40k patient records.
result Model achieved 93.45% accuracy and 0.271 utility score.
New theory explains consistency of kernel methods with non-i.i.d. data.
problem Consistency of kernel methods under non-i.i.d. data.
method Empirical weak convergence (EWC) as a general assumption.
result Established consistency of SVMs, kernel mean embeddings, and CKMEs with EWC data.
New results on inferring hidden states in trackable weak models.
problem Inferring hidden states in trackable weak models.
method Analyzing strongly-connected trackable weak models and reconstructing branch choices.
result The number of hypotheses in strongly-connected trackable models is bounded by a constant.
Deep neural nets learn from weakly dependent processes.
problem Learning from ψ-weakly dependent processes. method Deep neural networks for ψ-weakly dependent processes. result Established consistency of empirical risk minimization algorithm and generalization bound.
We improve random forest consistency and performance with DMRF, a new variant.
problem Improving the consistency and performance of random forest models.
method Developed DMRF, a data-driven multinomial random forest, by modifying proof methods and improving data utilization.
result DMRF achieves strong consistency with probability 1, surpassing previous models in classification tasks.
Enhances random forest consistency and introduces DMRF for improved performance.
problem Improving the consistency and efficiency of random forest algorithms.
method Strengthened proof methods and propose DMRF algorithm.
result DMRF achieves better theoretical and experimental performance than previous variants.
We focus on spectral clustering of unlabeled graphs and review some results on clustering methods which achieve weak or strong consistent identification in data generated by such models. We also present a new algorithm which appears to perform optimally both theoretically using asymptotic theory and empirically.
RaSE ensemble framework improves sparse classification accuracy.
problem Sparse classification challenges in high-dimensional data.
method Random Subspace Ensemble (RaSE) framework with subspace selection via RIC.
result RaSE achieves low misclassification rates and accurate feature ranking.
New method identifies causal relationships without strong assumptions.
problem Causal Representation Learning (CRL) is ill-posed due to representation and causal discovery issues.
method Identifiability based on grouping of observational variables, self-supervised estimation framework.
result Practical identifiability conditions without temporal structure, interventions, or weak supervision.
Localized SVMs maintain SVM's consistency properties for large datasets.
problem Inefficient computational requirements of global SVMs for large data sets.
method Localized SVMs apply different hyperparameters to different regions of the input space.
result Localized SVMs inherit Lp- and risk consistency from global SVMs. Proposes efficient bounds for causal effect estimation under weak confounding.
problem Estimating causal effects with weakly confounded variables.
method Develops an efficient linear program to derive upper and lower bounds on causal effect under small entropy of unobserved confounders.
result Bounds are consistent and tighter for weakly confounded variables.
The Efficient Market Hypothesis has been a staple of economics research for decades. In particular, weak-form market efficiency -- the notion that past prices cannot predict future performance -- is strongly supported by econometric evidence. In contrast, machine learning algorithms implemented to predict stock price h…
Study the limits of discrete DPPs to continuous DPPs as set size grows.
problem Characterize the behavior of discrete DPPs as they approach continuous DPPs.
method Non-asymptotic characterization of the limit in terms of weak coherency.
result Sufficient conditions for weak coherency are identified.
Benchmark evaluates financial misinformation detection models, revealing weaknesses without external context.
problem Detecting financial misinformation without external references.
method RFC Bench at paragraph level, two tasks: reference-free detection and comparison-based diagnosis.
result Performance improves with comparative context, revealing model weaknesses in reference-free settings.
We describe various equivalent ways of associating to an orbifold, or more generally a higher étale differentiable stack, a weak homotopy type. Some of these ways extend to arbitrary higher stacks on the site of smooth manifolds, and we show that for a differentiable stack X arising from a Lie groupoid G, the weak homo…
Introduces NL bialgebras combining Lie and Nijenhuis structures.
problem Developing algebraic structures for integrable systems.
method Introducing (weak) NL bialgebras with specific compatibility conditions.
result NL bialgebras generate compatible hierarchies of bialgebras.
Deep learning detects novel changes in time series data.
problem Detecting novel changes in time series with unknown probability structures.
method Causally extracts an innovations sequence for novelty detection.
result Minimax optimality established for the novelty detection method.
We define Conditional quasi concave Performance Measures (CPMs), on random variables bounded from below, to accommodate for additional information. Our notion encompasses a wide variety of cases, from conditional expected utility and certainty equivalent to conditional acceptability indexes. We provide the characteriza…
We address a class of schemes for the Euler equations with the following features: the space discretization is staggered, possible upwinding is performed with respect to the material velocity only and the internal energy balance is solved, with a correction term designed on consistency arguments. These schemes have bee…