In high-dimensional data, structured noise caused by observed and unobserved factors affecting multiple target variables simultaneously, imposes a serious challenge for modeling, by masking the often weak signal. Therefore, (1) explaining away the structured noise in multiple-output regression is of paramount importanc…
The paper extends SARMA models by relaxing independence assumptions on error terms.
problem Testing adequacy of SARMA models with non-independent errors.
method Study of asymptotic distributions of residual and normalized residual empirical autocovariances and autocorrelations under weak noise assumptions.
result Established asymptotic behavior of portmanteau tests for SARMA models.
Boosting algorithm reduces error in noisy data.
problem Improving weak learners in the presence of Massart noise.
method First computationally efficient boosting algorithm for Massart noise.
result Achieves misclassification error arbitrarily close to Massart noise threshold.
New theory for PCA under weak latent factors, improving inference and testing.
problem Statistical inference for PCA with weak latent factors and cross-sectional dependence.
method Comprehensive estimation and inference theory for PCA under nearly minimal factor strength, non-asymptotic.
result Asymptotic normality of PCA-based estimator for N≍T with SNR growth rate. There is increasing interest in learning algorithms that involve interaction between human and machine. Comparison-based queries are among the most natural ways to get feedback from humans. A challenge in designing comparison-based interactive learning algorithms is coping with noisy answers. The most common fix is to …
Paper solves robust convex problems with heavy-tailed noise.
problem Solving convex compositional problems with heavy-tailed noise.
method Sub-Gaussian confidence bounds under weak heavy-tailed noise assumptions, using boosting strategy.
result Achieves nearly optimal high probability convergence result.
Study agnostic feature-based dynamic pricing models with linear policies and noisy valuations.
problem Tackles dynamic pricing with unknown noise and no assumptions on data.
method Studies two agnostic models: linear policy and linear noisy valuation, presenting algorithms and regret bounds.
result Demonstrates no-regret learning is possible under weak assumptions, but noisy feedback is not significantly more useful than bandit feedback.
We construct explicitly a bridge process whose distribution, in its own filtration, is the same as the difference of two independent Poisson processes with the same intensity and its time 1 value satisfies a specific constraint. This construction allows us to show the existence of Glosten-Milgrom equilibrium and its as…
We consider the prediction of weak effects in a multiple-output regression setup, when covariates are expected to explain a small amount, less than ≈1, of the variance of the target variables. To facilitate the prediction of the weak effects, we constrain our model structure by introducing a novel Bayesian ap…
WSINDy for PDEs robustly identifies models from noisy data.
problem Identifying nonlinear dynamics from noisy partial differential equations data.
method Weak formulation of PDEs, Fourier-based model identification, sequential-thresholding least-squares.
result WSINDy enables robust identification of PDEs in noisy conditions.
Survey of weak form's role in equation learning, parameter estimation, and coarse graining.
problem Noise robustness, accuracy, and computational efficiency in weak form applications.
method Survey and recent developments in weak form versions of equation learning, parameter estimation, and coarse graining.
result Surprising noise robustness, accuracy, and computational efficiency in weak form applications.
Study improves robust nonparametric regression in heavy-tailed noise.
problem Robust nonparametric regression with heavy-tailed noise and unbounded functions.
method Huber regression in reproducing kernel Hilbert spaces (RKHS), probabilistic effective hypothesis space, new comparison theorems.
result Explicit finite-sample error bounds and convergence rates for Huber regression in RKHS under heavy-tailed noise.
The paper develops a method to model high-dimensional data with many variables and weak signals.
problem Modeling high-dimensional dependent data with many explanatory variables and low signal-to-noise ratio.
method Penalized regression for high-dimensional data, factor modeling of residuals, high-dimensional white noise testing, projected Principal Component Analysis.
result Established asymptotic properties of the proposed method for high-dimensional data.
Paper analyzes weak-to-strong generalization in CNNs, identifying data-scarce and data-abundant regimes.
problem Weak-to-strong generalization in CNNs trained on weak models.
method Formal analysis of gradient descent dynamics in data-scarce and data-abundant regimes.
result Identifies two regimes and distinct mechanisms of generalization in each.
New methods lift weak supervision to structured prediction, providing robustness guarantees.
problem Applying weak supervision techniques to structured prediction problems.
method Introducing pseudo-Euclidean embeddings, tensor decompositions, and invariants for consistent noise rate estimation.
result Generalization guarantees nearly identical to those for models trained on clean data.
We discuss the maximum modulus principle, and weak unique continuation, for CR functions on an abstract almost CR manifold M. We investigate these matters under the assumption of weak pseudoconcavity, and obtain sharp results about propagation along Sussmann leaves.
StoSOO optimistically maximizes noisy, locally smooth functions.
problem Global maximization of noisy, locally smooth functions with unknown semi-metric.
method StoSOO uses optimistic upper confidence bounds to iteratively decide on the next evaluation point.
result StoSOO performs almost as well as the best tuned algorithms, even without knowing the semi-metric.
Recently, Petrik et al. demonstrated that L1Regularized Approximate Linear Programming (RALP) could produce value functions and policies which compared favorably to established linear value function approximation techniques like LSPI. RALP's success primarily stems from the ability to solve the feature selection and va…
Deep learning reduces noise in weak lensing mass maps using GANs.
problem Noise reduction in weak lensing mass maps.
method Generative adversarial networks (GANs) applied to Subaru Hyper Suprime-Cam data.
result GANs successfully reproduce non-Gaussian information in denoised maps, showing stronger cosmological dependence.
To improve the off-sample generalization of classical procedures minimizing the empirical risk under potentially heavy-tailed data, new robust learning algorithms have been proposed in recent years, with generalized median-of-means strategies being particularly salient. These procedures enjoy performance guarantees in …
Formalizes weak and strong verification for LLMs, controlling errors without assumptions.
problem Balancing cost and reliability in reasoning with LLMs.
method Formalizes weak-strong verification policies, introduces metrics, develops online algorithm.
result Optimal policies admit a two-threshold structure, and calibration and sharpness govern value of weak verifiers.
Study tests adequacy of FARIMA models with uncorrelated but non-independent errors.
problem Testing adequacy of FARIMA models with specific error characteristics.
method Derive asymptotic distributions of residual autocovariances and autocorrelations, propose self-normalization approach.
result Asymptotic distributions of modified portmanteau statistics for weak FARIMA models.
New bounds for RFM-KRR with weak assumptions and easy verification.
problem Establishing accurate out-of-sample bounds for RFM-KRR.
method Elementary linear algebra and weak assumptions.
result Novel out-of-sample error upper and lower bounds with weak assumptions.
Study uses weak transport for non-convex costs in fixed-income markets.
problem Characterizing optimal caplet pricing in fixed-income markets.
method Introduced weak optimal transport for non-convex costs, reduced general costs to convex problems.
result Established robust super-replication results for fixed-income markets.
Sharp estimates lead to new comparison theorems in Riemannian and Kähler geometry.
problem Developing precise geometric inequalities for curvature assumptions.
method Quantitative Laplacian estimates and integral curvature assumptions.
result Derive quantitative comparison theorems for Riemannian and Kähler manifolds.
We consider the weak detection problem in a rank-one spiked Wigner data matrix where the signal-to-noise ratio is small so that reliable detection is impossible. We propose a hypothesis test on the presence of the signal by utilizing the linear spectral statistics of the data matrix. The test is data-driven and does no…
Bayesian priors improve neural network performance on weak signals.
problem Challenges in encoding domain knowledge for weak signals in neural networks.
method Proposed a new joint prior over local scale parameters for feature sparsity and signal-to-noise ratio, optimized with Stein gradient.
result Improved prediction accuracy on various datasets, including genetics applications with weak and sparse signals.
We study the task of online boosting--combining online weak learners into an online strong learner. While batch boosting has a sound theoretical foundation, online boosting deserves more study from the theoretical perspective. In this paper, we carefully compare the differences between online and batch boosting, and pr…
Paper develops a new algorithm for sparse signal recovery.
problem Sparse signal recovery from noisy observations.
method Iterative Stochastic Optimization using Stochastic Mirror Descent.
result Linear convergence during preliminary phase of the routine.
Boosting classifiers improve accuracy with noisy inputs.
problem Noisy communication or computation degrades boosting classifier accuracy.
method Optimize resource allocation for base classifiers based on importance metrics.
result Optimized noisy boosting classifiers are more robust than bagging.
WSINDy algorithm proves robust to noise in identifying differential equations.
problem Identifying differential equations from noisy data.
method Weak-form sparse identification of nonlinear dynamics (WSINDy) algorithm.
result WSINDy is asymptotically consistent for a wide class of models, including Navier-Stokes and Kuramoto-Sivashinsky equations.
Sharp regularity for Pfaff system leads to isometric immersions in arbitrary dimensions.
problem Existence and regularity of isometric immersions in arbitrary dimensions.
method Proving W1,2-regularity for Pfaff system with antisymmetric L2-coefficient matrix. result Equivalence between W2,2-isometric immersions and weak solubility of Gauss--Codazzi--Ricci equations. The paper proves properties of curves in Riemannian manifolds.
problem Characterizing curves in Riemannian manifolds.
method Analyzing locally minimizing and weak geodesics.
result Locally minimizing curves are weak geodesics under certain conditions.
Long time existence and convergence to a circle is proved for radial graph solutions to a mean curvature type curve flow in warped product surfaces (under a weak assumption on the warp potential of the surface). This curvature flow preserves the area enclosed by the evolving curve, and this fact is used to prove a gene…
Improved PINNs for solving PDEs with unknown measurement noise.
problem Handling non-Gaussian noise in physics-informed neural networks.
method Jointly train an EBM to learn the correct noise distribution.
result Improved performance in solving PDEs with non-Gaussian noise.
Ensemble techniques are powerful approaches that combine several weak learners to build a stronger one. As a meta learning framework, ensemble techniques can easily be applied to many machine learning techniques. In this paper we propose a neural network extended with an ensemble loss function for text classification. …
Algorithm learns halfspaces robust to Massart noise without distribution knowledge.
problem Learning halfspaces in noisy data with arbitrary marginal distribution.
method Poly-time algorithm for distribution-independent PAC learning.
result Achieves misclassification error of η+ε with poly(d, 1/ε) time complexity.
For binary classification we establish learning rates up to the order of n−1 for support vector machines (SVMs) with hinge loss and Gaussian RBF kernels. These rates are in terms of two assumptions on the considered distributions: Tsybakov's noise assumption to establish a small estimation error, and a new geometr…
We study Principal Component Analysis (PCA) in a setting where a part of the corrupting noise is data-dependent and, as a result, the noise and the true data are correlated. Under a bounded-ness assumption on the true data and the noise, and a simple assumption on data-noise correlation, we obtain a nearly optimal samp…
Enhanced consistency bounds derived for classification under a new noise condition.
problem Enhanced consistency bounds for classification under a new noise condition.
method Model Margin Noise (MM noise) assumption, derived enhanced H-consistency bounds.
result Enhanced H-consistency bounds under MM noise condition, interpolates between linear and square-root regimes.
Unified framework for policy learning using weak supervision.
problem High-quality supervision is often infeasible or expensive in practice.
method Treat weak supervision as imperfect peer information and evaluate policies based on correlated agreement.
result Substantial performance improvements, especially in complex or noisy environments.
New method bounds hardware noise without assumptions.
problem Estimating hardware noise without assumptions.
method Machine Learning and Conformal Prediction.
result Theoretical upper bounds of fidelity.
Improves SGM convergence bounds in W2-distance without strict assumptions.
problem Convergence bounds for SGMs in W2-distance require stringent assumptions.
method Novel framework using the OU process and PDE analysis.
result Log-concavity evolves from weak to strong over time.
We propose a projected semi-stochastic gradient descent method with mini-batch for improving both the theoretical complexity and practical performance of the general stochastic gradient descent method (SGD). We are able to prove linear convergence under weak strong convexity assumption. This requires no strong convexit…
AdaBoost improves binary classification in robust one-bit compressed sensing with adversarial errors.
problem Binary classification in robust one-bit compressed sensing with adversarial errors.
method AdaBoost and max-ℓ1-margin-classifier approach, with convergence rates improved under certain feature conditions. result Improved convergence rates and explanation for harmless interpolating adversarial noise.
We demonstrate the potential of Deep Learning methods for measurements of cosmological parameters from density fields, focusing on the extraction of non-Gaussian information. We consider weak lensing mass maps as our dataset. We aim for our method to be able to distinguish between five models, which were chosen to lie …
Study shows how noise can ensure solutions to fluid dynamics equations.
problem Ensuring unique solutions to stochastic fluid dynamics equations.
method Extended existing results to linear advection of k-forms, proving existence and uniqueness of weak L^p-solutions.
result Proved existence and uniqueness of weak L^p-solutions to stochastic linear advection equation of k-forms.
Proves inextendibility of weak null singularities from curvature blow-up.
problem Inextendibility of weak null singularities in the context of curvature blow-up.
method Introduces a new strategy to infer Cloc0,1-inextendibility from curvature blow-up. result Expected to contribute to the resolution of strong cosmic censorship conjecture.