New conditions ensure points can be uniquely represented by combinations of variety elements.
problem Ensuring points can be uniquely represented by combinations of variety elements.
method Conditions on contact locus of general linear spaces.
result Conditions ensuring non tangential weak defectiveness of projective varieties.
Paper studies identifiability and stability of drifting fields in generative modeling.
problem Identify and stabilize drifting fields in generative modeling.
method Introduces companion-elliptic kernel families to address limitations of Laplace kernel.
result Establishes field identifiability and demonstrates scalar observables for weak convergence.
The paper explores identifiability and stability in drifting fields using companion-elliptic kernels.
problem Identifying and stabilizing drifting fields in generative modeling.
method Introduces companion-elliptic kernel families and analyzes their properties to address identifiability and stability issues.
result Established field identifiability for arbitrary Borel probability measures and demonstrated that field convergence alone does not guarantee weak convergence.
Paper analyzes weak-to-strong generalization in CNNs, identifying data-scarce and data-abundant regimes.
problem Weak-to-strong generalization in CNNs trained on weak models.
method Formal analysis of gradient descent dynamics in data-scarce and data-abundant regimes.
result Identifies two regimes and distinct mechanisms of generalization in each.
Proposes efficient bounds for causal effect estimation under weak confounding.
problem Estimating causal effects with weakly confounded variables.
method Develops an efficient linear program to derive upper and lower bounds on causal effect under small entropy of unobserved confounders.
result Bounds are consistent and tighter for weakly confounded variables.
RAVEN improves weak-to-strong generalization under distribution shifts.
problem Weak models fail to supervise strong models effectively under distribution shifts.
method RAVEN dynamically learns optimal combinations of weak models and strong model parameters.
result RAVEN outperforms existing methods by over 30% on out-of-distribution tasks.
Boosting combines weak classifiers to form highly accurate predictors. Although the case of binary classification is well understood, in the multiclass setting, the "correct" requirements on the weak classifier, or the notion of the most efficient boosting algorithms are missing. In this paper, we create a broad and ge…
New method identifies latent variables with sparse perturbations.
problem Identifying latent variables with minimal supervision.
method Weakly supervised representation learning with sparse perturbations.
result Identification of latent variables up to specified blocks.
New model shows weak teachers can help strong students learn even with imperfect labels.
problem Improving strong student's performance with weak teacher's imperfect pseudolabels.
method Stylized overparameterized spiked covariance model with Gaussian covariates, proving two phases of generalization.
result Provable successful and random guessing phases of strong student's generalization.
Self-test loss functions improve data-driven modeling of weak-form operators and gradient flows.
problem Challenges in selecting test functions for data-driven modeling involving weak-form operators and gradient flows.
method Introducing self-test loss functions that depend on unknown parameters and are quadratic.
result Self-test loss functions conserve energy for gradient flows and coincide with log-likelihood ratios for stochastic differential equations.
The paper provides a framework for weakly supervised disentanglement guarantees.
problem Learning disentangled representations in real-world data.
method Theoretical framework for analyzing disentanglement guarantees with weak supervision.
result Empirical verification of weak supervision methods' predictive power and usefulness.
Online algorithm identifies PDEs from noisy data snapshots.
problem Identifying PDEs from sequential solution snapshots.
method Combines weak-form discretization with online proximal gradient descent.
result Efficiently identifies and tracks systems with time-varying coefficients.
Weak diffusion priors can still perform well in inverse problems.
problem Using mismatched or low-fidelity diffusion priors in inverse problems.
method Extensive experiments and theoretical analysis combining Bayesian-consistency theory and local-correlation analysis.
result Weak priors succeed when measurements are highly informative, and they fail in other regimes.
David Hilbert discovered in 1895 an important metric that is canonically associated to any convex domain Ω in the Euclidean (or projective) space. This metric is known to be Finslerian, and the usual proof assumes a certain degree of smoothness of the boundary of Ω and refers to a theorem by Busemann and Mayer that…
Framework identifies population quantities from MNAR feedback using weak shadow variables from pretrained models.
problem Estimating mean outcomes from MNAR user feedback with bias and lack of identification.
method Develops a partial identification framework using linear programs and weak shadow variables from pretrained models.
result Bounds on estimand are obtained by solving linear programs incorporating pretrained model predictions.
The study explores the strengths and weaknesses of models that generalize from weak to strong supervision.
problem Understanding the limitations and capabilities of models that generalize from weak to strong supervision.
method Theoretical analysis and experimental validation in both classification and regression settings.
result Theoretical bounds reveal the importance of strong generalization and calibration of the weak model and a careful balance in the training process.
New method identifies causal relationships without strong assumptions.
problem Causal Representation Learning (CRL) is ill-posed due to representation and causal discovery issues.
method Identifiability based on grouping of observational variables, self-supervised estimation framework.
result Practical identifiability conditions without temporal structure, interventions, or weak supervision.
Weak supervision enables learning causal representations from unstructured data.
problem Learning high-level causal representations from unstructured data like images.
method Weakly supervised setting with paired samples before and after interventions. Implicit latent causal models using variational autoencoders.
result Models can reliably identify causal structure and disentangle causal variables.
New proof shows how to identify DAGs with weakly increasing errors.
problem Identifying the true DAG in models with weakly increasing error variances.
method Minimum-trace DAG method and hill climbing algorithm with R2R neighborhood.
result Hill climbing algorithm without strict local optima under weakly increasing error variances.
Study reveals weak knotting in confined polymers, not dominated by any single knot type.
problem Characterizing knotting in open, confined polymers.
method Modeling open curves as virtual knots, comparing lattice walks and ideal chains in confined and unconfined conditions.
result Weak knotting is a common feature in confined polymers, not dominated by any single knot type.
WSINDy for PDEs robustly identifies models from noisy data.
problem Identifying nonlinear dynamics from noisy partial differential equations data.
method Weak formulation of PDEs, Fourier-based model identification, sequential-thresholding least-squares.
result WSINDy enables robust identification of PDEs in noisy conditions.
This paper focuses on density-based clustering, particularly the Density Peak (DP) algorithm and the one based on density-connectivity DBSCAN; and proposes a new method which takes advantage of the individual strengths of these two methods to yield a density-based hierarchical clustering algorithm. Our investigation be…
Study confirms weak-form market efficiency using machine learning on US stock data.
problem Validating weak-form market efficiency using machine learning.
method Conducted econometric tests and implemented algorithmic trading with five machine learning algorithms.
result No predictive power found in any machine learning model, reinforcing weak-form market efficiency.
New ensemble models classify mouse movement trajectories to assess survey question difficulty.
problem Assessing survey question difficulty based on respondents' interaction data.
method Ensemble models combining semi-metric-based weak learners to classify multivariate functional data.
result Improved survey data quality through better identification of respondent difficulty.
IdBench benchmarks semantic representations of identifiers, revealing strengths and weaknesses.
problem Evaluating semantic representations of identifiers in source code.
method Created a benchmark using developer ratings, evaluated natural language and source code embeddings, and compared lexical string distance functions.
result No single technique provides a satisfactory representation of semantic similarities, but ensemble models can improve performance.
Boosting improves accuracy with fewer calls to weak learners for certain concept classes.
problem Improving accuracy of learning algorithms with limited weak learner calls.
method Combines boosting and list-decodable codes to achieve better performance for specific concept classes.
result A new boosting algorithm that achieves strong learning with fewer calls to weak learners and additional samples.
New results on inferring hidden states in trackable weak models.
problem Inferring hidden states in trackable weak models.
method Analyzing strongly-connected trackable weak models and reconstructing branch choices.
result The number of hypotheses in strongly-connected trackable models is bounded by a constant.
Unified framework for best arm identification and dueling bandits regret minimization.
problem Best arm identification and dueling bandits regret minimization.
method Tree-Guided Identify-Then-Exploit (TG-ITE) framework.
result Unified approach achieving optimal sample complexity and regret guarantees.
Improves understanding of PWS by calculating influence of sources and data.
problem Understanding the influence of each component in PWS.
method Proposes source-aware Influence Function (IF) to decompose and calculate influence.
result Improves end model's generalization performance and identifies mislabeling.
Expands weak supervision by allowing partial labels from multiple noisy sources.
problem Creating models without labeled data using heuristic labelers.
method Probabilistic generative model estimating partial label accuracies.
result Improved model accuracy on various tasks (8.6% on text, comparable to zero-shot methods on images).
R2P method identifies homogeneous and heterogeneous subgroups for better treatment effect estimation.
problem Current subgroup analysis methods are weak in identifying homogeneous and heterogeneous subgroups and lack confidence estimates.
method R2P uses an arbitrary ITE estimator and quantifies uncertainty robustly.
result R2P produces more homogeneous and heterogeneous partitions than other methods.
Improved machine learning models outperform their simpler counterparts by using imperfect labels.
problem Improving model performance using imperfect labels.
method Random feature ridge regression (RFRR) with a deterministic equivalent for excess test error.
result The student model can outperform the teacher model regardless of the teacher's scaling law, achieving the minimax optimal rate.
New method identifies stable latent variables across different domains using weak distributional invariances.
problem Learning causal representations for multi-domain datasets.
method Autoencoders incorporating weak distributional invariances.
result Autoencoders can identify stable latent variables across different domains.
Machine learned models exhibit bias, often because the datasets used to train them are biased. This presents a serious problem for the deployment of such technology, as the resulting models might perform poorly on populations that are minorities within the training set and ultimately present higher risks to them. We pr…
We investigate the computational aspects of the basket CDS pricing with counterparty risk under a credit contagion model of multinames. This model enables us to capture the systematic volatility increases in the market triggered by a particular bankruptcy. The drawback of this problem is its analytical complication due…
We introduce a new paradigm that is important for community detection in the realm of network analysis. Networks contain a set of strong, dominant communities, which interfere with the detection of weak, natural community structure. When most of the members of the weak communities also belong to stronger communities, t…
Sharp analysis of knowledge distillation for high-dimensional regression.
problem Characterizing the risk of target models in high-dimensional settings.
method Sharp non-asymptotic bounds for ridgeless regression under model and distribution shifts.
result Identifies optimal surrogate models and reveals benefits and limitations of discarding weak features.
New method identifies latent sources from nonlinear mixtures without auxiliary variables.
problem Identifying latent sources from nonlinear mixtures without additional information.
method Structural Sparsity assumptions on the mixing process.
result Latent sources can be identified up to permutation and transformation.
Given an L2-acyclic connected finite CW-complex, we define its universal L2-torsion in terms of the chain complex of its universal covering. It takes values in the weak Whitehead group Whw(G). We study its main properties such as homotopy invariance, sum formula, product formula and Poincaré d…
Paper introduces a new identifiability criterion for DAGs using conditional variances.
problem Challenges in discovering causal relationships from observational data.
method Introduces a novel identifiability criterion for DAGs using conditional variances. Uses weak majorization on Cholesky factor of covariance matrix for learning DAGs.
result Demonstrates effectiveness of the new approach in recovering DAGs through simulations and real data analysis.
Clustering partitions a dataset such that observations placed together in a group are similar but different from those in other groups. Hierarchical and K-means clustering are two approaches but have different strengths and weaknesses. For instance, hierarchical clustering identifies groups in a tree-like structure b…
Efficiently clusters data with weak assumptions, robust to contamination.
problem General-shaped clustering under weak parametric assumptions with data contamination.
method Two-step hybrid robust clustering algorithm combining trimmed k-means and hierarchical agglomeration.
result Outperforms state-of-the-art methods in various applications.
The paper tackles performative risk optimization under weak convexity assumptions.
problem Optimizing performative risk in a closed-loop prediction system with weak convexity.
method Relaxing convexity assumptions to maintain optimization feasibility.
result Iterative optimization methods remain applicable even with weakened convexity conditions.
Modern bio-technologies have produced a vast amount of high-throughput data with the number of predictors far greater than the sample size. In order to identify more novel biomarkers and understand biological mechanisms, it is vital to detect signals weakly associated with outcomes among ultrahigh-dimensional predictor…
Paper proves large deviation principle for stochastic approximations.
problem Asymptotic estimates of learning algorithm deviations.
method Weak convergence approach to large deviations.
result Identifies appropriate scaling sequence and new representation for rate function.
Inspector Gadget uses crowdsourcing and data programming to label industrial images efficiently.
problem Securing enough labeled data for machine learning in industrial settings.
method Combines crowdsourcing, data augmentation, and data programming.
result Obtains better performance than other weak-labeling techniques.
Proves identifiability of deep latent variable models without auxiliary information.
problem Identify deep generative models without side information.
method Analyzes a broad class of deep latent variable models with universal approximation capabilities.
result Identifiability of generative models without side information u. New non-canonical flows found via parabolic Allen-Cahn equations.
problem Existence of non-canonical mean curvature flows inside fattening regions.
method Construction of non-canonical flows as limits of parabolic ε-Allen-Cahn solutions.
result First examples of non-outermost, non-canonical integral Brakke motions.