Extends curve theory to non-smooth data with finite curvature and torsion.
problem Applying classical curve theory to non-smooth data.
method Using distributional derivative measures of functions of bounded variation.
result Essentially unique non-smooth curve solution with finite total curvature and torsion.
Finite singular times for symmetric network curvature flow.
problem Formation of singularities in network curvature flow.
method Curvature flow of networks with symmetric initial data and two triple junctions.
result The set of singular times is finite.
Finite-gap solutions approximate jets of initial data for certain BKM systems.
problem Approximating jets of initial data for specific PDE systems.
method Using finite-reduction map to finite-gap solutions of Stäckel systems.
result Full jet-surjectivity for KdV and Kaup--Boussinesq, partial for Camassa--Holm.
Estimates causal effects in Gaussian Linear SCMs with finite data.
problem Estimating causal effects from observational data with latent confounders.
method Centralized Gaussian Linear SCMs (CGL-SCMs) and EM-based estimation algorithm.
result Learned CGL-SCM parameters accurately recover causal distributions from finite observational samples.
The paper develops finite knot theory using ropelength-filtered Reidemeister graphs.
problem Understanding knot types in bounded ropelength sublevel spaces.
method Study thick representatives in bounded ropelength sublevel spaces through lifted Reidemeister graphs.
result Define characteristic Reidemeister patterns and finite recognition length.
Proposes a new model for clustering with heavier tails.
problem Clustering with heavy-tailed data.
method Finite mixture of skewed sub-Gaussian stable distributions, maximum likelihood estimation, EM algorithm.
result The proposed model can robustly handle heavy-tailed data.
New guarantees for uniquely identifying transport maps and vector fields from finite measure-valued data.
problem Unique recovery of transport maps and vector fields from finite measure-valued data.
method Use of Whitney and Takens embedding theorems to establish conditions for unique identification.
result New metric for comparing diffeomorphisms and analogous results in infinitesimal settings.
The paper studies how more data affects prediction risk in high-dimensional models.
problem The impact of increasing data on prediction risk in high-dimensional models.
method Derives central limit theorem and provides finite-sample distribution and confidence interval for prediction risk.
result Demonstrates 'more data hurt' phenomenon in high-dimensional least squares estimation.
Paper extends conformal prediction to complex survey data.
problem Applying distribution-free prediction intervals to complex survey data.
method Design-based conformal prediction for non-exchangeable data.
result Empirical guarantees of finite-sample coverage for complex survey data.
New method reduces over-parametrization in neural networks, ensuring sparsity and finite network size.
problem Over-parametrization leads to too many active neurons in neural networks, especially with large data.
method Investigates a nonconvex regularization method for shallow ReLU networks.
result Locally optimal networks are finite even with infinite data, maintaining approximation guarantees and network size bounds.
We identify linear models from nonlinear systems with initialization constraints.
problem Identifying linear models from nonlinear systems with initialization constraints.
method Multiple trajectories-based deterministic data acquisition algorithm followed by regularized least squares.
result We provide a finite sample error bound on the learned linearized dynamics.
Gaussian processes (GPs) offer a flexible class of priors for nonparametric Bayesian regression, but popular GP posterior inference methods are typically prohibitively slow or lack desirable finite-data guarantees on quality. We develop an approach to scalable approximate GP regression with finite-data guarantees on th…
In this paper we prove that a complete, embedded minimal surface M in R3 with finite topology and compact boundary (possibly empty) is conformally a compact Riemann surface M with boundary punctured in a finite number of interior points and that M can be represented in terms of meromorphic …
Finite resources limit false discovery rate control in structured hypothesis spaces.
problem Controlling false discovery rate in hypothesis testing with finite data and structured hypothesis spaces.
method Framework for exact FDR control and adaptive power maximization.
result Exact FDR control and adaptive power maximization.
Adapts score matching for missing data in flexible settings.
problem Learning data distribution with missing data.
method Adapted score matching to handle missing data, providing two approaches: importance weighting and variational.
result Variational approach performs best in high-dimensional settings.
End-to-end algorithm for controlling bilinear systems with probabilistic noise.
problem Controlling bilinear systems with noisy data.
method Proposes an end-to-end algorithm using statistical learning theory and robust controller design.
result Derived finite sample identification error bounds and structurally suitable for control.
FMM fails to accurately determine the number of components even with consistent posterior.
problem Determining the number of subpopulations in a data set using FMM.
method Analysis of FMM component-count posterior under model misspecification.
result FMM component-count posterior diverges under model misspecification, contrary to intuition.
Kernelized convex clustering handles non-linear and non-convex data.
problem Lack of effective clustering methods for non-linear and non-convex data.
method Kernelized convex clustering in RKHS.
result Superior performance compared to state-of-the-art techniques.
We investigate representations of mapping class groups of surfaces that arise from the untwisted Drinfeld double of a finite group G, focusing on surfaces without marked points or with one marked point. We obtain concrete descriptions of such representations in terms of finite group data. This allows us to establish va…
Magnitude is not continuous but may be stable for most finite metric spaces.
problem Stability of magnitude invariant in finite metric spaces.
method Investigates the continuity properties of magnitude with respect to Gromov-Hausdorff topology.
result Magnitude is nowhere continuous but may be generically continuous.
We learn linear models from nonlinear systems using multiple trajectories and regularization.
problem Identifying linear models from data when the underlying dynamics are nonlinear.
method Multiple trajectories data acquisition followed by regularized least squares.
result Learn linearized dynamics with arbitrarily small error given enough samples.
FDNet learns PDEs from data with fast predictions.
problem Discovering complex systems behavior from data.
method Finite difference neural networks (FDNet) to learn PDEs from trajectory data.
result FDNet predicts future behavior with few trainable parameters.
Graph Laplacians and machine learning predict properties of finite graphs.
problem Understanding properties of finite graphs using spectral and topological methods.
method Combining graph Laplacians, spectral inequalities, machine learning, and topological data analysis.
result Neural networks can accurately predict graph properties like Ricci-flatness and spectral gaps.
Framework detects and mitigates data-poisoning attacks in causal effect estimation.
problem Vulnerability to append-only attacks in observational causal analyses.
method Develops a data-poisoning audit for augmented inverse-probability-weighted estimation.
result Proposes a greedy scan to compute exact worst-case movement at every append budget.
Stochastic optimization algorithms with variance reduction have proven successful for minimizing large finite sums of functions. Unfortunately, these techniques are unable to deal with stochastic perturbations of input data, induced for example by data augmentation. In such cases, the objective is no longer a finite su…
The paper sets sample complexity bounds for identifying LTI systems from a finite set.
problem Identifying an LTI system from a finite set of possible systems using trajectory data.
method Maximum likelihood estimator and information theory tools.
result Upper and lower bounds for sample complexity are derived, independent of stability assumption.
Paper provides unbiased spectral moment estimates from finite data.
problem Challenges in estimating spectral moments from limited data.
method Dynamic programming approach to estimate spectral moments of kernel integral operator.
result Demonstrates consistency with theoretical spectra and practical utility in neural networks.
We quantify Peter Scott's Theorem that surface groups are locally extended residually finite (LERF) in terms of geometric data. In the process, we will quantify another result by Scott that any closed geodesic in a surface lifts to an embedded loop in a finite cover.
Investigates the impact of finite VC dimension on neural network approximation and learning.
problem The influence of VC dimension on neural network approximation and learning from samples.
method Analysis of high-dimensional geometry and statistical learning theory, focusing on VC dimension.
result Finite VC dimension is beneficial for uniform convergence of empirical errors but not for approximation of functions from a probability distribution.
A neural network model predicts the critical point of the Ising phase transition.
problem Predicting the critical point of the Ising phase transition using supervised learning.
method Proposed a minimal one-free-parameter neural network model to describe the supervised learning problem for the Ising model.
result Just one free parameter is enough to describe the universal finite-size-scaling function in the network output.
The paper analyzes how the one-dimensional Wasserstein distance captures pointwise density differences in finite samples.
problem Uncertainty in identifying density differences when supports overlap and densities have substantial pointwise differences.
method Analysis using the Poisson process and neural spike train decoding.
result The one-dimensional Wasserstein distance highlights meaningful density differences related to both rate and support.
New algorithm improves causal discovery in biomedical data.
problem Stability and accuracy issues in causal discovery algorithms.
method Exploits temporal structure and tiered background knowledge.
result Increases accuracy in finite samples for causal structure estimation.
Finite mixture model is an important branch of clustering methods and can be applied on data sets with mixed types of variables. However, challenges exist in its applications. First, it typically relies on the EM algorithm which could be sensitive to the choice of initial values. Second, biomarkers subject to limits of…
Deep learning method improves regression accuracy.
problem Nonparametric regression challenges.
method Over-parametrized deep neural networks with logistic activation, gradient descent, special topology, random initialization, and data-dependent learning rate.
result Theoretical bound on L2 error and improved finite sample performance. Estimates barycenter in geodesic spaces with finite sample bounds.
problem Estimating the barycenter of a distribution in geodesic spaces.
method Finite sample error bounds, Hoeffding- and Bernstein-type concentration inequalities, efficient algorithms.
result Statistical guarantees for efficient barycenter computation.
We propose a kernel method to identify finite mixtures of nonparametric product distributions. It is based on a Hilbert space embedding of the joint distribution. The rank of the constructed tensor is equal to the number of mixture components. We present an algorithm to recover the components by partitioning the data p…
Spatially constrained Gaussian mixture models reduce covariance complexity.
problem High dimensionality in finite mixture models for spatial data.
method Spatial covariance constraint with only four free parameters.
result Improves clustering of multi-way spatial data and inference of spatial patterns.
Extends SGM to functional spaces for multimodal data.
problem Modeling densities in functional spaces.
method Represent data in spectral space, dissociate stochastic and space-time components, use SGM for sampling.
result Demonstrates effectiveness on multimodal datasets.
Study on Dirichlet process mixtures for clustering consistency.
problem Consistency of clustering with Dirichlet process mixtures.
method Analysis of posterior distribution as sample size increases, focusing on consistency for the number of clusters.
result Consistency for the number of clusters can be achieved with a properly adapted concentration parameter in a Bayesian setting.
Paper proposes a diagnostic tool for evaluating model performance out-of-sample.
problem Evaluating model performance on unseen data.
method Uses a finite calibration dataset to assess future losses.
result Provides guarantees under weak assumptions and quantifies distribution shifts.
We study minimal annuli in S2×R of finite type by relating them to harmonic maps C→S2 of finite type. We rephrase an iteration by Pinkall-Sterling in terms of polynomial Killing fields. We discuss spectral curves, spectral data and the geometry of the isospectral set…
Paper introduces FNM framework for learning finite-dimensional parametrized models.
problem Efficiently learning finite-dimensional parametrized models from limited data.
method Fourier Neural Mappings (FNMs) framework for operator learning.
result End-to-end learning of PtO maps can be less data-efficient than learning the solution operator first.
EbC learns equivariant embeddings from unlabeled group actions.
problem Learning equivariant embeddings from unlabeled group actions.
method Equivariance by Contrast (EbC) method to learn equivariant embeddings from observation pairs (y,g⋅y). result High-fidelity equivariance in latent space for diverse groups.
Privacy concerns have led to the development of privacy-preserving approaches for learning models from sensitive data. Yet, in practice, even models learned with privacy guarantees can inadvertently memorize unique training examples or leak sensitive features. To identify such privacy violations, existing model auditin…
We study the sample complexity of private synthetic data generation over an unbounded sized class of statistical queries, and show that any class that is privately proper PAC learnable admits a private synthetic data generator (perhaps non-efficient). Previous work on synthetic data generators focused on the case that …
Study shows DQN's performance degrades with temporal dependence in data.
problem Temporal dependence in replayed data affects DQN's performance.
method Modelled τ-mixing data, derived risk bounds, and empirical validation. result Temporal dependence leads to a degradation in DQN's performance rate.
We demonstrate how a 3-manifold, a Heegaard diagram, and a group presentation can each be interpreted as a pair of signed permutations in the symmetric group Sd. We demonstrate the power of permutation data in programming and discuss an algorithm we have developed that takes the permutation data as input and determi…
New algorithms improve spectral clustering for finite mixture models.
problem Issues with EM algorithm in spectral clustering.
method Spectral decomposition and non-parametric bootstrap sampling.
result Improved convergence and avoidance of poor solutions.