Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

3176359521,269 · Jun 202019922001200920172026
48 results for data minimization

In this work, we study data preconditioning, a well-known and long-existing technique, for boosting the convergence of first-order methods for regularized loss minimization. It is well understood that the condition number of the problem, i.e., the ratio of the Lipschitz constant to the strong convexity modulus, has a h…

2014-08-13abs ↗pdf ↗

New method estimates robust mean in high dimensions with minimized outliers.

problem Estimating the mean in high dimensions when a fraction of data is corrupted.
method Formulating the problem as 0\ell_0-norm minimization under second moment constraints, and using 1\ell_1 and p\ell_p minimization techniques.
result The proposed method achieves order optimal robust mean estimation and significantly outperforms existing methods.

We study the Dirichlet problem for minimal surface systems in arbitrary dimension and codimension via mean curvature flow, and obtain the existence of minimal graphs over arbitrary mean convex bounded C2C^2 domains for a large class of prescribed boundary data. This result can be seen as a natural generalization of the…

2017-01-06abs ↗pdf ↗

We construct geometric barriers for minimal graphs in H^n xR. We prove the existence and uniqueness of a solution of the vertical minimal equation in the interior of a convex polyhedron in H^n extending continuously to the interior of each face, taking infinite boundary data on one face and zero boundary value data on …

2009-08-28abs ↗pdf ↗

We assume i.i.d. data sampled from a mixture distribution with K components along fixed d-dimensional linear subspaces and an additional outlier component. For p>0, we study the simultaneous recovery of the K fixed subspaces by minimizing the l_p-averaged distances of the sampled data points from any K subspaces. Under…

2011-04-19abs ↗pdf ↗

Minimal existence time for Willmore flow established for smooth and weak Lipschitz initial data.

problem Existence time of the Willmore flow for various initial conditions.
method Established minimal existence time for Willmore flow using geometric data and conservation laws.
result Minimal existence time is a function of geometric data for general weak Lipschitz initial data.

Deep learning for HJB PDEs using synthetic data and residual minimization.

problem Solving Hamilton-Jacobi-Bellman PDEs for optimal control problems.
method Gradient-augmented synthetic dataset for supervised learning, residual minimization.
result Improves accuracy and efficiency of deep learning for HJB PDEs.

Constructs perturbations of a minimal surface with triple junctions.

problem Minimal surfaces with triple junctions in curved spaces.
method Constructs stationary perturbations with given boundary conditions.
result Constructs minimal surfaces with triple junctions in R2imesS1\mathbb{R}^2 imes \mathbb{S}^1.

Analyzes how learning algorithms affect and are affected by data manipulation.

problem Characterizing the closed-loop behavior of learning algorithms in the presence of decision-dependent data.
method Analyzes repeated risk minimization as perturbed gradient flows of performative risk minimization, considering multiple local minimizers.
result Characterizes the region of attraction for various equilibria and introduces performative alignment.

Paper tackles heavy-tailed data without finite variance, proposing robust risk minimization.

problem Empirical risk minimization under heavy-tailed data with finite pp-th moment.
method Minimizes risk values robustly estimated via Catoni's method, using generalized generic chaining.
result Shows better performance of optimizer based on empirical risks via Catoni-style estimation.

In this paper we investigate H-minimal graphs of lower regularity. We show that noncharactersitic C^1 H-minimal graphs whose components of the unit horizontal Gauss map are in W^{1,1} are ruled surfaces with C^2 seed curves. In a different direction, we investigate ways in which patches of C^1 H-minimal graphs can be g…

2005-05-13abs ↗pdf ↗

Paper extends SMM to weakly convex and multi-convex surrogates for non-convex optimization.

problem Non-convex optimization with weakly convex or multi-convex surrogates.
method Stochastic majorization-minimization with proximal regularization or block-minimization.
result Convergence rates for empirical and expected losses under non-i.i.d. data.

Study determines a minimal surface in a Riemannian manifold from boundary data.

problem Determining a minimal surface in a Riemannian manifold from boundary data.
method Analyzes the Dirichlet-to-Neumann map for the minimal surface equation.
result Knowledge of the Dirichlet-to-Neumann map determines the Riemannian manifold up to isometry.

We study the minimal surface equation in the Heisenberg space, Nil_3. A geometric proof of non existence of minimal graphs over non convex, bounded and unbounded domains is achieved (our proof holds in the Euclidean space as well). We solve the Dirichlet problem for the minimal surface equation over bounded and unbound…

2015-08-07abs ↗pdf ↗

Proposes a new method for kernel density estimation using stagewise minimization and a simple dictionary.

problem Kernel density estimation with data-adaptive weighting parameters and sparse representation.
method Stagewise minimization algorithm based on UU-divergence and a simple dictionary.
result Develops non-asymptotic error bound for the proposed estimator.

Constructs a family of genus three minimal surfaces with parallel ends.

problem Creating embedded doubly periodic minimal surfaces with specific topological properties.
method Constructs a one-parameter family of surfaces with given Weierstrass data and solves the period problem.
result Solves the two dimensional period problem for the constructed surfaces.

Unhinged loss minimization fails to improve classifier accuracy for simple data.

problem Accuracy of classifiers minimizing the unhinged loss.
method Minimizing the unhinged loss function.
result Minimizing the unhinged loss yields classifiers with accuracy no better than random guessing for simple data.

Discrete approximation solves Björling's minimal surface problem.

problem Constructing minimal surfaces from real-analytic curves with specified normal fields.
method Approximate solution by discrete minimal surfaces and discrete isothermic surfaces.
result Approximation error is proportional to the square of the mesh size.

Bayesian optimization (BO) aims to minimize a given blackbox function using a model that is updated whenever new evidence about the function becomes available. Here, we address the problem of BO under partially right-censored response data, where in some evaluations we only obtain a lower bound on the function value. T…

2013-10-07abs ↗pdf ↗

We suggest a new definition for discrete minimal surfaces in terms of sphere packings with orthogonally intersecting circles. These discrete minimal surfaces can be constructed from Schramm's circle patterns. We present a variational principle which allows us to construct discrete analogues of some classical minimal su…

2003-05-13abs ↗pdf ↗

Study of Dirichlet minimizers on manifolds with boundary and their asymptotic behavior.

problem Understanding the behavior of solutions to the Allen-Cahn equation on manifolds with boundary.
method Analyzing the asymptotic behavior of Dirichlet minimizers, relating Neumann data to boundary geometry, and using invertibility of the linearized Allen-Cahn operator.
result Computed expansions of the solution to high order and established a projection theorem about Allen-Cahn solutions near minimal surfaces.

New guarantees for ERM with adaptively collected data.

problem Failure of ERM guarantees with adaptively collected data.
method Importance sampling weighted ERM algorithm with maximal inequality.
result First generalization guarantees and fast convergence rates for adaptively collected data.

New framework minimizes model complexity for improved few-shot learning.

problem Empirical benefits of pre-training scale with data size but lack theoretical explanation.
method Complexity Minimization framework for meta-representation learning.
result Theoretical analysis shows error rate improves with more meta-training data.

We consider a general statistical learning problem where an unknown fraction of the training data is corrupted. We develop a robust learning method that only requires specifying an upper bound on the corrupted data fraction. The method minimizes a risk function defined by a non-parametric distribution with unknown prob…

2019-10-03abs ↗pdf ↗

We introduce Invariant Risk Minimization (IRM), a learning paradigm to estimate invariant correlations across multiple training distributions. To achieve this goal, IRM learns a data representation such that the optimal classifier, on top of that data representation, matches for all training distributions. Through theo…

2019-07-05abs ↗pdf ↗

We study the problem of finding the best linear model that can minimize least-squares loss given a data-set. While this problem is trivial in the low dimensional regime, it becomes more interesting in high dimensions where the population minimizer is assumed to lie on a manifold such as sparse vectors. We propose proje…

2019-07-03abs ↗pdf ↗

We study conditional risk minimization (CRM), i.e. the problem of learning a hypothesis of minimal risk for prediction at the next step of sequentially arriving dependent data. Despite it being a fundamental problem, successful learning in the CRM sense has so far only been demonstrated using theoretical algorithms tha…

2018-01-01abs ↗pdf ↗

Develops new methods to evaluate data influence in SAM for improved model training.

problem Challenges in mislabeled noisy data and privacy concerns in SAM.
method Two innovative data valuation methods based on influence functions (IF) for SAM.
result Demonstrates effectiveness in identifying mislabeled data and enhancing interpretability.

A new DP algorithm for weighted ERM protects sensitive data in predictive models.

problem Protecting sensitive personal information in predictive models trained via ERM.
method Proposes the first differentially private algorithm for weighted ERM with formal privacy guarantees.
result Demonstrates strong DP guarantees while maintaining robust performance in real-world data.

The Bartnik mass is a notion of quasi-local mass which is remarkably difficult to compute. Mantoulidis and Schoen [2016] developed a novel technique to construct asymptotically flat extensions of minimal Bartnik data in such a way that the ADM mass of these extensions is well-controlled, and thus, they were able to com…

2019-03-21abs ↗pdf ↗