New method defends against neural backdoors using generative modeling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The straight-line flow on almost every staircase and on almost every square tiled staircase is recurrent. For almost every square tiled staircase the set of periodic orbits is dense in the phase space.
Maximum entropy distributions with discrete support in dimensions arise in machine learning, statistics, information theory, and theoretical computer science. While structural and computational properties of max-entropy distributions have been extensively studied, basic questions such as: Do max-entropy distributio…
One of the core problems in variational inference is a choice of approximate posterior distribution. It is crucial to trade-off between efficient inference with simple families as mean-field models and accuracy of inference. We propose a variant of a greedy approximation of the posterior distribution with tractable bas…
Optimal DP mechanisms for vector queries are found to be staircase distributions.
The staircase property aids deep learning by guiding hierarchical feature learning.
Study Markov staircases in symplectic embeddings of rational homology ellipsoids.
New algorithms speed up inverse reinforcement learning by solving MDPs once.
Pointwise localization allows more precise localization and accurate interpretability, compared to bounding box, in applications where objects are highly unstructured such as in medical domain. In this work, we focus on weakly supervised localization (WSL) where a model is trained to classify an image and localize regi…
Efficient algorithms find optimal monotone transforms for calibration under strictly convex losses.
The paper analyzes RLHF with human feedback and provides convergence results for MLE and pessimistic MLE.
Neural networks outperform kernels by learning features better.
Study on connection points on double regular polygons, providing coordinates and proving non-connection points.
In this paper, we present a novel and general framework called {\it Maximum Entropy Discrimination Markov Networks} (MaxEnDNet), which integrates the max-margin structured learning and Bayesian-style estimation and combines and extends their merits. Major innovations of this model include: 1) It generalizes the extant …
AGF explains feature learning in neural networks through alternating steps.
What are the possible shapes of various things and why? For instance, when a closed wire or a frame is dipped into a soap solution and is raised up from the solution, the surface spanning the wire is a soap film. What are the possible shapes of soap films and why? Or, for instance, why is DNA like a double spiral stair…
New property helps SGD learn sparse functions efficiently in neural networks.
EBIL simplifies IL by estimating expert energy as reward, achieving effective performance.
Language recognition system is typically trained directly to optimize classification error on the target language labels, without using the external, or meta-information in the estimation of the model parameters. However labels are not independent of each other, there is a dependency enforced by, for example, the langu…
The paper uses Seshadri constants to construct symplectic ellipsoid embeddings.
The rising volume of datasets has made training machine learning (ML) models a major computational cost in the enterprise. Given the iterative nature of model and parameter tuning, many analysts use a small sample of their entire data during their initial stage of analysis to make quick decisions (e.g., what features o…
We define On-Average KL-Privacy and present its properties and connections to differential privacy, generalization and information-theoretic quantities including max-information and mutual information. The new definition significantly weakens differential privacy, while preserving its minimalistic design features such …
Minimal surfaces with uniform curvature (or area) bounds have been well understood and the regularity theory is complete, yet essentially nothing was known without such bounds. We discuss here the theory of embedded (i.e., without self-intersections) minimal surfaces in Euclidean 3-space without a priori bounds. The st…
Greedy training of recursive partitioning estimators faces a computational barrier when the true function doesn't satisfy a specific property.
SGD learns neural networks with a complexity measure called leap.
Max entropy exploration guides reinforcement learning agents to pursue achievable goals.
Study reveals efficient recovery of multi-modal signals via Bayesian methods and sequential learning.
This work connects Cramér distance to QR-DQN for DRL.
A Euclidean minimal torus with planar ends gives rise to an immersed Willmore torus in the conformal 3--sphere . The class of Willmore tori obtained this way is given a spectral theoretic characterization as the class of Willmore tori with reducible spectral curve. A spectral curve of this type…
For a Legendrian knot L in R^3 with a chosen Morse complex sequence (MCS) we construct a differential graded algebra (DGA) whose differential counts "chord paths" in the front projection of L. The definition of the DGA is motivated by considering Morse-theoretic data from generating families. In particular, when the MC…
Two-layer neural networks learn features through a few gradient descent steps, improving approximation capacity.
Let X be a proper CAT(0) cube complex admitting a proper cocompact action by a group G. We give three conditions on the action, any one of which ensures that X has a factor system in the sense of [BHS14]. We also prove that one of these conditions is necessary. This combines with results of Behrstock--Hagen--Sisto to s…
A homothety surface can be assembled from polygons by identifying their edges in pairs via homotheties, which are compositions of translation and scaling. We consider linear trajectories on a 1-parameter family of genus-2 homothety surfaces. The closure of a trajectory on each of these surfaces always has Hausdorff dim…
Characterizes a subset of links using quasipositive and homogeneous properties.
In this work, we consider the use of model-driven deep learning techniques for massive multiple-input multiple-output (MIMO) detection. Compared with conventional MIMO systems, massive MIMO promises improved spectral efficiency, coverage and range. Unfortunately, these benefits are coming at the cost of significantly i…
We investigate the random dynamics of rational maps on the Riemann sphere and the dynamics of semigroups of rational maps on the Riemann sphere. We show that regarding random complex dynamics of polynomials, in most cases, the chaos of the averaged system disappears, due to the cooperation of the generators. We investi…
Attention to entropic communication improves message decoding and cooperation.
Paper introduces RPWithPrior for efficient label differential privacy in regression.
This paper develops new tools for understanding surfaces with more than one end (and usually, of infinite topology) which properly minimally embed into Euclidean three-space. On such a surface, the set of ends forms a compact Hausdorff space, naturally ordered by the relative heights of the ends in space. One of our ma…
We develop a mean-field theory for multi-component ICA in high dimensions.
Two-layer networks learn faster with batch reuse, overcoming information and leap exponents.
We investigate the dynamics of -generator semigroups of polynomials with bounded planar postcritical set and associated random dynamics on the Riemann sphere. Also, we investigate the space of such semigroups. We show that for a parameter in the intersection of , the hyperbolicity locus ${\c…
This paper develops a cohomological hierarchy for bistable visual paradoxes.
This paper is the fifth and final in a series on embedded minimal surfaces. Following our earlier papers on disks, we prove here two main structure theorems for non-simply connected embedded minimal surfaces of any given fixed genus. The first of these asserts that any such surface without small necks can be obtained b…
Data repetition improves SGD's learning of high-dimensional functions.
We investigate the random dynamics of polynomial maps on the Riemann sphere and the dynamics of semigroups of polynomial maps on the Riemann sphere. In particular, the dynamics of a semigroup of polynomials whose planar postcritical set is bounded and the associated random dynamics are studied. In general, the Juli…
Optimization of neural networks scales with γ, revealing unique loss curves and optimal learning rates.
Study predicts lens performance using neural networks.