NT probability measures knotting in 3D arc systems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Non-trivialization probability of arc system in 3D space
Survey on negative transfer in machine learning.
A Dynamic Chain Event Graph (DCEG) provides a rich tree-based framework for modelling a dynamic process with highly asymmetric developments. An N Time-Slice DCEG (NT-DCEG) is a useful subclass of the DCEG class that exhibits a specific type of periodicity in its supporting tree graph and embodies a time-homogeneity ass…
NTS-NOTEARS learns DBNs from time-series data with prior knowledge.
The Dynamic Chain Event Graph (DCEG) is able to depict many classes of discrete random processes exhibiting asymmetries in their developments and context-specific conditional probabilities structures. However, paradoxically, this very generality has so far frustrated its wide application. So in this paper we develop an…
This paper presents a novel clustering concept that is based on jointly learned nonlinear transforms (NTs) with priors on the information loss and the discrimination. We introduce a clustering principle that is based on evaluation of a parametric min-max measure for the discriminative prior. The decomposition of the pr…
Uniform bounds for neural networks' generalization error in overparameterized settings.
Having a sequence-to-sequence model which can operate in an online fashion is important for streaming applications such as Voice Search. Neural transducer is a streaming sequence-to-sequence model, but has shown a significant degradation in performance compared to non-streaming models such as Listen, Attend and Spell (…
New algorithms achieve optimal regret in sliding window model with limited memory.
Kernel methods form a theoretically-grounded, powerful and versatile framework to solve nonlinear problems in signal processing and machine learning. The standard approach relies on the \emph{kernel trick} to perform pairwise evaluations of a kernel function, leading to scalability issues for large datasets due to its …
Neural networks can interpolate random data but still generalize well, studied in the NT regime.
New analysis of signSGD with random reshuffling shows faster convergence rates.
Paper discusses new stochastic algorithms for sparse signal recovery.
New algorithms reduce private bandit regret to nearly non-private levels.
This paper examines the impact of random initialization in neural networks using NTK theory.
Optimizes profit in targeted marketing across multiple markets with varying marketing expenditures.
Unified toolkit for comparing neural representations using SRTD and NTS.
In this paper, we study the algebraic properties of the higher analogues of Courant algebroid structures on the direct sum bundle for an -dimensional manifold. As an application, we revisit Nambu-Poisson structures and multisymplectic structures. We prove that the graph of an -vector fi…
Proportional transaction costs present difficult theoretical problems in trading algorithm design, on account of their lack of analytical tractability. The author derives a solution of DT-NT-DT form for an arbitrary model in which the the traded asset has diffusive dynamics described by one or more stochastic risk fact…
In this paper, we show that the spaces of sections of the -th differential operator bundle $\dev^n E$ and the -th skew-symmetric jet bundle $\jet_n E$ of a vector bundle are isomorphic to the spaces of linear -vector fields and linear -forms on respectively. Consequently, the -omni-Lie algebroi…
We consider assets for which price and squared volatility are jointly driven by Heston joint stochastic differential equations (SDEs). When the parameters of these SDEs are estimated from sub-sampled data , estimation errors do impact the classical option pricing PDEs. We estimate thes…
In this paper, we propose and analyze SPARQ-SGD, which is an event-triggered and compressed algorithm for decentralized training of large-scale machine learning models. Each node can locally compute a condition (event) which triggers a communication where quantized and sparsified local model parameters are sent. In SPA…
Knots parametrized in cylinder coordinates by t -> (st, 3 + cos(nt), cos(mt + φ)) share properties of Lissajous and billiard knots in a cylinder. We use these 'billiard knots in a flat solid torus' to study two topics: when is Z(s,n,m) equal to Z(s,m,n)? And: why are the determinants of certain Lissajous and billiard k…
We propose maximum likelihood estimation for learning Gaussian graphical models with a Gaussian (ell_2^2) prior on the parameters. This is in contrast to the commonly used Laplace (ell_1) prior for encouraging sparseness. We show that our optimization problem leads to a Riccati matrix equation, which has a closed form …
We study the supervised learning problem under either of the following two models: (1) Feature vectors are -dimensional Gaussians and responses are for an unknown quadratic function; (2) Feature vectors are distributed as a mixture of two $…
In this paper, we propose the first computationally efficient projection-free algorithm for bandit convex optimization (BCO). We show that our algorithm achieves a sublinear regret of (where is the horizon and is the dimension) for any bounded convex functions with uniformly bounded gradients. We …
New analysis shows Local SGD can achieve error scaling with only fixed number of communications.
While training a machine learning model using multiple workers, each of which collects data from their own data sources, it would be most useful when the data collected from different workers can be {\em unique} and {\em different}. Ironically, recent analysis of decentralized parallel stochastic gradient descent (D-PS…
We consider decentralized stochastic optimization with the objective function (e.g. data samples for machine learning task) being distributed over machines that can only communicate to their neighbors on a fixed communication graph. To reduce the communication bottleneck, the nodes compress (e.g. quantize or sparsi…
HD algorithm simulates dynamics on random matrix ensembles without generating full matrices.
Empirical study compares wide neural networks to kernel methods, resolving open questions.
We continue the study of the spectral theory associated to integrable metrics, started in our previous paper arXiv:1301.1793 [math.SP]. We introduce the notion of 1-integrable metric on line-bundles on a compact Riemann surface. We extend the spectral theory of generalized Laplacians to line-bundles equipped with 1-int…
In the present paper, we study globally framed f-manifolds in the particular setting of indefinite S-manifolds for both spacelike and timelike cases. We prove that if is a warped CR-submanifold such that is ?-anti-invariant and NT is ?-invariant, then M is a CR-product. We…
The goal of this paper is to offer a comprehensive exposition of the current knowledge about Heegaard splittings of exteriors of knots in the 3-sphere. The exposition is done with a historical perspective as to how ideas developed and by whom. Several new notions are introduced and some facts about them are proved. In …
New optimal rates for score estimation improve diffusion model performance.
A decentralized algorithm minimizes cumulative regret in stochastic linear bandits with safety constraints.
The systole length of hyperbolic n-manifolds is bounded by a function of n and t.
New method improves convergence of SPP for convex optimization problems.
Agents collaborate to minimize regret while keeping costs under a threshold.
In this study we suggest a portfolio selection framework based on option-implied information and multivariate non-Gaussian models. The proposed models incorporate skewness, kurtosis and more complex dependence structures among stocks log-returns than the simple correlation matrix. The two models considered are a multiv…
Decentralized training of deep learning models is a key element for enabling data privacy and on-device learning over networks, as well as for efficient scaling to large compute clusters. As current approaches suffer from limited bandwidth of the network, we propose the use of communication compression in the decentral…
Thompson Sampling is one of the oldest heuristics for multi-armed bandit problems. It is a randomized algorithm based on Bayesian ideas, and has recently generated significant interest after several studies demonstrated it to have better empirical performance compared to the state of the art methods. In this paper, we …
Study MNL-Bandit in non-stationary settings with optimal regret bound.
Thompson Sampling remains differentially private with minimal modifications.
Text-based analysis methods allow to reveal privacy relevant author attributes such as gender, age and identify of the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called Adver…
We analyze the convergence of gradient-based optimization algorithms that base their updates on delayed stochastic gradient information. The main application of our results is to the development of gradient-based distributed optimization algorithms where a master node performs parameter updates while worker nodes compu…
A new online learning setting for autoregressive processes with sublinear regret.