Paper introduces a geometric approach to model similar probability distributions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We analyse the structure of local martingale deflators projected on smaller filtrations. In a general continuous-path setting, we show that the local martingale part in the multiplicative Doob-Meyer decomposition of projected local martingale deflators are themselves local martingale deflators in the smaller informatio…
The paper explores how smaller data sets can lead to better model selection decisions.
This work proposes splitting deep neural networks into smaller sub-networks for faster and more efficient distillation.
Bayesian networks are now being used in enormous fields, for example, diagnosis of a system, data mining, clustering and so on. In spite of their wide range of applications, the statistical properties have not yet been clarified, because the models are nonidentifiable and non-regular. In a Bayesian network, the set of …
We consider discrete graphical models Markov with respect to a graph and propose two distributed marginal methods to estimate the maximum likelihood estimate of the canonical parameter of the model. Both methods are based on a relaxation of the marginal likelihood obtained by considering the density of the variable…
We propose a dynamic edge exchangeable network model that can capture sparse connections observed in real temporal networks, in contrast to existing models which are dense. The model achieved superior link prediction accuracy on multiple data sets when compared to a dynamic variant of the blockmodel, and is able to ext…
Smaller actor-critic models lead to performance degradation and overfitting, highlighting the critic's role in value underestimation.
Ensembling smaller models can outperform larger models in terms of accuracy and efficiency.
PixelHop++ improves image classification with a smaller model size.
We study the least squares regression problem \begin{align*} \min_{Θ\in \mathcal{S}_{\odot D,R}} \|AΘ-b\|_2, \end{align*} where is the set of for which for vectors for all and $d \in [D]…
We analyze coresets for regularized regression problems and propose a modified lasso that yields smaller coresets.
Adjoined Networks trains both base and compressed networks together for efficient model compression.
We look at a collection of conjectures with the unifying message that smaller social systems, tend to be less complex and can be aligned better, towards fulfilling their intended objectives. We touch upon a framework, referred to as the four pronged approach that can aid the analysis of social systems. The four prongs …
A new neural network model improves reinforcement learning efficiency.
New method improves quality and efficiency of generative models by using smaller diffusion times.
Develops a new method to model overlapping asymmetric datasets effectively.
Product Kanerva Machines dynamically combine smaller models for better memory organization.
The traditional Sznajd model, as well as its Ochrombel simplification for opinion spreading, are applied to marketing with the help of advertising. The larger the lattice is the smaller is the amount of advertising needed to convince the whole market
This work improves uncertainty estimates for LISTA estimators.
We give three lower bounds for the Morse index of a constant mean curvature torus in Euclidean 3-space in terms of its spectral genus g. The first two lower bounds grow linearly in g and are stronger for smaller values of g, while the third grows quadratically in g but is weaker for smaller values of g.
We prove the existence of a minimal diffeomorphism isotopic to the identity between two hyperbolic cone surfaces and when the cone angles of and are different and smaller than . When the cone angles of are strictly smaller than the ones of , this minimal diffeomorphism is u…
In this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of online attention-based encoder-decoder (AED) models. A three-stage training based on three levels of architectural granularity namely, character encoder, byte pair encoding (BPE) based encoder, and attention decod…
Given a classical channel---a stochastic map from inputs to outputs---the input can often be transformed to an intermediate variable that is informationally smaller than the input. The new channel accurately simulates the original but at a smaller transmission rate. Here, we examine this procedure when the intermediate…
Ensemble GP improves genetic programming by achieving better results with smaller models.
We present a machine learning-based approach to lossy image compression which outperforms all existing codecs, while running in real-time. Our algorithm typically produces files 2.5 times smaller than JPEG and JPEG 2000, 2 times smaller than WebP, and 1.7 times smaller than BPG on datasets of generic images across all …
With the aid of concrete examples, we consider the question of whether, in the presence of conformal curvature, a conformal geodesic can become trapped in smaller and smaller sets, or phrased informally: are spirals possible? We do not arrive at a definitive answer, but we are able to find situations where this behavio…
Coresets are compact representations of data sets such that models trained on a coreset are provably competitive with models trained on the full data set. As such, they have been successfully used to scale up clustering models to massive data sets. While existing approaches generally only allow for multiplicative appro…
HOTCAKE compresses CNNs by decomposing kernels into smaller parts.
Weight Squeezing transfers knowledge from large models to smaller ones, improving performance and speed.
Support vector regression (SVR) has been widely used to reduce the high computational cost of computer simulation. SVR assumes the input parameters have equal sample sizes, but unequal sample sizes are often encountered in engineering practices. To solve this issue, a new prediction approach based on SVR, namely as hig…
We show that there exists a universal constant C>0 such that the convex hull of any N points in the hyperbolic space H^n is of volume smaller than C N, and that for any dimension n there exists a constant C_n > 0 such that for any subset A of H^n, Vol(Conv(A_1)) < C_n Vol(A_1) where A_1 is the set of points of hyperbol…
Paper presents a method for geographic ratemaking using spatial embeddings.
This paper considers the subject of information losses arising from the finite datasets used in the training of neural classifiers. It proves a relationship between such losses as the product of the expected total variation of the estimated neural model with the information about the feature space contained in the hidd…
Study shows different trajectory prediction models generalize better under OoD conditions.
LEMON uses pre-trained models to scale neural networks efficiently.
The paper evaluates various machine learning models for predicting industrial aging processes.
Paper proposes a method to create smaller, more efficient deep generative audio models.
We propose a novel way to train ranking models, such as recommender systems, that are both effective and efficient. Knowledge distillation (KD) was shown to be successful in image recognition to achieve both effectiveness and efficiency. We propose a KD technique for learning to rank problems, called \emph{ranking dist…
A smaller, less-trained model guides image generation, improving quality without sacrificing variation.
In this paper, we integrate VAEs and flow-based generative models successfully and get f-VAEs. Compared with VAEs, f-VAEs generate more vivid images, solved the blurred-image problem of VAEs. Compared with flow-based models such as Glow, f-VAE is more lightweight and converges faster, achieving the same performance und…
We compress large neural networks for quick adaptation to specific contexts.
Fine-tuning normalization layers can reconstruct smaller networks.
Study finds symplectic fillings' properties for specific contact covers.
We introduce Independently Recurrent Long Short-term Memory cells: IndyLSTMs. These differ from regular LSTM cells in that the recurrent weights are not modeled as a full matrix, but as a diagonal matrix, i.e.\ the output and state of each LSTM cell depends on the inputs and its own output/state, as opposed to the inpu…
Let G=SO(n,1) and Gamma a geometrically finite Zariski dense subgroup of G which is contained in an arithmetic subgroup of G. Denoting by Gamma(q) the principal congruence subgroup of Gamma of level q, and fixing a positive number λ_0 strictly smaller than (n-1)^2/4, we show that, as q tends to infinity along primes, t…
In this paper we studied about the wavelet identification of the thresholds and time delay for more general case without the constraint that the time delay is smaller than the order of the model. Here we composed an empirical wavelet from the SETAR (Self-Exciting Threshold Autoregressive) model and identified the thres…
Study shows market quality improves with larger orders, not smaller tick sizes or higher trading frequencies.