Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3587161,0741,432 · Jun 202019922001200920172026
48 results for smaller model

This work proposes splitting deep neural networks into smaller sub-networks for faster and more efficient distillation.

problem Challenges in training deep neural networks, including local optima, gradient issues, and computational demands.
method Proposes a non-end-to-end distillation approach by splitting networks into smaller, independent sub-networks (neighbourhoods).
result Independent training of smaller sub-networks can speed up distillation and improve efficiency in various applications.

Bayesian networks are now being used in enormous fields, for example, diagnosis of a system, data mining, clustering and so on. In spite of their wide range of applications, the statistical properties have not yet been clarified, because the models are nonidentifiable and non-regular. In a Bayesian network, the set of …

2012-10-19abs ↗pdf ↗

We propose a dynamic edge exchangeable network model that can capture sparse connections observed in real temporal networks, in contrast to existing models which are dense. The model achieved superior link prediction accuracy on multiple data sets when compared to a dynamic variant of the blockmodel, and is able to ext…

2017-10-11abs ↗pdf ↗

Smaller actor-critic models lead to performance degradation and overfitting, highlighting the critic's role in value underestimation.

problem Performance degradation and overfitting in actor-critic models with smaller actors.
method Broad empirical investigations and analyses of asymmetric actor-critic setups, exploring techniques to mitigate value underestimation.
result Value underestimation is a key cause of performance degradation in smaller actor-critic models, and the critic plays a crucial role in mitigating this.

Ensembling smaller models can outperform larger models in terms of accuracy and efficiency.

problem The inefficiency of training larger models for boosting performance.
method Training ensembles of smaller models and comparing their performance to single larger models.
result Ensembles of smaller models outperform single larger models in accuracy and efficiency, especially as models grow large.

PixelHop++ improves image classification with a smaller model size.

problem Improving image classification models with smaller sizes.
method Decomposing input tensor, channel-wise Saab transform, successive subspace learning, feature ranking.
result PixelHop++ offers a flexible tradeoff between model size and performance.

We study the least squares regression problem \begin{align*} \min_{Θ\in \mathcal{S}_{\odot D,R}} \|AΘ-b\|_2, \end{align*} where SD,R\mathcal{S}_{\odot D,R} is the set of ΘΘ for which Θ=r=1Rθ1(r)θD(r)Θ= \sum_{r=1}^{R} θ_1^{(r)} \circ \cdots \circ θ_D^{(r)} for vectors θd(r)Rpdθ_d^{(r)} \in \mathbb{R}^{p_d} for all r[R]r \in [R] and $d \in [D]…

2017-09-20abs ↗pdf ↗

We analyze coresets for regularized regression problems and propose a modified lasso that yields smaller coresets.

problem Analyzing coresets for regularized regression problems.
method Examined coresets for ridge regression and proposed a modified lasso problem.
result No coreset for regularized regression can be smaller than the unregularized version when reqsr eq s.

Adjoined Networks trains both base and compressed networks together for efficient model compression.

problem Efficiently compressing deep neural networks while maintaining accuracy.
method Adjoined Networks (AN) trains both a base network and a smaller compressed network simultaneously, sharing parameters.
result AN achieves 71.8% top-1 accuracy with 1.8M parameters and 1.6 GFLOPs on ImageNet.

We look at a collection of conjectures with the unifying message that smaller social systems, tend to be less complex and can be aligned better, towards fulfilling their intended objectives. We touch upon a framework, referred to as the four pronged approach that can aid the analysis of social systems. The four prongs …

2016-03-19abs ↗pdf ↗

New method improves quality and efficiency of generative models by using smaller diffusion times.

problem Lack of theoretical understanding of diffusion time T in score-based diffusion models.
method Introduce an auxiliary model to bridge the gap between ideal and simulated dynamics, followed by reverse diffusion.
result Empirical results show competitive performance in image data compared to state-of-the-art models.

Product Kanerva Machines dynamically combine smaller models for better memory organization.

problem Limited organization in the Kanerva Machine.
method Introducing Product Kanerva Machines that dynamically combine multiple smaller Kanerva Machines.
result Product Kanerva Machines can discover spatial tunings that approximately factorize simple images by object.

The traditional Sznajd model, as well as its Ochrombel simplification for opinion spreading, are applied to marketing with the help of advertising. The larger the lattice is the smaller is the amount of advertising needed to convince the whole market

2002-07-06abs ↗pdf ↗

We give three lower bounds for the Morse index of a constant mean curvature torus in Euclidean 3-space in terms of its spectral genus g. The first two lower bounds grow linearly in g and are stronger for smaller values of g, while the third grows quadratically in g but is weaker for smaller values of g.

2004-10-06abs ↗pdf ↗

We prove the existence of a minimal diffeomorphism isotopic to the identity between two hyperbolic cone surfaces (Σ,g1)(Σ,g_1) and (Σ,g2)(Σ,g_2) when the cone angles of g1g_1 and g2g_2 are different and smaller than ππ. When the cone angles of g1g_1 are strictly smaller than the ones of g2g_2, this minimal diffeomorphism is u…

2014-11-10abs ↗pdf ↗

Ensemble GP improves genetic programming by achieving better results with smaller models.

problem Improving genetic programming for binary classification problems.
method Ensemble GP uses an evolved population structure, fitness evaluation, and genetic operators inspired by ensemble learning methods.
result Ensemble GP outperformed standard GP on eight binary classification problems, achieving better results with smaller models.

We present a machine learning-based approach to lossy image compression which outperforms all existing codecs, while running in real-time. Our algorithm typically produces files 2.5 times smaller than JPEG and JPEG 2000, 2 times smaller than WebP, and 1.7 times smaller than BPG on datasets of generic images across all …

2017-05-16abs ↗pdf ↗

With the aid of concrete examples, we consider the question of whether, in the presence of conformal curvature, a conformal geodesic can become trapped in smaller and smaller sets, or phrased informally: are spirals possible? We do not arrive at a definitive answer, but we are able to find situations where this behavio…

2012-04-27abs ↗pdf ↗

Coresets are compact representations of data sets such that models trained on a coreset are provably competitive with models trained on the full data set. As such, they have been successfully used to scale up clustering models to massive data sets. While existing approaches generally only allow for multiplicative appro…

2017-02-27abs ↗pdf ↗

HOTCAKE compresses CNNs by decomposing kernels into smaller parts.

problem Compressing deep CNNs without significant accuracy loss.
method Input channel decomposition, guided Tucker rank selection, higher order Tucker decomposition, fine-tuning.
result HOTCAKE produces highly compressed CNN models with good accuracy.

Weight Squeezing transfers knowledge from large models to smaller ones, improving performance and speed.

problem Transfer learning and model compression for faster and more efficient training.
method Reparameterization of weights from a large model to a smaller one.
result Weight Squeezing outperforms other methods on GLUE benchmark with faster training.

We show that there exists a universal constant C>0 such that the convex hull of any N points in the hyperbolic space H^n is of volume smaller than C N, and that for any dimension n there exists a constant C_n > 0 such that for any subset A of H^n, Vol(Conv(A_1)) < C_n Vol(A_1) where A_1 is the set of points of hyperbol…

2011-05-30abs ↗pdf ↗

This paper considers the subject of information losses arising from the finite datasets used in the training of neural classifiers. It proves a relationship between such losses as the product of the expected total variation of the estimated neural model with the information about the feature space contained in the hidd…

2019-02-15abs ↗pdf ↗

Study shows different trajectory prediction models generalize better under OoD conditions.

problem Comparing trajectory prediction models' robustness across different datasets.
method Training models on Argoverse 2 and testing on Waymo Open Motion, and vice versa, with various augmentation strategies.
result Smallest model with highest inductive bias performs best in OoD generalization.

LEMON uses pre-trained models to scale neural networks efficiently.

problem Efficiency in scaling deep neural networks, especially Transformers, which are resource-intensive to train from scratch.
method LEMON initializes scaled models using pre-trained weights and optimizes learning rates.
result Significant reduction in training time and computational costs for Vision Transformers and BERT.

The paper evaluates various machine learning models for predicting industrial aging processes.

problem Accurately predicting industrial aging processes to schedule maintenance efficiently.
method Compared traditional stateless models (linear and kernel ridge regression, feed-forward neural networks) to more complex recurrent neural networks (echo state networks and LSTMs) on synthetic and real-world data.
result Recurrent models produce near perfect predictions when trained on larger datasets and maintain good performance even with domain shifts, while simpler models perform comparably on smaller datasets.

Paper proposes a method to create smaller, more efficient deep generative audio models.

problem High computation cost and complexity of deep generative models in audio applications.
method Developed a method for structured trimming of deep generative audio models.
result 95% of model weights can be removed without significant degradation in accuracy.

A smaller, less-trained model guides image generation, improving quality without sacrificing variation.

problem Improving image quality and variation in diffusion models without compromising one for the other.
method Guiding a conditional model with a smaller, less-trained version of the same model.
result Significant improvements in ImageNet generation, setting record FIDs.

In this paper, we integrate VAEs and flow-based generative models successfully and get f-VAEs. Compared with VAEs, f-VAEs generate more vivid images, solved the blurred-image problem of VAEs. Compared with flow-based models such as Glow, f-VAE is more lightweight and converges faster, achieving the same performance und…

2018-09-16abs ↗pdf ↗

We compress large neural networks for quick adaptation to specific contexts.

problem How to quickly adapt a pretrained large neural network to specific contexts.
method Propose a Bayesian hypernetwork framework to compress the network and encourage sparsity.
result Generated compressed networks are significantly smaller than baseline methods.

Fine-tuning normalization layers can reconstruct smaller networks.

problem Understanding the expressive power of fine-tuning normalization layers.
method Random ReLU networks and sparsified networks were fine-tuned to reconstruct target networks.
result Fine-tuning normalization layers can reconstruct networks that are O(extwidth)O(\sqrt{ ext{width}}) times smaller.

We introduce Independently Recurrent Long Short-term Memory cells: IndyLSTMs. These differ from regular LSTM cells in that the recurrent weights are not modeled as a full matrix, but as a diagonal matrix, i.e.\ the output and state of each LSTM cell depends on the inputs and its own output/state, as opposed to the inpu…

2019-03-19abs ↗pdf ↗

Study shows market quality improves with larger orders, not smaller tick sizes or higher trading frequencies.

problem Impact of order book tick sizes, metaorders, and trading frequencies on market quality.
method Multi-agent reinforcement learning model to simulate stock market dynamics.
result Market quality benefits from larger orders but not from smaller tick sizes or higher trading frequencies.