Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

10.7%21.4%32.1%42.8% · Jun 202019922001200920172026
48 results for data accumulation

New data accumulation prevents model collapse in generative models.

problem Model collapse in generative models trained on their own outputs.
method Empirical study of language models, diffusion models, and variational autoencoders; analytically tractable framework for linear models.
result Accumulating synthetic data alongside real data avoids model collapse, preventing performance degradation.

In many large-scale machine learning applications, data are accumulated with time, and thus, an appropriate model should be able to update in an online paradigm. Moreover, as the whole data volume is unknown when constructing the model, it is desired to scan each data item only once with a storage independent with the …

2017-06-08abs ↗pdf ↗

Improves likelihood-free inference by using a new sampling approach to avoid biased data collection.

problem Efficient Bayesian inference without likelihood evaluation for real-world datasets.
method Introduces Neural Proposal (NP) to sample simulation inputs i.i.d. for unbiased posterior inference.
result Demonstrates improved performance, especially for multi-modal posteriors, through experiments.

The aim of the present article is to offer a strictly mathematical, statistical treatment of the current account balances in EU and in the Eurozone. Based on Eurostat data, an overview of the total and annual balances is first made for different collections among the EU countries. Then, using the Mathematica technical …

2013-02-19abs ↗pdf ↗

Wealth inequality is an important matter for economic theory and policy. Ongoing debates have been discussing recent rise in wealth inequality in connection with recent development of active financial markets around the world. Existing literature on wealth distribution connects the origins of wealth inequality with a v…

2018-09-23abs ↗pdf ↗

Study robustness of global feature effect explanations in machine learning models.

problem Vulnerability of global feature effect explanations to data and model perturbations.
method Theoretical bounds and experimental evaluation of partial dependence plots and accumulated local effects.
result Quantifies the gap between best and worst-case scenarios of misinterpreting machine learning predictions globally.

Study of historic stock returns distributions, highlighting asymmetry and outliers.

problem Understanding the asymmetry in accumulated gains and losses in stock returns over time.
method Analyzing decades-long historic distributions of S&P500 returns, comparing gains and losses, using statistical U-tests and fitting log-log scale linearly.
result The mean of de-trended distributions increases linearly with the number of days of accumulation, and the overall skew is negative, indicating heavier tails of losses.

Knowhow in societies accumulates as it gets transmitted from group to group, and from generation to generation. However, we lack of a unified quantitative formalism that takes into account the structured process for how this accumulation occurs, and this has precluded the development of a unified view of human developm…

2018-09-27abs ↗pdf ↗

In this paper, finite type domains with hyperbolic orbit accumulation points are studied. We prove, in case of C2\mathbb{C}^2, it has to be a (global) pseudoconvex domain, after an assumption of boundary regularity. Moreover, one of the applications will realize the classification of domains within this class, precisel…

2013-04-30abs ↗pdf ↗

Analyzes multi-day stock returns, showing linear volatility and mean dependence.

problem Linear dependence of volatility and mean in accumulated stock returns.
method Modified Jones-Faddy skew t-distribution analysis.
result Linear dependence of volatility and mean on the number of days of accumulation.

Lifelong learning can be viewed as a continuous transfer learning procedure over consecutive tasks, where learning a given task depends on accumulated knowledge --- the so-called knowledge base. Most published work on lifelong learning makes a batch processing of each task, implying that a data collection step is requi…

2018-10-26abs ↗pdf ↗

A multi-step framework tackles online unsupervised domain adaptation with novel mean-target subspace computation.

problem Online unsupervised domain adaptation with unlabelled target data arriving sequentially.
method Multi-step framework with a novel mean-target subspace computation and temporal coherency consideration.
result Improved performance over previous approaches on four datasets.

This paper presents a model of capital accumulation for a large number of heterogenous producer-consumers in an exchange space in which interactions depend on agents' positions. Each agent is described by his production, consumption, stock of capital, as well as the position he occupies in this abstract space. Each age…

2019-09-09abs ↗pdf ↗

WrapNet optimizes inference for low-resolution neural networks by using 8-bit additions.

problem Reducing multiplication complexity in low-resolution neural networks.
method Adapting neural networks to use low-resolution (8-bit) additions in accumulators, with a cyclic activation layer and overflow penalty regularizer.
result Achieves comparable classification accuracy to 32-bit counterparts using low-resolution additions.

This work addresses the instability in asynchronous data parallel optimization. It does so by introducing a novel distributed optimizer which is able to efficiently optimize a centralized model under communication constraints. The optimizer achieves this by pushing a normalized sequence of first-order gradients to a pa…

2017-10-06abs ↗pdf ↗

For any nonorientable closed surface, we determine the minimal dilatation among pseudo-Anosov mapping classes arising from Penner's construction. We deduce that the sequence of minimal Penner dilatations has exactly two accumulation points, in contrast to the case of orientable surfaces where there is only one accumula…

2018-07-24abs ↗pdf ↗

Improves matrix multiplication throughput for asymmetric bit-width operands.

problem Matrix multiplications between asymmetric bit-width operands, especially 8- and 4-bit, are not efficiently handled by existing SIMD instructions.
method Proposes a new SIMD matrix multiplication instruction that uses mixed precision on inputs (8- and 4-bit) and accumulates into 16-bit output, improving throughput.
result Offers 2x improvement in throughput compared to existing symmetric-operand-size instructions, with negligible overflow.

The paper improves self-training in semi-supervised learning by selecting more robust pseudo-labeled data.

problem Improving the reliability of pseudo-labeled data selection in self-training for semi-supervised learning.
method Proposes a multi-objective utility function to select pseudo-labeled data that maximizes reliability, considering model selection, accumulation of errors, and covariate shift uncertainties.
result Robustness towards model choice can lead to substantial accuracy gains in self-training.

Paper excludes the lowest energy level as an accumulation point for harmonic maps into analytic manifolds.

problem Analytic manifolds and their harmonic maps energy spectrum.
method Exclusion of the lowest energy level as an accumulation point using obstructions to the gluing of harmonic spheres and Lojasiewicz-estimates.
result Proves that the lowest energy level is not an accumulation point for generic 3-manifolds.

We consider a family of manifolds with a class of degenerating warped product metrics gε=ρ(ε,t)2adt2+ρ(ε,t)2bdsM2g_ε=ρ(ε,t)^{2a}dt^2 +ρ(ε,t)^{2b}ds_M^2, with MM compact, ρρ homogeneous degree one, a1a \le -1 and b>0b > 0. We study the Laplace operator acting on L2L^{2} differential pp-forms and give sharp accumulation rates for eigenvalues n…

2003-11-14abs ↗pdf ↗

Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning. However, in practice, commonly used off-policy approximate dynamic programming methods based on Q-learning and actor-critic methods are highly sensitive to the data distribution, and can make only …

2019-06-03abs ↗pdf ↗

In this paper we consider sparse approximation problems, that is, general l0l_0 minimization problems with the l0l_0-"norm" of a vector being a part of constraints or objective function. In particular, we first study the first-order optimality conditions for these problems. We then propose penalty decomposition (PD) me…

2012-05-10abs ↗pdf ↗

AdaX improves Adam by exponentially accumulating past gradients, leading to better performance in machine learning tasks.

problem Adam's fast convergence can lead to local minimums in non-convex problems.
method AdaX exponentially accumulates past gradients to adaptively tune the learning rate.
result AdaX outperforms Adam in various machine learning tasks, including computer vision and natural language processing.

FetchSGD reduces communication in federated learning with sketching.

problem Communication bottlenecks and convergence issues in federated learning.
method FetchSGD uses Count Sketch to compress and merge model updates efficiently.
result FetchSGD achieves high compression rates and good convergence without sparse client participation.

Twin-to-twin transfusion syndrome treatment requires fetoscopic laser photocoagulation of placental vascular anastomoses to regulate blood flow to both fetuses. Limited field-of-view (FoV) and low visual quality during fetoscopy make it challenging to identify all vascular connections. Mosaicking can align multiple ove…

2019-07-15abs ↗pdf ↗

The paper analyzes how synthetic data training degrades diffusion models, providing bounds and characterizing different drift regimes.

problem The degradation of performance in diffusion models trained on synthetic data.
method Theoretical analysis of score-based diffusion models, focusing on the accumulated divergence between generated and target distributions.
result Upper and lower bounds on the accumulated divergence, providing the first lower bound for diffusion models.

Generative diffusion models improve financial LOB simulation and forecasting.

problem High noise and complexity in financial LOB data makes deep generative models ineffective.
method Convert LOB data to images, apply diffusion models with inpainting for long-term sequence generation.
result Our method achieves state-of-the-art performance on LOB-Bench, improving coherence over local details.

New study shows MLE can avoid model collapse with gradual synthetic data addition.

problem Model collapse in generative models trained on synthetic data.
method Theoretical study of maximum likelihood estimation (MLE) under iterative training with accumulating synthetic data.
result Non-asymptotic bounds show MLE can avoid model collapse even as real data fraction vanishes.

The paper establishes criteria for spacetime inextendibility using asymptotic volume-distance-ratio analysis.

problem Determining inextendibility of spacetimes near singularities.
method Asymptotic analysis of volume-distance-ratio (VDR) to prove inextendibility criteria.
result Failure of VDR convergence to the Minkowski value implies inextendibility of spacetime.

Study compares price limit and circuit breaker effects in stock markets.

problem Preventing rapid and steep price drops in stock exchanges.
method Agent-based model for financial market simulation.
result Price limit and circuit breaker have similar effects under same conditions, but price limit less effective with shorter limit time range.

Study of homeomorphisms on infinite type surfaces with a classification theorem.

problem Classifying homeomorphisms on surfaces of infinite type.
method Introduce tame homeomorphisms and prove a Nielsen-Thurston type classification theorem.
result For tame homeomorphisms, surfaces decompose into invariant subsurfaces with canonical decompositions.