Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

69137206274 · Jun 202019922001200920182026
48 results for account size

Study assesses environmental management accounting practices in Bangladesh.

problem Low environmental management accounting practices in Bangladeshi manufacturing companies.
method Developed a compliance checklist and evaluated practices using binary scoring.
result Environmental management accounting practices are poor in Bangladeshi manufacturing companies.

A new scaling law predicts optimal batch size for training models.

problem Finding the optimal batch size for training models efficiently.
method Proposed a three-term scaling law that considers model size, training data, training steps, and batch size.
result The three-term law accurately recovers the optimal batch size and can be robustly fit with fewer training runs.

Study compares Islamic banks' accounting and market performance.

problem Assessing the relationship between Islamic banks' accounting and market performance.
method Selected six Islamic banks, collected data from 2009-2013, used random-effect models.
result Superior accounting performance does not correlate with superior market performance.

Investigates the impact of batch size on GPU and TPU performance.

problem Optimizing performance of GPUs and TPUs during training and inference phases.
method Investigated the impact of batch size on performance of GPUs and TPUs using standard MNIST and Fashion-MNIST datasets.
result Significant speedup was achieved even with low-scale usage of TPUv2 units, up to 10x for training and 2x for prediction.

Active learning performance degrades with larger batch sizes, but can be mitigated with smaller window sizes.

problem Impact of batch size on stopping active learning for text classification.
method Analyzed the impact of batch size on a stopping method for active learning in text classification, finding that larger batch sizes degrade performance and that using smaller window sizes mitigates this effect.
result Mitigating batch size degradation in active learning for text classification can be achieved by adjusting the window size parameter.

Muon optimizes training efficiency by improving data retention at large batch sizes.

problem Improving training efficiency and data retention at large batch sizes.
method Introducing Muon, a second-order optimizer, and combining it with muP for efficient hyperparameter transfer.
result Muon outperforms AdamW in retaining data efficiency at large batch sizes, enabling more economical training.

FIA misleads in small samples; lower-bound sample size NN' prevents errors.

problem Misleading model selection by FIA in small samples.
method Proposes a lower-bound sample size NN' to prevent FIA's misleadingness.
result Prevents inversion of model complexity ranks in FIA.

This paper tackles post-trade allocation inefficiencies and presents a uniform return allocation method.

problem Return divergence among accounts after trade allocation.
method Systematic treatment of trade allocation risk, presenting a uniform return allocation method.
result Uniform allocation of returns irrespective of the number of accounts and trade sizes.

One dimensional stylized model taking into account spatial activity of firms with uniformly distributed customers is proposed. The spatial selling area of each firm is defined by a short interval cut out from selling space (large interval). In this representation, the firm size is directly associated with the size of i…

2007-10-02abs ↗pdf ↗

SCS identifies a range of plausible equally weighted portfolios, quantifying selection uncertainty.

problem Uncertainty in selecting the best equally weighted portfolio subset.
method Introduces Selection Confidence Set (SCS) for EWPs, covering plausible portfolios with high probability.
result SCS quantifies selection uncertainty and covers the unknown optimal selection with high probability.

Size effect persists in equity markets, with CMH portfolios less correlated to Low-Vol anomaly.

problem The persistence and significance of the size effect in equity markets.
method Analysis of dollar-turnover, ββ-neutralisation, and Low-Vol neutralisation.
result Size-based portfolios are less anti-correlated to Low-Vol anomaly compared to market-cap based SMB.

New insights into SGD and SGD-M in high dimensions.

problem Understanding and comparing SGD and SGD-M in high-dimensional settings.
method Developed high-dimensional scaling limits for SGD-M and online SGD, examining their dynamics and performance.
result SGD-M amplifies high-dimensional effects, potentially degrading performance compared to online SGD.

Predicts optimal training dataset sizes per class for machine learning models.

problem Optimizing training dataset sizes for class-specific machine learning models.
method Algorithm based on space-filling design of experiments, models like powerlaw curves and generalized linear models.
result The algorithm predicts optimal training dataset sizes per class for improved model performance.

New algorithm SAGA++ improves on SAGA for faster convergence in stochastic batch size methods.

problem Optimizing convergence rate in stochastic variance reduction methods with batch size.
method Proposes SAGA++ algorithm with optimal average batch size considering cache/disk IO effects.
result SAGA++ outperforms SAGA and other solvers on real datasets.

Sample measures of top centile contributions to the total (concentration) are downward biased, unstable estimators, extremely sensitive to sample size and concave in accounting for large deviations. It makes them particularly unfit in domains with power law tails, especially for low values of the exponent. These estima…

2014-05-08abs ↗pdf ↗

Pruning improves model generalization in over-parameterized models, contradicting traditional theories.

problem Pruning's effect on generalization in over-parameterized models.
method Empirical study on standard pruning algorithms and additional regularization effects.
result Pruning leads to better training and regularization, improving generalization.

Model shows different trading behaviors during financial crisis.

problem Understanding trading dynamics during financial crises.
method Implemented a market microstructure model with informed, uninformed, and heuristic-driven traders.
result Heuristic-driven trading remains constant during financial crisis, while informed trading varies.

The key idea of this model is that firms are the result of an evolutionary process. Based on demand and supply considerations the evolutionary model presented here derives explicitly Gibrat's law of proportionate effects as the result of the competition between products. Applying a preferential attachment mechanism for…

2012-08-06abs ↗pdf ↗

Mixtures of neural operators reduce active complexity in operator learning.

problem Reduction of active complexity in operator learning models.
method Constructive comparison between routed mixtures of neural operators (MoNOs) and a fixed single-neural-operator construction.
result Every scalar uniformly continuous nonlinear operator can be approximated by a MoNO whose active expert has smaller depth, width, and rank scaling.

With recent advances in high throughput technology, researchers often find themselves running a large number of hypothesis tests (thousands+) and esti- mating a large number of effect-sizes. Generally there is particular interest in those effects estimated to be most extreme. Unfortunately naive estimates of these effe…

2013-11-15abs ↗pdf ↗

Estimates unknown population sizes using the hypergeometric distribution.

problem Estimating discrete distributions with unknown population sizes and category sizes.
method Proposes a novel solution using the hypergeometric likelihood, accounting for a data generating process with a latent variable.
result Empirically demonstrates superior performance in estimating population sizes and learning latent spaces compared to other methods.

New framework uses social media images to estimate wildlife populations.

problem Lack of basic data for wildlife species due to inadequate traditional methods.
method Developed a new computer vision tool to account for social media bias.
result Showed that wildlife population size estimates are learnable from social media.

Scalable GPLVM reduces complexity in scRNA-seq data, accounting for technical and biological confounders.

problem Complexity and confounders in scRNA-seq data hamper interpretation.
method Extended Gaussian process latent variable model (GPLVM) to handle large datasets.
result Framework reconstructs latent signatures and captures disease-specific gene expression.

Novel method uses Gaussian process to estimate particle sizes from scattering data.

problem Estimating particle size distributions from noisy optical scattering measurements.
method Constrained Gaussian process regression with normalization constraints.
result Accurately reconstructs particle size distributions from noisy data.

Tree-based models outperform deep learning on tabular data, especially for medium-sized datasets.

problem Understanding why tree-based models outperform deep learning on tabular data.
method Extensive benchmarks of tree-based and deep learning models on 45 datasets, accounting for hyperparameters.
result Tree-based models remain state-of-the-art on medium-sized tabular data, even without hyperparameter optimization.

DA-MLNs improve MLNs by scaling feature weights based on domain size.

problem Extreme probabilities in MLNs when testing on different domain sizes.
method DA-MLNs divide ground feature weights by a scaling factor that depends on the number of connections.
result DA-MLNs achieve significantly higher accuracy on domains with different sizes compared to standard MLNs.

Paper presents adaptive minimax risk classifiers for multidimensional concept drift.

problem Multidimensional concept drift in supervised classification.
method Adaptive minimax risk classifiers (AMRCs) tracking multivariate and high-order distribution changes.
result AMRCs provide computable tight performance guarantees and improve classification.

We consider the problem of portfolio optimization in the presence of market impact, and derive optimal liquidation strategies. We discuss in detail the problem of finding the optimal portfolio under Expected Shortfall (ES) in the case of linear market impact. We show that, once market impact is taken into account, a re…

2010-04-23abs ↗pdf ↗

Optimal decision trees are constructed via integer programming for better accuracy and interpretability.

problem Overfitting and loss of interpretability in decision trees.
method Mixed integer programming formulation to construct optimal decision trees of a prespecified size, considering categorical and numerical features.
result Very good accuracy can be achieved with small trees using moderately-sized training sets.

A new test for volatility in clustered time series data, robust to distributional assumptions.

problem Volatility issues in clustered multiple time series data, especially in stock market indicators.
method Bootstrap method for multiple time series, accounting for contagion effect.
result The test is correctly sized and powerful, especially for stationary mean and contained volatility in fewer clusters.

Study non-monotonic loss functions in CRC, achieving valid risk control with large calibration samples.

problem Non-monotonic loss functions in CRC, violating existing theory's monotonicity assumption.
method Finite grid selection, calibration sample size analysis, Lipschitz continuity, monotonicity, distribution shift.
result Valid CRC achieved with large calibration samples, optimal excess risk rate of log(m)/n\sqrt{\log(m)/n}.

Network analysis reveals regional banking clusters during financial crisis.

problem Understanding how financial institutions react to systemic crises.
method Extracting Accounting Network from financial statements, applying quality checks, community detection, PCA.
result Regional banking clusters emerge, with US and Japanese banks dominating, reflecting global practices.

Deep neural networks predict earthquake locations with high accuracy.

problem Predicting the location of earthquakes with high precision.
method Recurrent Convolutional Neural Networks (R-CNN) model that accounts for spatio-temporal dependencies.
result Neural networks model outperforms baseline models in predicting earthquakes with ROC AUC 0.975 and PR AUC 0.0890.

New TVBO algorithm optimizes time-varying functions with varying sampling frequencies.

problem Optimizing time-varying, expensive, noisy functions with constant frequency assumption.
method Formulated practical recommendations and derived upper regret bound for varying sampling frequencies.
result BOLT algorithm outperforms state-of-the-art TVBO algorithms in experiments.

A new criterion selects models in overparameterized settings.

problem Model selection for overparameterized models with more parameters than data.
method Establishes Bayesian duality and introduces the Interpolating Information Criterion.
result The Interpolating Information Criterion selects models in overparameterized settings.

The paper improves regret bounds for admission control in queueing systems.

problem Improving regret bounds for admission control in queueing systems.
method Proposes an algorithm inspired by UCRL2 and uses problem structure to bound regret.
result Proves an upper bound on the expected total regret of O(SlogT+mTlogT)O(S\log T + \sqrt{mT \log T}).