Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

146293439585 · Jun 202019922001200920172026
48 results for small-scale datasets

CST-YOLO improves blood cell detection with YOLOv7 and CNN-Swin Transformer.

problem Small-scale object detection in blood cells.
method YOLOv7 architecture enhanced with CNN-Swin Transformer, W-ELAN, MCS, CatConv.
result CST-YOLO achieves 92.7%, 95.6%, and 91.1% mAP@0.5 on three blood cell datasets.

Study how large-scale flows align small-scale vortices in 3D Euler equations.

problem Understanding how large-scale flows align small-scale vortices in 3D Euler equations.
method Constructing a Lagrangian coordinate to identify when the Lie bracket is zero and investigating the locality of the pressure term.
result Clarified conditions under which small-scale vortices are aligned by large-scale flows.

Study of splitting maps in Type I Ricci flows for understanding singular set structure.

problem Understanding the structure of the singular set in non-collapsed Ricci limit spaces.
method Construction and investigation of almost splitting maps on Ricci flows that are almost self-similar.
result Sharp splitting maps remain splitting maps at smaller scales under certain conditions.

The goal of this article is to draw new applications of small scale quantum ergodicity in nodal sets of eigenfunctions. We show that if quantum ergodicity holds on balls of shrinking radius r(λ)0r(λ) \to 0, then one can achieve improvements on the recent upper bounds of Logunov and Logunov-Malinnikova on the size of nodal…

2016-06-07abs ↗pdf ↗

PePR scores assess DL model performance per resource unit, promoting smaller, more efficient models.

problem Limited access to large-scale resources hinders medical image analysis research.
method Introduced PePR score to measure DL model performance per resource unit.
result Small-scale, specialized models outperform large-scale models in resource-constrained settings.

In 1960 Reifenberg proved the topological disc property. He showed that a subset of RnR^n which is well approximated by mm-dimensional affine spaces at each point and at each (small) scale is locally a bi-Hölder image of the unit ball in RmR^m. In this paper we prove that a subset of R3R^3 which is well approximated b…

2006-07-18abs ↗pdf ↗

Employing profits data of Japanese companies in 2002 and 2003, we confirm that Pareto's law and the Pareto index are derived from the law of detailed balance and Gibrat's law. The last two laws are observed beyond the region where Pareto's law holds. By classifying companies into job categories, we find that companies …

2005-06-08abs ↗pdf ↗

Study on scalar curvature bounds and manifold topological complexity.

problem Understanding the topological complexity of manifolds with scalar curvature constraints.
method Introduced a small scale index theorem to establish bounds for Gromov's simplicial norm.
result Upper bound for Gromov's simplicial norm established in terms of scalar curvature, volume, and injectivity radius.

Transfer learning improves chaotic dynamics predictions with less data.

problem Efficiently predicting chaotic dynamics with limited data.
method Transfer learning for nonlinear dynamics, optimizing transfer rate and leveraging small-scale turbulence universality.
result Significantly more accurate inference of chaotic dynamics achieved.

We study the problem of instance segmentation in biological images with crowded and compact cells. We formulate this task as an integer program where variables correspond to cells and constraints enforce that cells do not overlap. To solve this integer program, we propose a column generation formulation where the prici…

2017-09-21abs ↗pdf ↗

A topology on a set XX is the same as a projection (i.e. an idempotent linear operator) cl:2X2Xcl:2^X\to 2^X satisfying Acl(A)A\subset cl(A) for all AXA\subset X. That's a good way to summarize Kuratowski's closure operator. Basic geometry on a set XX is a dot product :2X×2X2Y\cdot:2^X\times 2^X\to 2^Y. Its equivalent form is an or…

2018-03-24abs ↗pdf ↗

This work explores how neural architecture search can improve adversarial robustness without adversarial training.

problem Improving adversarial robustness of neural networks without adversarial training.
method Experimented with hand-crafted and NAS-based architectures to compare robustness to PGD attacks.
result NAS-based architectures are more robust for small-scale attacks, but hand-crafted architectures are more robust for larger datasets and tasks.

In this paper, we propose a game theoretical adversarial intervention detection mechanism for reliable smart road signs. A future trend in intelligent transportation systems is ``smart road signs" that incorporate smart codes (e.g., visible at infrared) on their surface to provide more detailed information to smart veh…

2019-01-30abs ↗pdf ↗

In this short note we show that the lower bounds of Mangoubi on the inner radius of nodal domains can be improved for quantum ergodic sequences of eigenfunctions, according to a certain power of the radius of shrinking balls on which the eigenfunctions equidistribute. We prove such improvements using a quick applicatio…

2016-06-10abs ↗pdf ↗

We study the classification of ultrametric spaces based on their small scale geometry (uniform homeomorphism), large scale geometry (coarse equivalence) and both (all scale uniform equivalences). We prove that these equivalences can be characterized with parallel constructions using a combinatoric tool called common zi…

2009-09-01abs ↗pdf ↗

Meta-learn Bayesian inference for task-specific BNNs using amortised inference.

problem Efficiently learning Bayesian inference for small-scale probabilistic meta-learning.
method Replace global inducing points with actual data to create a set of approximate likelihoods, train a meta-model to learn these parameters across related datasets.
result Meta-learned inference can be applied to task-specific BNNs, improving efficiency and scalability.

In this paper we investigate the performance of different types of rectified activation functions in convolutional neural network: standard rectified linear unit (ReLU), leaky rectified linear unit (Leaky ReLU), parametric rectified linear unit (PReLU) and a new randomized leaky rectified linear units (RReLU). We evalu…

2015-05-05abs ↗pdf ↗

Study tests rough fractional volatility model across different time scales, revealing new volatility patterns.

problem Testing robustness of rough fractional volatility model over various time scales.
method Used large dataset on FX rates, included smoothing and measurement errors, analyzed log-log plots of realized variance increments.
result Found new stylized facts in volatility patterns, including convexity and nonlinear behavior.

Fidel-TS creates a new benchmark for time series forecasting models.

problem Lack of high-quality benchmarks for time series forecasting models.
method Formalized high-fidelity benchmark principles, including data sourcing integrity, leak-free design, and structural clarity. Created Fidel-TS, a new large-scale benchmark.
result Demonstrated the limitations of prior benchmarks and potential discrepancies in model evaluation.

Shifts dataset evaluates uncertainty in real-world tasks across modalities.

problem Lack of standard datasets for evaluating uncertainty estimation and robustness to distributional shift.
method Proposes Shifts Dataset for evaluation of uncertainty estimates and robustness to distributional shift across tabular, audio, text, and sensor data.
result Baseline results for tabular weather prediction, machine translation, and SDC vehicle motion prediction.

This paper is devoted to dualization of dimension-theoretical results from the small scale to the large scale. So far there are two approaches for such dualization: one consisting of creating analogs of small scale concepts and the other amounting to the covering dimension of the Higson corona ν(X)ν(X) of XX. The first …

2013-04-22abs ↗pdf ↗

Stochastic Sparse Subspace Clustering improves subspace clustering by reducing over-segmentation through dropout.

problem Over-segmentation in subspace clustering.
method Introducing dropout regularization to enforce denser connections between points from the same subspace.
result Stochastic Sparse Subspace Clustering effectively handles large datasets and reduces over-segmentation.

DD-SP uses ML to improve SP for Lorenz 96 systems, outperforming LR and DD-P.

problem Improving computational efficiency in weather/climate modeling.
method Data-driven super-parameterization using recurrent neural networks.
result DD-SP is more accurate and cheaper than SP, especially with scale separation.

Paper proposes data quality measures for large-scale high-dimensional data.

problem Lack of practical data quality measures for large-scale high-dimensional data.
method Proposes two data quality measures: class separability and in-class variability. Efficient algorithms based on random projections and bootstrapping are provided.
result Efficient algorithms for computing data quality measures on large-scale high-dimensional data.

Proposes a method for differentially private linear regression and synthetic data generation.

problem Lack of valid inference and synthetic data generation methods for small-scale datasets in privacy-aware settings.
method Gaussian differentially private linear regression with bias-corrected estimator and SDG procedure.
result Improves accuracy and provides valid confidence intervals for downstream tasks.

Efficiently applies NTK to large-scale datasets using random features.

problem Computational limitations of kernel methods for large-scale datasets.
method Proposes a sketching-based algorithm combining random features of arc-cosine kernels to construct an efficient feature map of the NTK.
result Achieves comparable error bounds to exact kernel methods but with significantly reduced feature dimensionality.

Investigates offline RL in factorisable action spaces, overcoming overestimation bias.

problem Overestimation bias in value estimates for unseen state-action pairs.
method Value-decomposition approach in DecQN, adapted for factorised discrete action spaces.
result Demonstrates the effectiveness of factorised approach in offline RL.

In this study we examine the evolution of price, volume, and the bid-ask spread after extreme 15 minute intraday price changes on the NYSE and the NASDAQ. We find that due to strong behavioral trading there is an overreaction. Furthermore we find that volatility which increases sharply at the event decays according to …

2004-01-06abs ↗pdf ↗

This work introduces a protocol to automatically select the correct range of scales for meaningful Intrinsic Dimension estimation.

problem The Intrinsic Dimension (ID) varies with scale in real-world datasets, leading to erroneous results.
method The protocol selects the correct range of scales by ensuring constant density of data points.
result The method provides a robust and scale-adaptive approach to estimating meaningful Intrinsic Dimension.

Generative Adversarial Networks (GANs) are an elegant mechanism for data generation. However, a key challenge when using GANs is how to best measure their ability to generate realistic data. In this paper, we demonstrate that an intrinsic dimensional characterization of the data space learned by a GAN model leads to an…

2019-05-02abs ↗pdf ↗

Deep linear networks minimize sharpness, avoiding large eigenvalues.

problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.

Deep reinforcement learning (deep RL) has been successful in learning sophisticated behaviors automatically; however, the learning process requires a huge number of trials. In contrast, animals can learn new tasks in just a few trials, benefiting from their prior knowledge about the world. This paper seeks to bridge th…

2016-11-09abs ↗pdf ↗

Neuro-inspired recurrent neural network algorithms, such as echo state networks, are computationally lightweight and thereby map well onto untethered devices. The baseline echo state network algorithms are shown to be efficient in solving small-scale spatio-temporal problems. However, they underperform for complex task…

2018-08-01abs ↗pdf ↗

We focus in this work on the estimation of the first kk eigenvectors of any graph Laplacian using filtering of Gaussian random signals. We prove that we only need kk such signals to be able to exactly recover as many of the smallest eigenvectors, regardless of the number of nodes in the graph. In addition, we address…

2016-11-03abs ↗pdf ↗

Generalizes bits back coding for time-series models with latent Markov structures.

problem Efficiently compressing time-series data with latent Markov structures.
method Extends bits back coding to time-series models with latent Markov structures, including HMMs and LGSSMs.
result Effective for small scale models, promising for larger scale settings like video compression.