Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2935868791,172 · Jun 202019922001200920172026
48 results for data aspect ratio

Paper shows how to embed Möbius bands with many twists and small aspect ratios.

problem Finding the smallest aspect ratio for Möbius bands with many twists.
method Constructs a folded paper ribbon knot to bound the aspect ratio.
result Paper Möbius bands and annuli with any number of half-twists can be embedded with aspect ratio less than 8.

New equivalences found between subsampling and ridge regularization methods.

problem Establishing precise structural and risk equivalences between subsampling and ridge regularization.
method Proved structural and risk equivalences between subsample ridge estimators and different ridge regularization levels and subsample aspect ratios.
result Optimally tuned ridge regression exhibits a monotonic prediction risk in the data aspect ratio.

Gradient descent dynamics in nonconvex models explained with universality.

problem Understanding long-time behavior of nonconvex gradient descent.
method Developed a state evolution system for tracking gradient descent iterates.
result Gradient descent iterates are approximately independent of data and strongly incoherent with feature vectors.

Study optimal ridge regularization for out-of-distribution prediction.

problem Optimal ridge regularization for predicting out-of-distribution data.
method Established conditions for optimal regularization under covariate and regression shifts, proving monotonic risk in data aspect ratio.
result Negative regularization can be optimal under shifts, even with isotropic or underparameterized training features.

Optimizes kernel density ratios for better predictions and information measures.

problem Improving accuracy of kernel density estimates for density ratios.
method Derives an optimal weight function using calculus of variations.
result Reduces bias in kernel density estimates, leading to improved prediction posteriors and information-theoretic measures.

Noise increases the Rashomon ratio, leading simpler models to perform similarly to complex ones.

problem Why simpler models perform similarly to complex models on noisy datasets.
method Analyzed the data generation process and model training choices, introduced pattern diversity.
result Noisier datasets lead to larger Rashomon ratios, explaining simpler models' performance.

Paper introduces lexical ratio to measure portfolio diversification.

problem Traditional diversification metrics overlook non-numerical relationships.
method Uses textual data to capture diversification dimensions through entropy-based insights.
result Lexical ratio (LR) outperforms traditional metrics in optimizing portfolio returns.

Aggregation defenses improve deep learning models' robustness against data poisoning attacks.

problem Data poisoning attacks manipulate deep learning models with malicious training samples.
method Deep Partition Aggregation, efficiency improvements, data-to-complexity ratio, poisoning overfitting phenomenon.
result Aggregation defenses boost poisoning robustness through the poisoning overfitting phenomenon.

We present an elementary analysis of the dynamical aspects of the GDP / government surplus multiplier with relevance to the assessment of a country's debt repayment policy. We show the (at first) counter intuitive result that in order to reduce the Debt/GDP ratio, countries with high Debt to GDP should go into further …

2013-10-11abs ↗pdf ↗

Importance weighting is a general way to adjust Monte Carlo integration to account for draws from the wrong distribution, but the resulting estimate can be highly variable when the importance ratios have a heavy right tail. This routinely occurs when there are aspects of the target distribution that are not well captur…

2015-07-09abs ↗pdf ↗

The paper analyzes the benefit-cost ratio for feature selection in machine learning.

problem Tackling the challenge of distinguishing relevant features from noise in feature selection.
method Simulation study with different cost and data settings to analyze the benefit-cost ratio.
result The benefit-cost ratio can overemphasize cheap noise features in scenarios with large cost differences and small effect sizes.

Framework mitigates risk non-monotonicity in high-dimensional predictions.

problem Risk non-monotonicity in high-dimensional predictions.
method Model-agnostic framework using cross-validation and data-driven methodologies (zero- and one-step).
result Modified prediction procedures achieve monotonic asymptotic risk behavior.

We present a reinforcement learning approach for detecting objects within an image. Our approach performs a step-wise deformation of a bounding box with the goal of tightly framing the object. It uses a hierarchical tree-like representation of predefined region candidates, which the agent can zoom in on. This reduces t…

2018-10-15abs ↗pdf ↗

Solves constant pre-factor problem for tt*-Toda equations using asymptotic data and symplectic structures.

problem Constant pre-factor problem for the tt*-Toda equations.
method Explicit evaluation using asymptotic data and introduction of symplectic structures.
result Preservation of symplectic structures by Riemann-Hilbert correspondence for wider class of solutions.

It had been believed in the conventional practice that the risk of a bank going bankrupt is lessened in a straightforward manner by transferring the risk of loan defaults. But the failure of American International Group in 2008 posed a more complex aspect of financial contagion. This study presents an extension of the …

2014-09-25abs ↗pdf ↗

In this paper, we propose a novel learning method for image classification called Between-Class learning (BC learning). We generate between-class images by mixing two images belonging to different classes with a random ratio. We then input the mixed image to the model and train the model to output the mixing ratio. BC …

2017-11-28abs ↗pdf ↗

New algorithm for MDS with quasi-polynomial dependency on aspect ratio.

problem Finding an embedding that minimizes a specific objective function for given dissimilarities.
method A novel geometry-aware analysis of a conditional rounding of the Sherali-Adams LP hierarchy.
result Achieved a solution with cost \(O(\log Δ) \cdot extrm{OPT}^{Ω(1)} + ε\) in quasi-polynomial time.

Data augmentations are important ingredients in the recipe for training robust neural networks, especially in computer vision. A fundamental question is whether neural network features encode data augmentation transformations. To answer this question, we introduce a systematic approach to investigate which layers of ne…

2020-02-29abs ↗pdf ↗

PolyModel theory and iTransformer improve hedge fund portfolio construction.

problem Sparse financial time series data makes portfolio construction challenging.
method Identify asset pool, select risk factors, create quantitative and classical measures, and use iTransformer for trend capture.
result Improved Sharpe ratio and annualized return compared to benchmarks.

Study ridge ensembles in proportional feature-to-sample size regime, proving risk equivalence and GCV consistency.

problem Characterizing and optimizing ridge ensembles in proportional feature-to-sample size regimes.
method Proportional asymptotics analysis, GCV for tuning, proving risk equivalence.
result Risk of optimal full ridgeless ensemble matches optimal ridge predictor's risk.

DeepCTRL integrates rules into deep learning models, allowing flexible control at inference.

problem Lack of flexibility in incorporating rules into deep learning models.
method Integrates rule representations into deep neural networks, enabling flexible control at inference.
result Improves rule verification ratio and accuracy gains at downstream tasks.

The book explains deep learning theory and how networks learn nontrivial representations.

problem Understanding and optimizing deep neural networks.
method Developed RG flow to characterize signal propagation, solved layer-to-layer equations, and analyzed representation learning.
result Predictions of trained networks are nearly-Gaussian, with depth-to-width ratio controlling deviations.

Estimates Gaussian location model with ridge regularization, comparing variational and spectral methods.

problem Estimating parameters in Gaussian location model with regularization.
method Ridge-regularized log-density-ratio estimation, variational and spectral approaches.
result Regularized variational estimator has lower risk with many observations, spectral estimator with fewer observations.

Paper proposes new strategies for better portfolio estimation in long-term investments with unknown distributions.

problem Worse out-of-sample performance of estimated portfolios due to unknown future data distribution.
method Online learning framework, dynamic sequential portfolios, updating risk aversion coefficient.
result Dynamic strategies achieve asymptotically optimal utility, Sharpe ratio, and growth rate of true portfolios.

FS-GCLSTM predicts stock returns by leveraging value-chain relationships.

problem Traditional time series models fail to capture complex interdependencies in modern markets.
method FS-GCLSTM integrates value-chain networks and graph convolutions to predict stock returns.
result FS-GCLSTM consistently delivers superior portfolio performance compared to traditional models.

We study the problem of detecting the presence of a single unknown spike in a rectangular data matrix, in a high-dimensional regime where the spike has fixed strength and the aspect ratio of the matrix converges to a finite limit. This setup includes Johnstone's spiked covariance model. We analyze the likelihood ratio …

2018-02-20abs ↗pdf ↗

Researchers validate ML scenario generators by checking dependencies and detecting memorization effects.

problem Validation of machine learning-based scenario generators differs from classical methods due to data-driven dependencies.
method Two novel validation aspects: checking dependencies and detecting memorization effects. Novel memorization ratio introduced.
result Validation methods successfully detect dependencies and memorization effects in ML-based scenario generators.

The paper proposes an asset allocation strategy using the Sortino ratio for better performance.

problem Traditional asset allocation methods like the Sharpe ratio do not penalize negative returns adequately.
method The Sortino ratio is used to maximize asset allocation, penalizing only negative return variances.
result The Sortino ratio-based strategy outperforms traditional methods like the Kelly criterion.

The paper proposes using density ratio estimation to evaluate synthetic data quality.

problem Improving the quality and utility of synthetic data for analysis.
method Density ratio estimation to measure synthetic data quality.
result Density ratio estimation yields more accurate global utility estimates than existing methods.

This study argues for pruning trees in random forests to improve performance in low signal-to-noise scenarios.

problem Improving random forest performance in scenarios with low signal-to-noise ratio.
method Using regularization theory, the study re-examines the depth of trees in random forests and provides evidence that shallow trees are advantageous.
result Random forests with shallow trees are advantageous when the signal-to-noise ratio is low.

Hedge Funds are considered as one of the portfolio management sectors which shows a fastest growing for the past decade. An optimal Hedge Fund management requires an appropriate risk metrics. The classic CAPM theory and its Ratio Sharpe fail to capture some crucial aspects due to the strong non-Gaussian character of He…

2006-10-20abs ↗pdf ↗

Proposes a method to cluster multi-aspect data using manifold learning with NMF.

problem Clustering multi-aspect data with diverse features and views.
method Includes inter-manifold learning in NMF framework to handle different data types.
result The method improves clustering accuracy and efficiency on various datasets.

This article considers algorithmic and statistical aspects of linear regression when the correspondence between the covariates and the responses is unknown. First, a fully polynomial-time approximation scheme is given for the natural least squares optimization problem in any constant dimension. Next, in an average-case…

2017-05-19abs ↗pdf ↗

Proposes a robust method for predicting missing outcomes in covariate shift adaptation.

problem Predicting missing outcomes in test data with covariate shift.
method Doubly robust estimator for covariate shift adaptation via importance weighting, incorporating an additional estimator for the regression function.
result Shows robustness against density-ratio estimation errors, maintaining consistency if either estimator is consistent.

New financial ratios using compositional data improve analysis of firm health.

problem Statistical issues with standard financial ratios, especially skewness and outliers.
method Compositional data (CoDa) methodology to analyze financial statements.
result Outliers and skewness reduced, results invariant to numerator and denominator permutation.