Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

136273409545 · Jun 202019922001200920182026
48 results for very large

The paper examines the co-rank of weakly parafree 3-manifold groups and provides counterexamples.

problem Investigating the co-rank of weakly parafree 3-manifold groups and their relationship to free groups.
method Constructing specific examples of homology handlebodies and analyzing their fundamental groups.
result The fundamental group of W. Thurston's tripus manifold is not very large, showing that weakly parafree groups of rank at least 3 are not necessarily very large.

We present a parallelized bijective graph matching algorithm that leverages seeds and is designed to match very large graphs. Our algorithm combines spectral graph embedding with existing state-of-the-art seeded graph matching procedures. We justify our approach by proving that modestly correlated, large stochastic blo…

2013-10-04abs ↗pdf ↗

The paper proposes an efficient method to scale Bayesian inference for mixed multinomial logit models to very large datasets.

problem Efficiency in Bayesian inference for mixed multinomial logit models on large datasets.
method Amortized Variational Inference with stochastic backpropagation, automatic differentiation, and GPU acceleration.
result The proposed method achieves significant computational speedups over traditional methods for large datasets.

Deep GNNs and self-supervision boost graph learning at scale.

problem Efficiently deploying GNNs at large scale remains challenging.
method Two large-scale GNNs: a deep transductive node classifier and a very deep inductive graph regressor.
result Award-level performance on MAG240M and PCQM4M benchmarks.

Efficiently trains large corpora models without sampling.

problem Training neural network embedding models on very large corpora using SGD is expensive.
method Proposes new methods to train models without sampling unobserved pairs, using Gramian estimation and variance reduction schemes.
result Significant improvement in training time and generalization quality compared to traditional methods.

Stochastic EP improves memory efficiency for large datasets in Gaussian process classification.

problem Memory limitations in EP for large datasets.
method Stochastic Expectation Propagation (EP) for large scale Gaussian process classification.
result Stochastic EP avoids memory scaling with dataset size, improving scalability.

Study volatility models with rough paths, focusing on large deviations and option behavior.

problem Analyzing volatility in financial markets with very rough paths.
method Introduced time-inhomogeneous stochastic volatility models with Volterra Gaussian processes.
result Obtained large deviation principles for log-price processes in super rough Gaussian models.

New LAMB optimizer reduces BERT training time from 3 days to 76 minutes.

problem Training large deep neural networks on massive datasets is computationally challenging.
method Developed a new layerwise adaptive large batch optimization technique called LAMB.
result LAMB reduces BERT training time from 3 days to 76 minutes using very large batch sizes.

DAC improves associative classification for very large datasets with high scalability and quality.

problem Handling large datasets with many categorical features.
method DAC uses ensemble learning, parallel processing, and pruning techniques.
result DAC outperforms state-of-the-art solutions in prediction quality and execution time.

ESCHER avoids importance sampling to estimate regret in large games.

problem Estimating Nash equilibria in large games with high variance.
method Computes a history value function to estimate regret without importance sampling.
result ESCHER reduces regret estimation variance significantly compared to existing methods.

In this paper, we construct a family of asymptotically hyperbolic manifolds with horizons and with scalar curvature equal to -6. The manifolds we constructed can be arbitrary close to anti-de Sitter-Schwarzschild manifolds at infinity. Hence, the mass of our manifolds can be very large or very small. The main arguments…

2006-05-30abs ↗pdf ↗

LS-RPCA reduces dimensionality of large datasets better than random projections.

problem Reducing dimensionality of very large datasets for better classification performance.
method Developed LS-RPCA, an extension of RPCA for large datasets, to compare with random projections.
result LS-RPCA significantly improves classification performance over random projections.

A faster method for visualization recommendations on large datasets.

problem Infeasibility of state-of-the-art vis-rec models on large datasets due to high computational time.
method Reinforcement-learning (RL) framework that identifies optimal statistics within a time budget.
result Significantly reduces time-to-visualize with minimal error compared to baseline approaches.

Optimal student loan repayment strategies vary based on loan size.

problem Finding the most cost-effective repayment strategy for federal student loans.
method Analyzing the impact of different repayment strategies on total cost for varying loan sizes.
result Optimal repayment strategies depend on the loan balance, with different approaches for small, large, and intermediate balances.

This paper evaluates LLMs on large graph property estimation tasks.

problem Limited context length of LLMs limits their evaluation on large graphs.
method Developed EstGraph dataset and introduced four tasks for LLMs to estimate large graph properties.
result LLMs perform better on graph property estimation tasks when provided with context-rich prompts based on random walks.

Efficient methods for sparse random projections improve classification accuracy in very high-dimensional data.

problem Handling very high-dimensional sparse data efficiently.
method Non-iterative and iterative classification methods using sparse random projections and Jaccard kernel.
result Non-iterative methods yield larger, more accurate models than iterative methods.

New method reduces SBL complexity from cubic to linear, improving scalability.

problem Sparse Bayesian Learning's high computational complexity for large feature spaces.
method DQN-SBL, a diagonal Quasi-Newton method for SBL.
result DQN-SBL achieves competitive generalization with sparse models, scaling well to large-scale problems.

We show that the Gromov boundary of the free factor graph for the free group Fn with n>2 generators is the space of equivalence classes of minimal very small indecomposable projective Fn-trees without point stabilizer containing a free factor equipped with a quotient topology. Here two such trees are equivalent if the …

2012-11-07abs ↗pdf ↗

Performing signal processing tasks on compressive measurements of data has received great attention in recent years. In this paper, we extend previous work on compressive dictionary learning by showing that more general random projections may be used, including sparse ones. More precisely, we examine compressive K-mean…

2015-04-05abs ↗pdf ↗

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the effect sizes corresponding to the rejected hypotheses. For instance, this setting arises in the analysis of gene expression or DNA sequencing data. However, naive estimates of the effect sizes …

2014-05-16abs ↗pdf ↗

A new algorithm MBMF improves recommendation accuracy and speed for sparse datasets.

problem Sparse and fluctuating predictions in recommender systems.
method MBMF uses magnitude constraints and Spherical coordinates to optimize faster than existing methods.
result MBMF outperforms existing algorithms in accuracy and speed on synthetic and real datasets.

New insights into neural network training show some interpolating methods can generalize well, while others fail catastrophically.

problem Understanding why neural networks trained to interpolate can still generalize well or fail catastrophically.
method Analyzing empirical risk minimization (ERM) over large hypotheses classes, focusing on interpolating methods.
result Some interpolating ERM-like methods for large hypotheses classes provide good statistical guarantees, while others fail catastrophically.

This paper uses Bayesian ARD to automatically determine utility functions for discrete choice models.

problem Challenging and time-consuming task in identifying optimal utility function specifications.
method Bayesian framework and automatic relevance determination (ARD) for data-driven utility function specification.
result The proposed DCM-ARD model accurately recovers true utility function specifications and outperforms previous methods.

The paper proposes scalable methods for selecting prototypes from large dissimilarity datasets.

problem Selecting good prototypes from large dissimilarity datasets.
method Genetic algorithms, dissimilarity-based hashing, unsupervised and supervised criteria.
result The methods select good prototypes efficiently from large datasets.

The paper shows how to efficiently generate large Gaussian process samples with reliability guarantees.

problem Generating large-scale Gaussian process samples efficiently and with reliability.
method Demonstrates scaling data generation to large \(n\) while providing high probability guarantees.
result Efficiently generates large Gaussian process samples with reliability guarantees.

LVM and LVMB manifolds are a large family of examples of non kähler manifolds. For instance, Hopf manifolds and Calabi-Eckmann manifolds can be seen as LVMB manifolds. The LVM manifolds have a very natural action of the real torus and the quotient of this action is a simple polytope. This quotient allows us to relate c…

2010-06-09abs ↗pdf ↗

LVM and LVMB manifolds are a large family of examples of non kahler manifolds. For instance, Hopf manifolds and Calabi-Eckmann manifolds can be seen as LVMB manifolds. The LVM manifolds have a very natural action of the real torus and the quotient of this action is a simple polytope. This quotient allows us to relate c…

2010-06-09abs ↗pdf ↗

Paper develops machine learning methods to identify thermal models for HPC clusters.

problem Accurate thermal modeling for high-power HPC systems with diverse workloads.
method Advanced system identification algorithm combined with machine learning for data selection.
result Very accurate thermal models generated for HPC systems (average error < 1°C).