Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1223 · Jul 201919922001200920182026
42 results for CUDA Fortran

GPU accelerates Bayesian inference of RSV model up to 17x faster.

problem Bayesian inference of realized stochastic volatility model.
method Hybrid Monte Carlo (HMC) algorithm parallelized on GPU (GTX 760) and CPU (Intel i7-4770 3.4GHz).
result GPU can achieve up to 17 times faster computation compared to CPU.

A parallel Fortran framework for neural networks and deep learning.

problem Developing efficient parallel Fortran for neural networks and deep learning.
method Simple interface, activation functions, stochastic gradient descent, Fortran 2018 collective subroutines, parallelism with derived types and collective operations.
result Ease of use and computational performance similar to existing machine learning frameworks, suitable for production.

CUDA optimized neural network predicts HbA1c from joint mobility and anthropometrics.

problem Early detection and accurate diagnosis of diabetes.
method Parallelized neural network using CUDA and C++ on Nvidia GPUs.
result Achieved high accuracy (95.65% on training, 86.67% on testing for males; 97.73% on training, 66.67% on testing for females).

ParaMonte simplifies Monte Carlo simulations for various scientific fields.

problem Efficiently performing Monte Carlo simulations for complex models.
method Unified, high-performance, parallelized library for C, C++, Fortran.
result Automates and streamlines Monte Carlo sampling for arbitrary-dimensional functions.

A quantum walk-based method for generating precise probability distributions efficiently.

problem Generating high-precision probability distributions for various applications.
method Integrates variational quantum circuits with split-step quantum walks to dynamically tune coin parameters and evolve quantum states.
result Achieves high simulation fidelity and reduces computational overhead compared to conventional methods.

New algorithms cluster high-dimensional polygonal curves efficiently.

problem Clustering high-dimensional polygonal curves with many vertices.
method Johnson-Lindenstrauss projection for polygonal curves, subsampling, probabilistic reduction of dependency on vertices.
result Achieves sublinear dependency on the number of input curves.

CUDA CTDR tackles unsupervised domain adaptation without domain alignment.

problem Lack of direct methods for unlabeled target domain classification.
method Jointly learns CTDR on source and target distributions using contradistinguish loss and supervised loss.
result CUDA CTDR achieves state-of-the-art results on various domain adaptation datasets.

Fast-vollib offers high-performance option pricing and IV computation.

problem Efficiently pricing and computing implied volatility for financial models.
method Open-source Python library with PyTorch, JAX, and CUDA backends, implementing Halley and LBR algorithms.
result High-performance option pricing and IV computation with vectorized implementations.

Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.

problem Inaccurate and slow Birkhoff projection in mHC implementations.
method Dual formulation, Newton's method, implicit differentiation, warp-level CUDA kernel.
result Substantial speedups and accuracy improvements in doubly stochastic projections.

KineticSim accelerates financial market simulations 3406x over CPU.

problem Simulating financial markets at scale with multi-agent models is bottlenecked by sequential processing and GPU kernel overhead.
method Formalized and implemented a reusable parallel design pattern for iterative multi-agent reductions in thread-block shared memory.
result Achieved a peak throughput of over 54.7 billion agent-events per second, delivering 3406x speedup over CPU.

KineticSim: A lightweight, high-performance execution engine for real-time market simulators

problem Simulating financial markets at scale with multi-agent models
method Reusable parallel design pattern: persistent, state-carrying clearing for iterative multi-agent reductions
result Reduces per-step critical-path depth from Theta(L+A) to Theta(log L + ceil(A/L))

Random forests handle categorical predictors natively but overlook 'absent levels' can bias models.

problem Bias in decision tree models due to 'absent levels' problem.
method Examined with Leo Breiman and Adele Cutler's random forests FORTRAN code and the randomForest R package.
result Simple heuristics can help mitigate the effects of the absent levels problem.

DiffTaichi enables fast, differentiable physical simulations with shorter code.

problem Building efficient differentiable physical simulators.
method Differentiable programming language (DiffTaichi) that generates gradients using source code transformations and a light-weight tape.
result Differentiable physical simulators written in DiffTaichi are faster and more concise than existing methods.

Deep-MacroFin uses neural networks to solve complex economic models efficiently.

problem Solving high-dimensional partial differential equations in continuous time economics.
method Leverages deep learning, specifically Multi-Layer Perceptrons and Kolmogorov-Arnold Networks, optimized with HJB equations.
result Offers a more efficient solution (5imes imes less memory, 40imes imes fewer FLOPs) for 50D economic models.

SPFlow simplifies SPN-based probabilistic learning with a Python library.

problem Creating and manipulating deep probabilistic models efficiently.
method Provides a Python library with DSL for SPN creation, efficient inference routines, and structure learning.
result SPFlow enables quick and efficient probabilistic inference and learning for SPNs.

The paper develops a new simulation technique for estimating conditional expectations in financial models.

problem Estimating conditional expectations in financial models with expensive simulation of endogenous variables.
method Introduces a hierarchical simulation scheme with oversimplified defaults to address variance issues.
result The hierarchical simulation technique significantly improves the success of neural net regression for conditional expectation estimation.

A new meta-learning method improves deep neural net training efficiency.

problem Efficient training of complex deep neural networks with long training processes.
method Meta-learning with Hessian-Free (MLHF) approach based on Hessian-Free optimization.
result MLHF shows good and continuous training performance in deep convolution neural nets.

Performance-aware channel pruning improves CNN on embedded GPUs.

problem Inefficient channel pruning on embedded GPUs leads to performance slowdowns.
method Evaluate higher-level libraries that analyze input characteristics for optimized code generation.
result Performance-aware pruning can achieve significant performance speedups, up to 10x.