Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Nov 199319922001200920182026
48 results for massively distributed

DeepCMC compresses CSI for massive MIMO systems, reducing overhead and improving performance.

problem High CSI overhead in massive MIMO systems limits spectral efficiency.
method Deep learning-based fully convolutional neural network with residual layers and entropy coding.
result DeepCMC outperforms state-of-the-art schemes in CSI reconstruction quality for the same compression rate.

Improved predictions for rare labels using neural networks and ontologies.

problem Long-tailed frequency distribution in multi-label prediction problems.
method Modified neural network output layer with a Bayesian network of sigmoids leveraging ontology relationships.
result Significant improvements in per-label AUROC and average precision for less common labels.

This paper introduces a new framework for collective online learning of Gaussian processes in massive multi-agent systems.

problem The inefficiency of centralized communication in distributed machine learning systems.
method A novel Collective Online Learning Gaussian Process framework that allows each agent to build its local model and exchange it with others via peer-to-peer communication.
result Empirical results demonstrate the efficiency of the framework on both synthetic and real-world datasets.

We prove a Goldberg-Sachs theorem in dimension three. To be precise, given a three-dimensional Lorentzian manifold satisfying the topological massive gravity equations, we provide necessary and sufficient conditions on the tracefree Ricci tensor for the existence of a null line distribution whose orthogonal complement …

2015-02-01abs ↗pdf ↗

Paper proposes a GPU-based system for training massive deep learning models in ads systems.

problem Training massive deep learning models with terabyte-scale parameters in ads systems.
method Hierarchical GPU parameter server with 3-layer storage (GPU High-Bandwidth Memory, CPU main memory, SSD).
result 4-node hierarchical GPU parameter server trains a model 2X faster than a 150-node in-memory system.

A new protocol for private averaging protects data privacy in a crowd of users.

problem Protecting privacy in a crowd of users sharing personal data.
method Massively distributed algorithm for private averaging with malicious adversaries.
result Privacy is preserved even with malicious users, and the algorithm can find arbitrary accuracy solutions.

DFRot improves LLMs by reducing outlier and massive activation effects.

problem Reducing outlier and massive activation effects in rotated LLMs.
method Weighted loss function and orthogonal Procrustes transforms for rotation matrix refinement.
result DFRot achieves dual free (Outlier-Free and Massive Activation-Free) with significant improvements in perplexity.

This paper optimizes subsampling for large datasets using Poisson distribution.

problem Efficiently subsample large datasets for quasi-likelihood estimation.
method Derives optimal Poisson subsampling probabilities and develops a distributed subsampling framework.
result Consistent and asymptotically normal estimators are obtained.

Stochastic Gradient Descent (SGD) has become one of the most popular optimization methods for training machine learning models on massive datasets. However, SGD suffers from two main drawbacks: (i) The noisy gradient updates have high variance, which slows down convergence as the iterates approach the optimum, and (ii)…

2015-12-09abs ↗pdf ↗

New framework improves fraud prediction with incremental data balancing for massive data streams.

problem Class imbalance problem in massive imbalanced data streams.
method Incremental data balancing framework using Racing Algorithm for automated balancing and Random Forest for classification.
result Better results than Batch mode on European Credit Card dataset.

SIMD operations boost Bayesian computations up to 6x faster.

problem Expensive Bayesian computations are computationally intensive and parallelizable.
method Demonstrated the utility of SIMD operations for Bayesian applications using standard libraries.
result Up to 6x improvement in floating point arithmetic performance.

In recent years, a rich variety of shrinkage priors have been proposed that have great promise in addressing massive regression problems. In general, these new priors can be expressed as scale mixtures of normals, but have more complex forms and better properties than traditional Cauchy and double exponential priors. W…

2011-07-25abs ↗pdf ↗

Distributed computing offers a high degree of flexibility to accommodate modern learning constraints and the ever increasing size of datasets involved in massive data issues. Drawing inspiration from the theory of distributed computation models developed in the context of gradient-type optimization algorithms, we prese…

2014-07-16abs ↗pdf ↗

Random orthogonalization improves FL in massive MIMO systems without CSI.

problem Efficient model aggregation in FL with minimal channel estimation overhead.
method Combining FL with massive MIMO's channel hardening and favorable propagation, random orthogonalization reduces channel estimation overhead.
result Achieves model aggregation without CSI, significantly reducing channel estimation overhead.

Deep learning optimizes user association in Massive MIMO networks.

problem Optimizing user cell association for maximum sum-rate in Massive MIMO networks.
method Training a deep neural network to learn optimal association rules based on user positions.
result The neural network achieves the same performance as traditional optimization methods with reduced computational complexity.

COMET is a single-pass MapReduce algorithm for learning on large-scale data. It builds multiple random forest ensembles on distributed blocks of data and merges them into a mega-ensemble. This approach is appropriate when learning from massive-scale data that is too large to fit on a single machine. To get the best acc…

2011-03-10abs ↗pdf ↗

Develops an online nonparametric classifier for massive data.

problem Challenges of batch kernel-based nonparametric classifiers in massive data.
method Online principle components analysis to reduce dimensionality, followed by stochastic approximation algorithm for real-time calculation.
result Online classifier provides the best trade-off between accuracy and computation cost.

Study evaluates topological contributions in massive SQCD on compact 4-manifolds.

problem Analyzing topological path integrals for massive SQCD with up to 3 massive hypermultiplets.
method Decouples hypermultiplets, evaluates massless limit, and merges singularities at Argyres-Douglas points. Uses mass expansions for P2\mathbb{P}^2 and K3K3.
result Physical partition functions match mathematical results on Segre numbers of instanton moduli spaces.

Unified framework for photon and massive particle hypersurfaces in stationary spacetimes.

problem Understanding photon and massive particle hypersurfaces in stationary spacetimes.
method Unified framework using Killing-invariant timelike hypersurfaces and associated Finsler structures.
result Conditions for a hypersurface to be a photon or massive particle hypersurface are established.

In this work we simulate null geodesics for the Bonnor massive dipole metric by implementing a symbolic-numerical algorithm in Sage and Python. This program is also capable of visualizing in 3D, in principle, the geodesics for any given metric. Geodesics are launched from a common point, collectively forming a cone of …

2015-04-19abs ↗pdf ↗

The paper studies a method to sample nodes from a massive graph using personalized PageRank.

problem Sampling from a massive network is expensive and impractical; the paper provides an alternative.
method The paper introduces a crawling method to approximate the personalized PageRank vector without querying the entire graph.
result The adjusted personalized PageRank vector can effectively select nodes within the same block as the seed node.

Introduces a massive variant of Ray-Singer Torsion to avoid zero modes in topological field theories.

problem Avoiding zero modes in the evaluation of path integrals for topological field theories.
method Introduces a massive variant of the Ray-Singer Torsion, involving determinants of the twisted Laplacian with mass but without zero modes.
result Explicitly evaluates the massive Ray-Singer Torsion on product manifolds and mapping tori.

QEM uses parallel importance weighting for fast approximate Bayesian inference.

problem Bayesian inference challenges in large models with many observations and latent variables.
method Expectation Maximization (EM) with massively parallel importance weighting.
result QEM is faster and more scalable than RWS and VI.

We consider the problem of learning classifiers for labeled data that has been distributed across several nodes. Our goal is to find a single classifier, with small approximation error, across all datasets while minimizing the communication between nodes. This setting models real-world communication bottlenecks in the …

2012-02-27abs ↗pdf ↗

Deep learning reduces noise in weak lensing mass maps using GANs.

problem Noise reduction in weak lensing mass maps.
method Generative adversarial networks (GANs) applied to Subaru Hyper Suprime-Cam data.
result GANs successfully reproduce non-Gaussian information in denoised maps, showing stronger cosmological dependence.

The paper addresses frequency-dependent distortions in massive MIMO systems and proposes a method to recover covariance matrices.

problem Frequency-dependent distortions in the covariance matrix of massive MIMO systems.
method Proposes a novel UL-DL covariance interpolation technique under a mild reciprocity condition.
result The proposed method can recover the covariance matrix in the DL from an estimate in the UL, especially in FDD massive MIMO systems.

Study improves scalability of cell-free massive MIMO networks by optimizing UE-AP association.

problem Optimizing UE-AP association in cell-free massive MIMO networks.
method Deep learning algorithm using Bidirectional Long Short-Term Memory cells and hybrid probabilistic weight updating.
result Enhanced scalability without retraining, robust against pilot contamination.