Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

146292437583 · Jun 202019922001200920182026
48 results for scalable computing

The study identifies influential bioinformatics algorithms for scalable computing.

problem Data deluge in bioinformatics and need for scalable computing solutions.
method Identifying and analyzing influential data mining and machine learning algorithms.
result Guiding scalable computing experts to focus on specific bioinformatics algorithms.

New scalable algorithm improves linear regression accuracy and efficiency.

problem Improving linear regression accuracy and efficiency for large datasets.
method Developed new theory to model significance and multicollinearity as constraints, scaling with nn in the 10,000s.
result Significantly improves accuracy, reduces false detection rate, and speeds up computation.

New method for scalable barycenter computation using Wasserstein gradient flows.

problem Scalability and integration of label information in barycenter computation.
method Gradient flows in Wasserstein space, time discretization, mini-batch optimal transport, modular regularization, task-aware functions, supervised information integration.
result Empirically validated new state-of-the-art barycenter solver with labeled barycenters outperforming unlabeled ones.

Scalable3-BO tackles scalability issues in Bayesian optimization for big data and high dimensions.

problem Bayesian optimization scalability issues in big data and high dimensions.
method Sparse Gaussian process, random embedding, asynchronous parallelization.
result Scalable3-BO framework optimizes high-dimensional problems with 1 million data points and 10,000 dimensions.

Unified framework for scalable black-box optimization.

problem Expensive black-box evaluations in scientific and engineering domains.
method Integrates active learning, multi-armed bandits, and distributed computing.
result Consistently outperforms state-of-the-art black-box optimizers.

This dissertation advances scalable Gaussian processes using iterative methods and pathwise conditioning.

problem The classical Gaussian process formulation is not scalable for large datasets and modern hardware.
method Combining iterative methods and pathwise conditioning to improve scalability.
result Significantly reduced memory requirements and facilitated application to larger datasets.

NodeSig efficiently computes binary node embeddings for scalable graph analysis.

problem Scalability issues in graph representation learning models.
method NodeSig uses random walk diffusion probabilities and stable random projections to compute binary node embeddings efficiently.
result NodeSig achieves a good balance between accuracy and efficiency on node classification and link prediction tasks.

Meta-learning interpretable decision trees with synthetic data.

problem Lack of efficient, scalable methods for generating synthetic data for decision tree meta-learning.
method Synthetic generation of near-optimal decision trees using the MetaTree transformer architecture.
result Meta-learning of decision trees achieves performance comparable to real-world data or optimal decision trees, with significant computational cost reduction.

A scalable Gaussian process model using a mixture-of-experts approach.

problem Training Gaussian process models is computationally expensive and not scalable.
method A mixture-of-experts model with low-dimensional matrix inversions and importance sampling.
result The model offers comparable performance to Gaussian process regression at a lower computational cost.

This work improves scalability of Wasserstein distances in high dimensions.

problem Scalability issues in computing Wasserstein distances in high dimensions.
method Empirical convergence rates, robustness to data contamination, and computational methods.
result Established fast rates and robust estimation risks for sliced Wasserstein distances.

New method for scalable DPMM estimation in distributed data.

problem Efficiently handling new components in distributed Dirichlet Process Mixture Models.
method Locally creating new components, probabilistically consolidating them, and maintaining consistency with low communication cost.
result High scalability and consistent estimation in distributed and asynchronous environments.

Paper combines scalable BMF algorithms for web-scale datasets.

problem High computational cost of Bayesian Matrix Factorization.
method Combines Posterior Propagation and asynchronous distributed implementation.
result Substantial improvements in scalability on web-scale datasets.

Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…

2014-11-24abs ↗pdf ↗

Efficiently computes tree-Wasserstein barycenter for large-scale multilevel clustering and scalable Bayes.

problem Large-scale multilevel clustering and scalable Bayes problems.
method Proposes an efficient algorithm for tree-Wasserstein barycenter and variants.
result Significantly improves efficiency in computation and memory usage for large-scale applications.

Bayesian deep learning methods improve uncertainty estimation in computer vision.

problem Estimating uncertainty in deep learning models for robust computer vision.
method Comprehensive evaluation framework for scalable epistemic uncertainty estimation methods.
result Ensembling provides more reliable uncertainty estimates than MC-dropout.

GNet uses Gaussian processes for scalable, flexible neural networks.

problem Large-scale predictive modeling with high computational and storage costs.
method GNet employs Gaussian processes with nonparametric activation functions and a fast algorithm for training and predictions.
result GNet achieves competitive performance across various test problems, including nonlinear function prediction and real-world data regression.

GNet uses Gaussian processes for scalable, flexible neural networks.

problem Large-scale predictive modeling with high computational and storage costs.
method GNet employs Gaussian processes with nonparametric activation functions and a fast algorithm for efficient training and predictions.
result GNet achieves competitive performance across various test problems, including nonlinear function prediction and real-world data regression.

Estimates complex dependency structures in multi-omics data.

problem Graphical model estimation from multi-omics data with scalability and consistency.
method Pseudolikelihood-based graphical model framework with 1\ell_1-penalized empirical risk.
result Estimates partial correlation network from dual-omic liver cancer data.

This paper introduces a neural sampler for scalable sampling from complex distributions.

problem Efficiently sampling from high-dimensional un-normalized distributions.
method Neural implicit sampler trained with KL and Fisher divergence methods.
result The neural sampler generates large batches of samples with low computational costs.

A new method for scalable spectral clustering using random binning features.

problem Scalability issues in spectral clustering for large-scale problems.
method Random Binning features to accelerate similarity graph construction and eigendecomposition.
result Achieves similar accuracy to standard spectral clustering but with linear computational cost.

This work connects BNNs to GPs, providing scalable inference and identifying key properties.

problem Scaling and inference challenges in Bayesian neural networks.
method General convergence from BNNs to GPs, new covariance function, and scalable Nyström approximation.
result Established a scalable maximum a posterior (MAP) training and prediction procedure.

How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…

2013-09-30abs ↗pdf ↗

COPML framework securely trains models across multiple data owners without revealing individual data.

problem Privacy-preserving collaborative machine learning with multiple data owners.
method Securely encodes data, distributes computation, performs distributed training.
result Achieves up to 16x speedup in training time while maintaining strong privacy.

Scalable verifier for recurrent neural networks using polyhedral abstractions.

problem Certifying the correctness of recurrent neural networks.
method Combining sampling, optimization, and Fermat's theorem for polyhedral abstractions; gradient descent for refinement.
result Successfully verified challenging recurrent models in various domains.

New scalable methods for log determinant computations speed up Gaussian process kernel learning.

problem Prohibitive computational cost of log determinant calculations for Gaussian process kernel learning.
method Stochastic approximations based on Chebyshev, Lanczos, and surrogate models.
result Lanczos method is superior for kernel learning, and surrogate models are highly efficient and accurate.

A new framework SPOT efficiently solves large scale optimal transport problems.

problem Heavy computational burden in optimal transport limits its use.
method Implicit generative learning framework (SPOT) approximates optimal transport plan and solves it using stochastic gradient algorithms.
result SPOT efficiently solves optimal transport problems and can recover the density of the plan.

New scalable Lipschitz bounds improve neural network robustness analysis.

problem Computing tight Lipschitz bounds for deep neural networks is challenging and computationally expensive.
method Derived new closed-form Lipschitz bounds using more general feasible points of LipSDP, avoiding SDP solvers.
result Improved scalability and precision of Lipschitz estimation for large neural networks.

Improves scalability of Gaussian process for large datasets.

problem Scalability issues in Gaussian process for large datasets.
method Developed variational sparse inference algorithms (VSHGP, SVSHGP, DVSHGP) and distributed computing techniques.
result Demonstrated superior performance on various datasets compared to existing methods.

A scalable GPVAE method using local adjacencies to approximate GP inference.

problem Scalability issues in exact GP inference for large-scale GPVAEs.
method Neighbour-driven approximation strategy that confines computations to nearest neighbours.
result Outperforms other GPVAE variants in predictive performance and computational efficiency.

This work proposes a scalable framework for trusted multi-party computations using blockchain.

problem Ensuring trust in results from multi-agent computational experiments.
method Combining distributed validation and blockchain for immutable audits, reducing storage and communication costs.
result Guaranteed verifiability and validity of local computations in a scalable multi-agent environment.

Framework for applying GPs to real-world data with scalability guidelines.

problem Deployment of Gaussian Processes (GPs) is hindered by computational costs and lack of guidelines.
method Proposed a framework for identifying GP suitability and setting up robust models, formalizing decisions of experienced practitioners.
result More accurate results at test time for glacier elevation change case study.

SVB method provides scalable Bayesian proportional hazards model for high-dimensional gene expression data.

problem Bayesian methods for high-dimensional sparse survival data often sacrifice uncertainty quantification or computational scalability.
method Mean-field variational approximation for scalable Bayesian proportional hazards model.
result SVB method offers posterior distribution for parameters and variable selection via posterior inclusion probabilities.

New Krylov subspace methods speed up mixed-effects models with crossed random effects.

problem Slow computations for high-dimensional crossed random effects in mixed-effects models.
method Krylov subspace-based methods for generalized mixed-effects models with cross effects.
result Speedups by factors of up to 10,000 in computations for mixed-effects models.