New 3D protein analysis methods improve accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Root cause analysis in a large-scale production environment is challenging due to the complexity of services running across global data centers. Due to the distributed nature of a large-scale system, the various hardware, software, and tooling logs are often maintained separately, making it difficult to review the logs…
Method analyzes large-scale network data to detect communication pattern shifts.
In recent years, ideas from statistics and scientific computing have begun to interact in increasingly sophisticated and fruitful ways with ideas from computer science and the theory of algorithms to aid in the development of improved worst-case algorithms that are useful for large-scale scientific and Internet data an…
This paper analyzes convergence of large-scale Transformers with weight decay.
Training deep neural networks using a large batch size has shown promising results and benefits many real-world applications. However, the optimizer converges slowly at early epochs and there is a gap between large-batch deep learning optimization heuristics and theoretical underpinnings. In this paper, we propose a no…
Study analyzes price response and spread impact in foreign exchange markets.
Kernel methods provide a principled way to perform non linear, nonparametric learning. They rely on solid functional analytic foundations and enjoy optimal statistical properties. However, at least in their basic form, they have limited applicability in large scale scenarios because of stringent computational requireme…
Many modern big data applications feature large scale in both numbers of responses and predictors. Better statistical efficiency and scientific insights can be enabled by understanding the large-scale response-predictor association network structures via layers of sparse latent factors ranked by importance. Yet sparsit…
This paper surveys large-scale machine learning methods for efficient data analysis.
This paper proposes a new method for learning covers of geometric datasets to improve topological inference and visualization.
Real time large scale streaming data pose major challenges to forecasting, in particular defying the presence of human experts to perform the corresponding analysis. We present here a class of models and methods used to develop an automated, scalable and versatile system for large scale forecasting oriented towards saf…
PePR scores assess DL model performance per resource unit, promoting smaller, more efficient models.
We analyze the Hessian spectra of large models up to 100B parameters.
The scaling properties of the time series of asset prices and trading volumes of stock markets are analysed. It is shown that similarly to the asset prices, the trading volume data obey multi-scaling length-distribution of low-variability periods. In the case of asset prices, such scaling behaviour can be used for risk…
The paper analyzes Indian stock sectors using multifractal analysis for long and short-term investment.
Based on the new type of random walk process called the Potentials of Unbalanced Complex Kinetics (PUCK) model, we theoretically show that the price diffusion in large scales is amplified 2/(2 + b) times, where b is the coefficient of quadratic term of the potential. In short time scales the price diffusion depends on …
New method speeds up analysis of computer experiments.
Paper presents a fast, private MH algorithm for large-scale Bayesian inference.
This paper tackles hyperparameter tuning for large-scale kernel ridge regression.
Large learning rates work surprisingly well in standard parameterization, contrary to theory.
The scale of functional magnetic resonance image data is rapidly increasing as large multi-subject datasets are becoming widely available and high-resolution scanners are adopted. The inherent low-dimensionality of the information in this data has led neuroscientists to consider factor analysis methods to extract and a…
dnamite simplifies NAMs for feature selection and survival analysis.
New method combines FMEA and Bayesian Network for root cause analysis in lithium-ion battery production.
Constructing an efficient parameterization of a large, noisy data set of points lying close to a smooth manifold in high dimension remains a fundamental problem. One approach consists in recovering a local parameterization using the local tangent plane. Principal component analysis (PCA) is often the tool of choice, as…
We propose a new two stage algorithm LING for large scale regression problems. LING has the same risk as the well known Ridge Regression under the fixed design setting and can be computed much faster. Our experiments have shown that LING performs well in terms of both prediction accuracy and computational efficiency co…
We investigate the large-volatility dynamics in financial markets, based on the minute-to-minute and daily data of the Chinese Indices and German DAX. The dynamic relaxation both before and after large volatilities is characterized by a power law, and the exponents usually vary with the strength of the large vo…
This work extends the scaling law to multiple and kernel regression, challenging traditional machine learning principles.
Analyzing deep neural networks (DNNs) via information plane (IP) theory has gained tremendous attention recently as a tool to gain insight into, among others, their generalization ability. However, it is by no means obvious how to estimate mutual information (MI) between each hidden layer and the input/desired output, …
Adaptive regularization prevents overfitting in large-scale sparse feature models.
We determine the critical batch size for large language models and find it scales with data size, not model size.
Inverse depth scaling found in LLMs due to similar layers averaging error.
Novel LRMC tackles missing data and outliers in large-scale low-rank data recovery.
We investigate the large-fluctuation dynamics in financial markets, based on the minute-to-minute and daily data of the Chinese Indices and German DAX. The dynamic relaxation both before and after the large fluctuations is characterized by a power law, and the exponents usually vary with the strength of the lar…
Tensor decomposition is a well-known tool for multiway data analysis. This work proposes using stochastic gradients for efficient generalized canonical polyadic (GCP) tensor decomposition of large-scale tensors. GCP tensor decomposition is a recently proposed version of tensor decomposition that allows for a variety of…
This paper enhances privacy-preserving randomized power method for large datasets.
This study analyzes quantization in deep learning models using statistical physics methods.
Water pollution is a major global environmental problem, and it poses a great environmental risk to public health and biological diversity. This work is motivated by assessing the potential environmental threat of coal mining through increased sulfate concentrations in river networks, which do not belong to any simple …
Gradient boosting decision tree (GBDT) is a widely-used machine learning algorithm in both data analytic competitions and real-world industrial applications. Further, driven by the rapid increase in data volume, efforts have been made to train GBDT in a distributed setting to support large-scale workloads. However, we …
Unified CCA methods for large-scale data with fast SGD algorithms.
This study examines cores within superclusters, highlighting their transitional nature and dynamical state.
SEMASIA provides a large dataset of latent representations for model comparison.
This work analyzes actor-critic methods for faster convergence.
New method speeds up learning of complex dynamical systems.
Theoretical analysis of data quality and synergies in LLMs.
The sum-of-correlations (SUMCOR) formulation of generalized canonical correlation analysis (GCCA) seeks highly correlated low-dimensional representations of different views via maximizing pairwise latent similarity of the views. SUMCOR is considered arguably the most natural extension of classical two-view CCA to the m…
Much recent research aims to identify evidence for Drug-Drug Interactions (DDI) and Adverse Drug reactions (ADR) from the biomedical scientific literature. In addition to this "Bibliome", the universe of social media provides a very promising source of large-scale data that can help identify DDI and ADR in ways that ha…
New method SF-AdamW trains large models without decay phases or memory overhead.