This work explores variably scaled kernels to improve non-stationary Gaussian processes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This study presents a new lossy image compression method that utilizes the multi-scale features of natural images. Our model consists of two networks: multi-scale lossy autoencoder and parallel multi-scale lossless coder. The multi-scale lossy autoencoder extracts the multi-scale image features to quantized variables a…
New method infers causal factors from large-scale data without full graph reconstruction.
We develop a new statistical test for comparing variables with varying scales.
We propose a new method for input variable selection in nonlinear regression. The method is embedded into a kernel regression machine that can model general nonlinear functions, not being a priori limited to additive models. This is the first kernel-based variable selection method applicable to large datasets. It sides…
Identifying causal direction in location-scale noise models with hidden variables
Improves speaker verification for variable-duration utterances using a feature pyramid module.
Properties of low-variability periods in the time series are analysed. The theoretical approach is used to show the relationship between the multi-scaling of low-variability periods and multi-affinity of the time series. It is shown that this technically simple method is capable of reveling more details about time-seri…
The scaling properties of the time series of asset prices and trading volumes of stock markets are analysed. It is shown that similarly to the asset prices, the trading volume data obey multi-scaling length-distribution of low-variability periods. In the case of asset prices, such scaling behaviour can be used for risk…
In this paper we combine two important extensions of ordinary least squares regression: regularization and optimal scaling. Optimal scaling (sometimes also called optimal scoring) has originally been developed for categorical data, and the process finds quantifications for the categories that are optimal for the regres…
Study uses CNNs to upscale wind speed data from 100 km to 3 km, improving subgrid-scale variability.
New method improves uncertainty quantification in latent variable models.
Paper introduces slow kill for efficient large-scale variable screening.
Understanding the structure of financial markets deals with suitably determining the functional relation between financial variables. In this respect, important variables are the trading activity, defined here as the number of trades , the traded volume , the asset price , the squared volatility , the bid…
We propose a probabilistic graphical model realizing a minimal encoding of real variables dependencies based on possibly incomplete observation and an empirical cumulative distribution function per variable. The target application is a large scale partially observed system, like e.g. a traffic network, where a small pr…
Model detects epileptic seizures in EEG with high sensitivity.
Extracts coarse-grained PDEs from microscopic simulations.
Paper proposes data quality measures for large-scale high-dimensional data.
Generative model improves wind field downscaling from coarse climate models.
The authors study the method of scaling in the context of the study of automorphism groups of complex domains in multiple dimensions. Various types of scaling techniques are compared and contrasted. Applications are given in a number of areas of complex geometric analysis. Relations with other parts of mathematics are …
Paper shows noisy labels can improve PLR variable selection.
In this paper, we consider the classic measurement error regression scenario in which our independent, or design, variables are observed with several sources of additive noise. We will show that our motivating example's replicated measurements on both the design and dependent variables may be leveraged to enhance a spa…
We extend Kirman's model by introducing variable event time scale. The proposed flexible time scale is equivalent to the variable trading activity observed in financial markets. Stochastic version of the extended Kirman's agent based model is compared to the non-linear stochastic models of long-range memory in financia…
Fractal neural networks play SimCity and Conway's Game of Life on varying scales.
A new method treats all variables equally in fitting data.
CPI overcomes limitations of permutation importance by providing accurate variable selection.
When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large scale data applications like genomics, connectomics, and eco-informatics the dataset…
Neural population activity often exhibits rich variability and temporal structure. This variability is thought to arise from single-neuron stochasticity, neural dynamics on short time-scales, as well as from modulations of neural firing properties on long time-scales, often referred to as "non-stationarity". To better …
Collective phenomena with universal properties have been observed in many complex systems with a large number of components. Here we present a microscopic model of the emergence of scaling behavior in such systems, where the interaction dynamics between individual components is mediated by a global variable making the …
We develop a scale-invariant truncated Lévy (STL) process to describe physical systems characterized by correlated stochastic variables. The STL process exhibits Lévy stability for the probability density, and hence shows scaling properties (as observed in empirical data); it has the advantage that all moments are fini…
We introduce a few variants on Frank-Wolfe style algorithms suitable for large scale optimization. We show how to modify the standard Frank-Wolfe algorithm using stochastic gradients, approximate subproblem solutions, and sketched decision variables in order to scale to enormous problems while preserving (up to constan…
New method reduces memory usage for high-dimensional variable selection.
We propose a framework combining detrended fluctuation analysis with standard regression methodology. The method is built on detrended variances and covariances and it is designed to estimate regression parameters at different scales and under potential non-stationarity and power-law correlations. The former feature al…
Proposes a non-parametric method for deep discrete latent variable models.
New algorithm tackles big data Bayesian problems with latent variables.
We propose a variable decomposition algorithm -greedy block coordinate descent (GBCD)- in order to make dense Gaussian process regression practical for large scale problems. GBCD breaks a large scale optimization into a series of small sub-problems. The challenge in variable decomposition algorithms is the identificati…
A scalable method to learn causal graphs from large data.
Bayesian optimization speeds up bioprocess development across scales.
CRN learns causal models using neural networks, scaling with variables and leveraging prior knowledge.
Variable selection for Gaussian process models is often done using automatic relevance determination, which uses the inverse length-scale parameter of each input variable as a proxy for variable relevance. This implicitly determined relevance has several drawbacks that prevent the selection of optimal input variables i…
Study homogenizes equations on parallelizable manifolds using tensor localization and periodicity.
The article proposes modified Gower's coefficients for handling mixed type variables in nearest neighbor methods.
Productions functions map the inputs of a firm or a productive system onto its outputs. This article expounds generalizations of the production function that include state variables, organizational structures and increasing returns to scale. These extensions are needed in order to explain the regularities of the empiri…
New method identifies latent causal variables from observed data, overcoming indeterminacies.
Study counterfactuals in cyclic systems with shifts and scales.
With this study we want to test the validity of the well known "Verdoorn's Law" which considers the relationship between the growth of productivity and output in the case of the Portuguese economy at a regional and sectoral levels (NUTs II) for the period 1995-1999. The importance of some additional variables in the or…
A new method combines predictors and their lags using supervised PCA for dynamic forecasting.
Network analysis improves stock return forecasting.