We solve a high-dimensional nonlinear model using Gaussian regressors.
problem Recovering a structured signal from high-dimensional data with a nonlinear link function.
method Proposes and analyzes an alternative convex recovery method that treats certain nonlinear link functions as linear in a lifted space.
result Our method successfully recovers the signal when previous methods fail due to a zero proportionality constant.
Gaussian processes improve system identification models.
problem Improving system identification models for non-linear dynamics.
method Using Gaussian processes to create time series prediction models.
result Gaussian processes enhance model accuracy in system identification.
Adaptive Gaussian kernel filtering with updated parameters.
problem Improving kernel adaptive filtering for better performance.
method Adaptive updating of Gaussian kernel parameters on an SPD manifold.
result Validation of the proposed method through experimental results.
Study improves flood loss risk models using historical data and rainfall data.
problem Predicting financial losses from flooding events.
method Used neural networks, decision trees, and kernel-based regressors on NFIP dataset, incorporating rainfall data.
result Extreme Gradient Boosting provided the best results, and bias correction improved model performance.
Active learning improves GP regression on complex, high-dimensional data.
problem Improving Gaussian Process regression in high-dimensional spaces with discontinuous functions.
method Combines manifold learning with active learning to optimize data selection and reduce dimensionality.
result Superior performance over random learning in synthetic data experiments.
Compact audio event descriptor from regressor confidence scores.
problem Efficiently representing audio events for classification.
method Class-specific regressors trained with random forests to estimate event onset and offset.
result Simple linear classifiers perform better than state-of-the-art on audio event classification.
We present the first tree-based regressor whose convergence rate depends only on the intrinsic dimension of the data, namely its Assouad dimension. The regressor uses the RPtree partitioning procedure, a simple randomized variant of k-d trees.
GaussDetect-LiNGAM eliminates Gaussianity tests for causal discovery.
problem Causal direction identification without Gaussianity assumptions.
method Leverages the equivalence between noise Gaussianity and residual independence in reverse regression.
result Gaussianity tests replaced with robust kernel-based independence tests.
Paper evaluates competence measures for DRS systems.
problem Choosing the best measure to quantify competence in DRS systems is challenging.
method Reviewed and adapted eight competence measures for regression problems, compared them on 15 datasets, and evaluated three DRS systems.
result DRS systems outperform individual regressors and static systems, but competence measure choice depends on the problem.
A new method aggregates predictions from decentralized learners using Gaussian copulas.
problem Learning from different data sets without sharing data.
method DELCO (Decentralized Ensemble Learning with COpulas) method using Gaussian copulas to aggregate predictions.
result Increased robustness and competitive accuracy in case of dependent classifiers.
This paper introduces Kernel-based Information Criterion (KIC) for model selection in regression analysis. The novel kernel-based complexity measure in KIC efficiently computes the interdependency between parameters of the model using a variable-wise variance and yields selection of better, more robust regressors. Expe…
SPACR trains uncertainty-aware regressors directly within a single pass, improving efficiency and validity.
problem Training uncertainty-aware regressors while maintaining efficiency and validity.
method Joint optimization of efficiency and validity during training.
result SPACR consistently provides tighter intervals and better coverage-efficiency trade-offs compared to standard CP and DOICR.
Safe active learning for multi-output Gaussian processes reduces data acquisition costs and ensures safety.
problem Expensive data acquisition and safety concerns in multi-output regression problems.
method Proposes a safe active learning approach considering data informativeness and safety constraints.
result Improved convergence compared to competitors on simulated and real-world datasets.
High-precision machine learning reduces particle physics simulations by orders of magnitude.
problem Reducing computational burden in particle physics simulations.
method Developed optimal training strategies and tuned machine learning regressors, including Deep Neural Networks with skip connections and boosted decision trees.
result Significantly reduced computational time by factors of 10^3 to 10^6 over first-principles simulations.
Paper develops a probabilistic regressor chain method using Monte Carlo methods.
problem Improving multi-output regression with probabilistic chains.
method Develops a sequential Monte Carlo scheme for probabilistic regressor chains.
result Monte Carlo scheme for probabilistic regressor chains can be effective and useful.
Error correction for correlated regressors using compressed sensing.
problem Measuring errors in regression tasks with sparse correlations.
method Compressed sensing applied to error signals, with ℓ 1 \ell_1 ℓ 1 -minimization for uniqueness. result Correct solution recovery in settings of error correction.
Study shows k k k -NN regressor consistency in complex survey designs.
problem Lack of consistency results for k k k -NN regressor in complex survey data. method Analysis of regularity conditions on sampling design and data distribution.
result Consistency of k k k -NN regressor under complex survey designs. New TSER algorithms outperform existing methods in time series extrinsic regression.
problem Improving time series extrinsic regression models.
method Extended TSER archive, introduced two new algorithms (FreshPRINCE and DrCIF), compared with rotation forest.
result DrCIF and FreshPRINCE models significantly outperform existing methods.
Improved kernel ridge regression using conjugate gradients.
problem Efficiently solving large-scale kernel ridge regression problems.
method Structured Gaussian regression model with low-rank approximation and conjugate gradients.
result Enhanced approximation of kernel ridge regressor/Gaussian process posterior mean.
The study evaluates nine machine learning regressors for predicting NASDAQ stock opening prices.
problem Predicting stock market opening prices for profitable trading strategies.
method Nine different machine learning regressors were applied to NASDAQ stock market data.
result The study found that certain regressors outperform others in predicting stock opening prices.
Improved predictive uncertainties in Gaussian Process regression.
problem Substantially underestimated uncertainties in GP predictive distributions.
method Two methods for scalable GP regression: variational inference for FITC and direct posterior predictive distribution.
result Significantly better calibrated uncertainties and higher log likelihoods.
DBKs enable scalable GPs with tractable inference for large datasets.
problem Scaling Gaussian processes to large and complex datasets while maintaining tractable inference.
method DBKs constructed from neural-network-parameterized basis functions with explicit low-rank structure, enabling linear-complexity inference.
result DBKs provide a unified perspective and improve predictive accuracy, uncertainty quantification, and computational efficiency.
Faster and versatile sampler for Bayesian linear regression.
problem Efficiently sampling from Bayesian linear regression models with arbitrary priors.
method Slice sampler exploiting linear regression likelihood structure.
result Better effective sample size per second than alternatives.
We introduce GP-FNARX: a new model for nonlinear system identification based on a nonlinear autoregressive exogenous model (NARX) with filtered regressors (F) where the nonlinear regression problem is tackled using sparse Gaussian processes (GP). We integrate data pre-processing with system identification into a fully …
Paper introduces ℓ \ell ℓ -DER for regression tasks using morphological operators and convex-concave procedure.
problem Developing a universal approximator for regression tasks.
method Introduces ℓ \ell ℓ -DER model, trains it using a convex-concave procedure (CCP) to minimize least-squares. result Outperforms other hybrid morphological models and state-of-the-art approaches.
Proposes a new regularizer for multi-task regression using Wasserstein geometry.
problem Lack of spatial information in multi-task regression models.
method Wasserstein geometry for multi-task regression, using unbalanced optimal transport.
result Improved statistical power and flexibility in multi-task regression models.
A robust algorithm for forecasting vector time series with seasonal components.
problem Forecasting vector time series with seasonal patterns and handling missing data.
method Auto-regression with seasonal annual, weekly, and daily baselines, and a Gaussian process for residuals. Custom truncated eigendecomposition and low-rank plus block-diagonal Gaussian kernel. Schur complement and Tikhonov regularization for efficient inference.
result The model can scale to very large datasets and is efficient in terms of memory and computation.
A new algorithm estimates NARMAX models with L1 regularization using coordinate descent.
problem Estimating NARMAX models with interpretability and error regressors.
method Cyclical coordinate descent for L1-regularized NARMAX models with error regressors.
result The method provides interpretable models with fewer important regressors.
A deep learning method reduces chemometric data size and improves analysis accuracy.
problem The curse of dimensionality in chemometric data analysis.
method L2 regularized sparse autoencoder with automatic node selection and Gaussian process regression.
result Significant improvement in regression accuracy compared to state-of-the-art methods.
Near-optimal algorithms for mean estimation and linear regression with Gaussian covariates and Huber contamination.
problem Gaussian mean estimation and linear regression with Gaussian covariates in the presence of Huber contamination.
method Near-optimal algorithms with optimal error guarantees, achieving sample complexity n = i l d e O ( d / ε 2 ) n = ilde{O}(d/ε^2) n = i l d e O ( d / ε 2 ) and almost linear runtime. result First sample near-optimal and almost linear-time algorithms with optimal error guarantees for both problems.
This work simplifies Gaussian process regression for multiple outputs.
problem Exponential computational complexity in Gaussian process regression.
method Approximating the covariance kernel using eigenvalues and functions.
result Significant reduction in training and regression complexity.
Proposes an ensemble loss function for robust regression.
problem Improving robustness of simple regression models in noisy environments.
method Ensemble techniques applied to a simple regressor with a half-quadratic learning algorithm.
result Significantly improves performance of simple regressors in noisy environments.
Two methods use DNN-HMM for global SNR estimation of speech signals.
problem Estimating global SNR of speech signals in various noise conditions.
method Dropout approximation for uncertainty estimation and noise-specific regressors.
result Improved SNR estimation accuracy compared to existing methods.
Two new methods improve neural network regressor calibration.
problem Improper calibration of neural network regressor predictions.
method Empirical calibration and temperature scaling.
result Both methods produce calibrated prediction intervals for neural network regressors.
Boosting ridge regression for high-dimensional data classification reduces computational cost and improves learning time.
problem High computational demand of inverting regularised covariance matrix in ridge regression for high-dimensional problems.
method Train an ensemble of ridge regressors in randomly projected subspaces, then combine them using adaptive boosting.
result Effective in terms of learning time and improved predictive performance in some cases.
Echo state network (ESN) is viewed as a temporal non-orthogonal expansion with pseudo-random parameters. Such expansions naturally give rise to regressors of various relevance to a teacher output. We illustrate that often only a certain amount of the generated echo-regressors effectively explain the variance of the tea…
A new approach simplifies multitask Gaussian processes without rank approximations.
problem Handling multioutput regression problems with conditionally dependent tasks.
method Introduces a novel approach to reduce multitask learning to univariate GPs, eliminating the need for rank approximations.
result Accurately recovers multitask covariance and noise matrices with fewer parameters, improving performance and reducing overfitting risk.
Research forecasts electricity spot prices using stochastic volatility models.
problem Forecasting day-ahead electricity prices in a spot market.
method Exploring and enriching a baseline stochastic volatility model with exogenous regressors.
result A better fitting model confirmed by out-of-sample forecasts.
Many models for sparse regression typically assume that the covariates are known completely, and without noise. Particularly in high-dimensional applications, this is often not the case. This paper develops efficient OMP-like algorithms to deal with precisely this setting. Our algorithms are as efficient as OMP, and im…
DTOR explains anomalies with rule-based explanations.
problem Need to explain anomalies in data effectively.
method Applies Decision Tree Regressor to estimate anomaly scores and generate rule-based explanations.
result DTOR produces robust and consistent rule-based explanations.
We study nonlinear regression of real valued data in an individual sequence manner, where we provide results that are guaranteed to hold without any statistical assumptions. We address the convergence and undertraining issues of conventional nonlinear regression methods and introduce an algorithm that elegantly mitigat…
Paper reduces sample complexity for bilinear systems identification to nearly constant.
problem Identifying discrete-time bilinear systems under bounded disturbances.
method Uses trajectory-dependent regressors and polynomial mean-square state growth analysis.
result Proves sample complexity of O ~ ( 1 / ε ) \widetilde{\mathcal O}(1/ε) O ( 1/ ε ) for estimation error ε ε ε . Novel algorithm identifies nonlinear Granger causal relationships using kernel ridge regression.
problem Identification of nonlinear Granger causal relationships.
method Flexible plug-in architecture with kernel ridge regression using radial basis function.
result Kernel ridge regression in mlcausality achieves competitive AUC scores and more finely calibrated p-values.
Proposes a new test for validating multivariate dynamic regression models.
problem Inadequate exogeneity conditions for conventional model specification tests in dynamic systems.
method Develops a generalized Durbin estimator for multiple-equation systems with dynamic dependencies, and constructs Wald tests.
result Bootstrap-based Wald tests improve finite-sample size control and validate the null hypothesis in multifactor models.
New algorithm tackles self-selection bias in estimating linear regressors.
problem Estimating k k k linear regressors with self-selection bias in d d d dimensions. method First local convergence algorithm for self-selection, reducing to coarsening problem.
result Improves running time of previous algorithms by a poly(d, k, 1/ε) factor.
The paper explores the tradeoff between fairness and accuracy in regression models.
problem Characterizing the tradeoff between fairness and accuracy in regression models.
method Provided a lower bound on the error of any fair regressor and extended the result to joint error using Wasserstein distance.
result Lower bounds on the error of fair regressors and their connection to Wasserstein distance.
This study examines the relationship between PLS and OLS regression using eigenvalue distributions.
problem Analyzing the difference between PLS and OLS regression in terms of eigenvalue distributions.
method Examined the distance between PLS and OLS regression coefficients using the Mahalanobis distance and eigenvalue distributions of the regressor covariance matrix.
result Provided a bound on the distance between PLS and OLS regression coefficients that depends only on the eigenvalue distribution of the regressor covariance matrix.
Meta-learning improves hyperparameter tuning for XGBoost.
problem Improving hyperparameter tuning for XGBoost models.
method Proposed MeSH algorithm using meta-regressors to guide hyperparameter selection.
result MeSH often finds superior hyperparameter configurations compared to SH and random search.