Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

65131196261 · Jun 202019922001200920172026
48 results for data-driven regression

A new method for support vector regression using a data-driven insensitive parameter.

problem Determining an optimal insensitive parameter in support vector regression.
method A data-driven approach to approximate the insensitive parameter by minimizing a generalized loss function based on the likelihood principle.
result The proposed method outperforms traditional support vector regression methods and has lower computational costs.

Improves robustness of high-dimensional regression with rank objective and group lasso regularization.

problem Heavy-tailed noise and outliers in high-dimensional regression.
method Non-smooth Wilcoxon score based rank objective, group lasso regularization, data-driven tuning rule, proximal augmented Lagrangian method.
result Robust estimator with finite-sample error bound and efficient computational method.

This work evaluates and benchmarks calibration metrics for data-driven regression models.

problem Conflicting results from different calibration metrics make it hard to compare and interpret model performance.
method Systematically extracted and benchmarked 14 regression calibration metrics across various data types and recalibration methods.
result Many metrics disagree on the same recalibration result, highlighting the need for careful metric selection.

Gaussian Process Regression and Kernel Ridge Regression are popular nonparametric regression approaches. Unfortunately, they suffer from high computational complexity rendering them inapplicable to the modern massive datasets. To that end a number of approximations have been suggested, some of them allowing for a distr…

2019-12-13abs ↗pdf ↗

We present a regression technique for data-driven problems based on polynomial chaos expansion (PCE). PCE is a popular technique in the field of uncertainty quantification (UQ), where it is typically used to replace a runnable but expensive computational model subject to random inputs with an inexpensive-to-evaluate po…

2018-08-09abs ↗pdf ↗

Develops a new method for uncertainty quantification in high-dimensional learning.

problem Challenges in uncertainty quantification in high-dimensional regression or learning problems.
method Data-driven approach for UQ that corrects bias terms from training data.
result Non-asymptotic confidence intervals that avoid overestimating uncertainty.

We improve random forest consistency and performance with DMRF, a new variant.

problem Improving the consistency and performance of random forest models.
method Developed DMRF, a data-driven multinomial random forest, by modifying proof methods and improving data utilization.
result DMRF achieves strong consistency with probability 1, surpassing previous models in classification tasks.

Framework improves data-driven ROMs for complex systems using Bayesian operator inference.

problem Improving the quality of data-driven reduced-order models for complex dynamical systems.
method Develops an active learning framework using Bayesian operator inference to identify and select training parameters.
result The proposed adaptive sampling strategy consistently yields more stable and accurate ROMs than random sampling.

The study explores how machine learning can enhance scientific research.

problem Improving scientific models with machine learning.
method Analysis of data-driven models versus manually added variables in regression.
result Complex models may not always improve over simpler ones in scientific contexts.

Local laGPR speeds up multiscale mechanics simulations without neural networks.

problem High computational costs in multiscale mechanics simulations.
method Local approximate Gaussian process regression (laGPR) combined with FE schemes.
result laGPR offers better accuracy than neural networks for stress predictions.

ADML combines debiased learning with data-driven model selection for efficient inference.

problem Debiased machine learning estimators can be unstable and biased in nonparametric models.
method Data-driven model selection techniques combined with debiased machine learning.
result ADML estimators yield superefficient inference for pathwise differentiable parameters.

Study integrates machine learning with SAA for optimizing decisions based on uncertain parameters and covariates.

problem Optimizing decisions under uncertain parameters and covariates.
method Data-driven frameworks integrating machine learning prediction models within SAA for scenario generation.
result Consistent and asymptotically optimal solutions under certain conditions, with finite sample guarantees.

Data-driven approaches, most prominently deep learning, have become powerful tools for prediction in many domains. A natural question to ask is whether data-driven methods could also be used to predict global weather patterns days in advance. First studies show promise but the lack of a common dataset and evaluation me…

2020-02-02abs ↗pdf ↗

The paper tackles data-driven optimal control of unknown nonlinear systems using RKHS.

problem Unknown nonlinear dynamics and stage cost functions.
method Embed state densities into RKHS, learn Markov operators, solve Hamilton-Jacobi-Bellman recursions.
result Solves a wide range of nonlinear control problems, including depth regulation.

Method estimates treatment effects with continuous values, correcting for confounding.

problem Estimating treatment effects with continuous values, dealing with confounding.
method Two-stage kernel ridge regression: first stage learns response, second stage corrects for distribution shift.
result Optimal learning bounds achieved without estimating treatment density, adapts to unknown overlap and kernel spectral decay.

New algorithm tackles regression on manifold data using diffusion and semi-supervised learning.

problem Regression on high-dimensional manifold data with complex structures.
method Diffusion-based spectral algorithm using graph Laplacian and heat kernel.
result Algorithm achieves convergence rate dependent on intrinsic manifold dimension, avoiding curse of dimensionality.

This paper tackles the problem of selecting among several linear estimators in non-parametric regression; this includes model selection for linear regression, the choice of a regularization parameter in kernel ridge regression, spline smoothing or locally weighted regression, and the choice of a kernel in multiple kern…

2009-09-10abs ↗pdf ↗

Framework learns physics-informed continuum models from molecular data.

problem Discovering accurate and robust data-driven continuum models from molecular simulation data.
method Operator regression framework using neural networks in modal space with physical inductive biases.
result Learned operators generalize to unseen system characteristics.

Automated digital twin discovery from biological data improves drug discovery and personalized medicine.

problem Developing reliable digital twins from noisy, incomplete biological data.
method Symbolic and sparse regression, Bayesian frameworks, deep learning, and large language models.
result Sparse regression generally outperforms symbolic regression, especially with Bayesian frameworks.

Self-test loss functions improve data-driven modeling of weak-form operators and gradient flows.

problem Challenges in selecting test functions for data-driven modeling involving weak-form operators and gradient flows.
method Introducing self-test loss functions that depend on unknown parameters and are quadratic.
result Self-test loss functions conserve energy for gradient flows and coincide with log-likelihood ratios for stochastic differential equations.

Study compares data-driven vs model-based MRS quantification strategies, focusing on resilience to out-of-distribution effects.

problem Resilience to out-of-distribution effects in data-driven MRS quantification.
method Compared three data-driven strategies (supervised regression, self-supervised learning, test-time adaptation) against model-based fitting tools.
result Test-time adaptation proved most resilient to out-of-distribution effects, while self-supervised learning achieved intermediate performance.

A new line search rule improves support recovery in high-dimensional data.

problem Support recovery in high-dimensional data analysis with 0\ell_0 penalty.
method Data-driven line search rule for adaptive step size determination.
result Proves 2\ell_2 error bound without restrictions on cost functional.

This work improves SINDy-type algorithms for system identification using score-guided dictionary selection.

problem Improving accuracy and interpretability in dynamical system identification.
method Score-guided library selection to refine dictionary terms in sparse regression.
result Score-guided methods enhance SINDy's robustness in discovering governing equations.

Random Forest kernels improve performance in various regression and survival tasks.

problem Improving performance of Random Forest in high-dimensional data with noisy features.
method Developed and evaluated data-driven RF kernels for regression, classification, and survival tasks.
result RF kernels are competitive or superior to RF in most scenarios, especially for survival tasks.

Discovering quasipotential equations from data using machine learning.

problem Understanding escape mechanisms from metastable states in nonlinear systems.
method Combining neural networks and sparse regression to symbolically reconstruct quasipotential equations.
result Model-unbiased analytical forms of quasipotential discovered directly from data.

Bayesian nonparametrics improves data-driven risk optimization under distributional uncertainty.

problem Improving out-of-sample performance in machine learning models due to distributional uncertainty.
method Combining Bayesian nonparametric theory and decision-theoretic preferences to propose a robust optimization criterion.
result The proposed robust optimization procedure provides favorable statistical guarantees and tractable approximations.