The paper develops predictors for functional data on manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We analyze the (unconditional) distribution of a linear predictor that is constructed after a data-driven model selection step in a linear regression model. First, we derive the exact finite-sample cumulative distribution function (cdf) of the linear predictor, and a simple approximation to this (complicated) cdf. We t…
We study online linear regression problems in a distributed setting, where the data is spread over a network. In each round, each network node proposes a linear predictor, with the objective of fitting the \emph{network-wide} data. It then updates its predictor for the next round according to the received local feedbac…
New method converts LVAs into linear projections for better understanding of complex models.
Recent results in the literature indicate that a residual network (ResNet) composed of a single residual block outperforms linear predictors, in the sense that all local minima in its optimization landscape are at least as good as the best linear predictor. However, these results are limited to a single residual block …
In this work, we study the problem of aggregating a finite number of predictors for nonstationary sub-linear processes. We provide oracle inequalities relying essentially on three ingredients: (1) a uniform bound of the norm of the time varying sub-linear coefficients, (2) a Lipschitz assumption on the predict…
New method selects sparse predictors in large LMMs.
New bounds show linear predictors rarely overfit with certain optimization methods.
A predictor that is deployed in a live production system may perturb the features it uses to make predictions. Such a feedback loop can occur, for example, when a model that predicts a certain type of behavior ends up causing the behavior it predicts, thus creating a self-fulfilling prophecy. In this paper we analyze p…
PARC uses piecewise linear predictors for regression and classification.
A residual network (or ResNet) is a standard deep neural net architecture, with state-of-the-art performance across numerous applications. The main premise of ResNets is that they allow the training of each layer to focus on fitting just the residual of the previous layer's output and the target output. Thus, we should…
New sample complexity bounds for linear predictors and neural networks, focusing on initialization.
This paper proposes a method to reduce complexity in GLMs with categorical predictors.
LESS combines local predictors for subsets to learn from heterogeneous input-output pairs.
We provide a new approach to training neural models to exhibit transparency in a well-defined, functional manner. Our approach naturally operates over structured data and tailors the predictor, functionally, towards a chosen family of (local) witnesses. The estimation problem is setup as a co-operative game between an …
Method constructs confidence regions for linear models with arbitrary predictors.
LIT-LVM improves linear predictors by estimating interaction terms with latent vectors.
Linear properties are either universal or absent across language models.
Optimizing proper loss yields calibrated models under specific conditions.
Paper addresses linear regression with partially mismatched data using local search with theoretical guarantees.
Overparameterized MLR fits hyper-curves, improving model robustness.
BART and MOTR-BART improve tree-based predictions with local linear models.
Predict covariance from features using convex optimization.
Study on merging predictors in causal and anticausal directions using CMAXENT.
In machine learning and data mining, linear models have been widely used to model the response as parametric linear functions of the predictors. To relax such stringent assumptions made by parametric linear models, additive models consider the response to be a summation of unknown transformations applied on the predict…
Novel strategy for federated learning with privacy-preserving predictors and nonvacuous generalization bounds.
Least squares estimator fails to achieve optimal risk in bounded distributions, but non-linear predictors can.
The study finds a trade-off between model size, test loss, and training loss for linear predictors.
Random imputation is surprisingly effective for linear predictors in missing data scenarios.
A new Bayesian approach to linear system identification has been proposed in a series of recent papers. The main idea is to frame linear system identification as predictor estimation in an infinite dimensional space, with the aid of regularization/Bayesian techniques. This approach guarantees the identification of stab…
We present a method to stop the evaluation of a prediction process when the result of the full evaluation is obvious. This trait is highly desirable in prediction tasks where a predictor evaluates all its features for every example in large datasets. We observe that some examples are easier to classify than others, a p…
Improved classification model for high-cardinality categorical predictors.
Study tail risk in high-frequency finance using -regularized regression.
This study compares various superlearner and deep learning architectures (machine-learning-based and neural-network-based) for classification problems across several simulated and industrial datasets to assess performance and computational efficiency, as both methods have nice theoretical convergence properties. Superl…
The problem of forecasting conditional probabilities of the next event given the past is considered in a general probabilistic setting. Given an arbitrary (large, uncountable) set C of predictors, we would like to construct a single predictor that performs asymptotically as well as the best predictor in C, on any data.…
In this short note, we provide a sample complexity lower bound for learning linear predictors with respect to the squared loss. Our focus is on an agnostic setting, where no assumptions are made on the data distribution. This contrasts with standard results in the literature, which either make distributional assumption…
Study examines how body segments respond to random vibrations.
New insights into when benign overfitting occurs in linear and classification tasks.
Neural Local Wasserstein Regression models distribution-on-distribution regression with flexible, localized transport maps.
Study fairness in ordinal regression using threshold models.
Generalized Linear Models (GLMs) and Single Index Models (SIMs) provide powerful generalizations of linear regression, where the target variable is assumed to be a (possibly unknown) 1-dimensional function of a linear predictor. In general, these problems entail non-convex estimation procedures, and, in practice, itera…
New probabilistic complexity measures for linear and kernel methods.
Ensemble methods that average over a collection of independent predictors that are each limited to a subsampling of both the examples and features of the training data command a significant presence in machine learning, such as the ever-popular random forest, yet the nature of the subsampling effect, particularly of th…
We consider building predictors when the data have missing values. We study the seemingly-simple case where the target to predict is a linear function of the fully-observed data and we show that, in the presence of missing values, the optimal predictor may not be linear. In the particular Gaussian case, it can be writt…
Deep neural networks improve AFT model for non-linear predictors.
We analyze the local Rademacher complexity of empirical risk minimization (ERM)-based multi-label learning algorithms, and in doing so propose a new algorithm for multi-label learning. Rather than using the trace norm to regularize the multi-label predictor, we instead minimize the tail sum of the singular values of th…
This paper provides estimation and inference methods for the best linear predictor (approximation) of a structural function, such as conditional average structural and treatment effects, and structural derivatives, based on modern machine learning (ML) tools. We represent this structural function as a conditional expec…
Functional PLS improves prediction and inference for scalar responses from functional predictors.