Study predicts academic achievement using students' support networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes how the one-dimensional Wasserstein distance captures pointwise density differences in finite samples.
The paper improves support recovery in high-dimensional precision matrix estimation using meta learning.
Support Vector Machine (SVM) is an efficient classification approach, which finds a hyperplane to separate data from different classes. This hyperplane is determined by support vectors. In existing SVM formulations, the objective function uses L2 norm or L1 norm on slack variables. The number of support vectors is a me…
In this paper, we investigate a multivariate multi-response (MVMR) linear regression problem, which contains multiple linear regression models with differently distributed design matrices, and different regression and output vectors. The goal is to recover the support union of all regression vectors using -reg…
Meta-learning improves support recovery in high-dimensional PCA.
The paper analyzes kNN density estimation's convergence rates under different conditions.
Efficient algorithms for sparse parameter recovery in mixture models.
We consider the homogeneous and the non-homogeneous convex relaxations for combinatorial penalty functions defined on support sets. Our study identifies key differences in the tightness of the resulting relaxations through the notion of the lower combinatorial envelope of a set-function along with new necessary conditi…
Constructs a support-preserving homotopy for differential forms with boundary decay estimates.
Tuning SVM and boosting models using optimization algorithms.
Gromov-Wasserstein (GW) is a powerful tool to compare probability measures whose supports are in different metric spaces. GW suffers however from a computational drawback since it requires to solve a complex non-convex quadratic program. We consider in this work a specific family of cost metrics, namely \textit{tree me…
Neural networks perform differently when regression is treated as classification.
This paper presents novel Gaussian process decentralized data fusion algorithms exploiting the notion of agent-centric support sets for distributed cooperative perception of large-scale environmental phenomena. To overcome the limitations of scale in existing works, our proposed algorithms allow every mobile sensing ag…
Global Chern currents and Baum Bott currents defined on arbitrary complex manifolds.
The main goal of statistical learning theory is to provide a fundamental framework for the problem of decision making and model construction based on sets of data. Here, we present a brief introduction to the fundamentals of statistical learning theory, in particular the difference between empirical and structural risk…
Study examines how bank holding structures affect financial stress spread.
The support vector machine (SVM) is a widely used machine learning tool for classification based on statistical learning theory. Given a set of training data, the SVM finds a hyperplane that separates two different classes of data points by the largest distance. While the standard form of SVM uses L2-norm regularizatio…
Classification and regression tasks in overparameterized models show different generalization properties.
This paper examines how voter concentration affects election outcomes in district-based systems.
A smooth diffeomorphism is said to be distributionally uniquely ergodic (DUE for short) when it is uniquely ergodic and its unique invariant probability measure is the only invariant distribution (up to multiplication by a constant). Ergodic translations on tori are classical examples of DUE diffeomorphisms. In this ar…
SVM and linear regression models coincide in high dimensions.
Proposes a novel SVM model for binary classification with different misclassification costs.
k Nearest Neighbor (kNN) method is a simple and popular statistical method for classification and regression. For both classification and regression problems, existing works have shown that, if the distribution of the feature vector has bounded support and the probability density function is bounded away from zero in i…
The -support norm is a regularizer which has been successfully applied to sparse vector prediction problems. We show that it belongs to a general class of norms which can be formulated as a parameterized infimum over quadratics. We further extend the -support norm to matrices, and we observe that it is a special …
Solving different types of optimization models (including parameters fitting) for support vector machines on large-scale training data is often an expensive computational task. This paper proposes a multilevel algorithmic framework that scales efficiently to very large data sets. Instead of solving the whole training s…
The massive amount of available data potentially used to discover patters in machine learning is a challenge for kernel based algorithms with respect to runtime and storage capacities. Local approaches might help to relieve these issues. From a statistical point of view local approaches allow additionally to deal with …
New research challenges the idea that counterfactual explanations should be sparse.
Analyzes SVM classifier behavior with different parameters and data types.
For a learning task, data can usually be collected from different sources or be represented from multiple views. For example, laboratory results from different medical examinations are available for disease diagnosis, and each of them can only reflect the health state of a person from a particular aspect/view. Therefor…
The study examines representations of compactly supported diffeomorphisms with a positive energy condition.
Given a training set with binary classification, the Support Vector Machine identifies the hyperplane maximizing the margin between the two classes of training data. This general formulation is useful in that it can be applied without regard to variance differences between the classes. Ignoring these differences is not…
Tensor, a multi-dimensional data structure, has been exploited recently in the machine learning community. Traditional machine learning approaches are vector- or matrix-based, and cannot handle tensorial data directly. In this paper, we propose a tensor train (TT)-based kernel technique for the first time, and apply it…
Classifiers label data as belonging to one of a set of groups based on input features. It is challenging to obtain accurate classification performance when the feature distributions in the different classes are complex, with nonlinear, overlapping and intersecting supports. This is particularly true when training data …
Background and objectives: Parkinson's disease is a neurological disorder that affects the motor system producing lack of coordination, resting tremor, and rigidity. Impairments in handwriting are among the main symptoms of the disease. Handwriting analysis can help in supporting the diagnosis and in monitoring the pro…
Paper supports robust estimation in regression with heavy-tailed errors.
In this work, we take a closer look at the evaluation of two families of methods for enriching information from knowledge graphs: Link Prediction and Entity Alignment. In the current experimental setting, multiple different scores are employed to assess different aspects of model performance. We analyze the informative…
This work proposes a new method to match distributions across different spaces using cycle-consistent maps.
Study on recovering supports of multiple sparse vectors from mixed linear measurements.
DAT-CGAN improves time series generation for better decision support.
MetaPhysiCa tackles robust physics-informed machine learning for OOD tasks.
We define a generalized likelihood function based on uncertainty measures and show that maximizing such a likelihood function for different measures induces different types of classifiers. In the probabilistic framework, we obtain classifiers that optimize the cross-entropy function. In the possibilistic framework, we …
Deep Learning Library (DLL) is a new library for machine learning with deep neural networks that focuses on speed. It supports feed-forward neural networks such as fully-connected Artificial Neural Networks (ANNs) and Convolutional Neural Networks (CNNs). It also has very comprehensive support for Restricted Boltzmann …
Lapse-supported life insurance exacerbates adverse selection risks.
Study supports recovery of PDEs from noisy data using a specific regularization method.
In nonparametric classification and regression problems, regularized kernel methods, in particular support vector machines, attract much attention in theoretical and in applied statistics. In an abstract sense, regularized kernel methods (simply called SVMs here) can be seen as regularized M-estimators for a parameter …
It is known that for a certain class of single index models (SIMs) , support recovery is impossible when and a model complexity adjusted sample size is below a critical threshold. Recen…
Study semiclassical measures on complex hyperbolic quotients, identifying measure supports.