New method generates diverse EHR data types while maintaining privacy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Tree ensemble kernels improve Bayesian optimization for mixed features and constraints.
Paper explores how feature interactions improve XGBoost models.
FedCONST adapts update magnitudes to enhance feature generalization in FL.
The paper introduces logic constraints to improve AI model interpretability.
Distributed sensors compress and send features to a fusion center for linear regression.
Automates model selection for GLMs using optimization.
A new model improves uncertainty estimation in deep learning.
This work interprets SFA through variational inference, relaxing linearity constraints.
This paper proposes an active metric learning method for clustering with pairwise constraints.
ParamBoost uses gradient boosting to create interpretable non-linear models with constraints.
A new metric learning scheme for structured data combining graph and feature-space information.
Unsupervised domain adaptation studies the problem of utilizing a relevant source domain with abundant labels to build predictive modeling for an unannotated target domain. Recent work observe that the popular adversarial approach of learning domain-invariant features is insufficient to achieve desirable target domain …
Most learning methods with rank or sparsity constraints use convex relaxations, which lead to optimization with the nuclear norm or the -norm. However, several important learning applications cannot benefit from this approach as they feature these convex norms as constraints in addition to the non-convex rank a…
Proposes a semi-supervised K-Means algorithm for better feature selection.
We demonstrate a new deep learning autoencoder network, trained by a nonnegativity constraint algorithm (NCAE), that learns features which show part-based representation of data. The learning algorithm is based on constraining negative weights. The performance of the algorithm is assessed based on decomposing data into…
Multi-omic data provides multiple views of the same patients. Integrative analysis of multi-omic data is crucial to elucidate the molecular underpinning of disease etiology. However, multi-omic data has the "big p, small N" problem (the number of features is large, but the number of samples is small), it is challenging…
One primary focus in multimodal feature extraction is to find the representations of individual modalities that are maximally correlated. As a well-known measure of dependence, the Hirschfeld-Gebelein-Rényi (HGR) maximal correlation becomes an appealing objective because of its operational meaning and desirable propert…
One question central to Reinforcement Learning is how to learn a feature representation that supports algorithm scaling and re-use of learned information from different tasks. Successor Features approach this problem by learning a feature representation that satisfies a temporal constraint. We present an implementation…
The high dimensionality of hyperspectral images often results in the degradation of clustering performance. Due to the powerful ability of deep feature extraction and non-linear feature representation, the clustering algorithm based on deep learning has become a hot research topic in the field of hyperspectral remote s…
Study optimal consumption and portfolio strategies with no-borrowing constraint in financial markets.
Efficient learning of minimax risk classifiers in high dimensions.
We propose to prune a random forest (RF) for resource-constrained prediction. We first construct a RF and then prune it to optimize expected feature cost & accuracy. We pose pruning RFs as a novel 0-1 integer program with linear constraints that encourages feature re-use. We establish total unimodularity of the constra…
Critical incident stages identification and reasonable prediction of traffic incident duration are essential in traffic incident management. In this paper, we propose a traffic incident duration prediction model that simultaneously predicts the impact of the traffic incidents and identifies the critical groups of tempo…
Graph Attention Networks (GATs) are the state-of-the-art neural architecture for representation learning with graphs. GATs learn attention functions that assign weights to nodes so that different nodes have different influences in the feature aggregation steps. In practice, however, induced attention functions are pron…
New RL algorithm achieves sublinear regret and constraint violation without simulators.
Paper proposes a method to extract style features from unlabeled data.
We propose a general method for deformation quantization of any second-class constrained system on a symplectic manifold. The constraints determining an arbitrary constraint surface are in general defined only locally and can be components of a section of a non-trivial vector bundle over the phase-space manifold. The c…
Simplifies neural network constraints with computationally efficient method.
Many machine learning approaches are characterized by information constraints on how they interact with the training data. These include memory and sequential access constraints (e.g. fast first-order methods to solve stochastic optimization problems); communication constraints (e.g. distributed learning); partial acce…
COMET learns monotonic neural networks by incorporating counterexamples.
Decomposes bias in linear models under demographic parity constraints.
In this paper, we propose a Ward-like hierarchical clustering algorithm including spatial/geographical constraints. Two dissimilarity matrices and are inputted, along with a mixing parameter . The dissimilarities can be non-Euclidean and the weights of the observations can be non-uniform. The fi…
This paper describes a new approach, based on linear programming, for computing nonnegative matrix factorizations (NMFs). The key idea is a data-driven model for the factorization where the most salient features in the data are used to express the remaining features. More precisely, given a data matrix X, the algorithm…
When applied to high-dimensional datasets, feature selection algorithms might still leave dozens of irrelevant variables in the dataset. Therefore, even after feature selection has been applied, classifiers must be prepared to the presence of irrelevant variables. This paper investigates a new training method called Co…
Two new methods solve large-scale stochastic convex problems with linear constraints.
Sharp inequalities in unit ball with constraints on moments.
Proposes a method to enforce fairness in machine learning models without sensitive data.
We propose a mixed integer programming (MIP) model and iterative algorithms based on topological orders to solve optimization problems with acyclic constraints on a directed graph. The proposed MIP model has a significantly lower number of constraints compared to popular MIP models based on cycle elimination constraint…
Semantic segmentation is an established while rapidly evolving field in medical imaging. In this paper we focus on the segmentation of brain Magnetic Resonance Images (MRI) into cerebral structures using convolutional neural networks (CNN). CNNs achieve good performance by finding effective high dimensional image featu…
We propose learning flexible but interpretable functions that aggregate a variable-length set of permutation-invariant feature vectors to predict a label. We use a deep lattice network model so we can architect the model structure to enhance interpretability, and add monotonicity constraints between inputs-and-outputs.…
FairMixRep learns fair representations from mixed data types.
We consider the problem of metric learning subject to a set of constraints on relative-distance comparisons between the data items. Such constraints are meant to reflect side-information that is not expressed directly in the feature vectors of the data items. The relative-distance constraints used in this work are part…
Dynamic risk constraints help limit risky behavior in financial portfolios.
New algorithm solves minimax games with linear constraints.
Develops an algorithm for bilevel optimization with coupled constraints.
This paper explores the potential of Lagrangian duality for learning applications that feature complex constraints. Such constraints arise in many science and engineering domains, where the task amounts to learning optimization problems which must be solved repeatedly and include hard physical and operational constrain…
Framework optimizes model performance and interpretability for tabular data.