Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Feb 199419922001200920182026
48 results for probabilistic feature selection

Proposes PFCVMLP for feature selection and classification.

problem Performance degradation and low efficiency of traditional sparse Bayesian classifiers in high-dimensional data.
method Sparse Bayesian embedded feature selection method using truncated Gaussian distributions.
result PFCVMLP improves classification performance and feature selection effectiveness.

A new feature selection method for semi-supervised learning with imperfect labels.

problem Feature selection for semi-supervised learning with imperfectly labeled data.
method Genetic algorithm for proposing feature subsets, probabilistic error model for mislabeling, multi-class C-bound selection criterion.
result Empirical results show the effectiveness of the proposed framework compared to state-of-the-art approaches.

Paper proposes a framework for probabilistic load forecasting by integrating point forecasts.

problem Short-term load forecasting for power systems energy management.
method Two-stage framework: first stage for point forecasting, second stage for probabilistic forecasting using feature integration.
result Numerical results show effectiveness of the proposed approach in hour-ahead load forecasting.

New method enforces encoder sparsity in HPF for more interpretable feature selection.

problem Lack of encoder sparsity in HPF leads to lack of column-clustering property.
method Enforces encoder sparsity using a generalized additive model (GAM).
result Gains ability to perform feature selection and relates each representation to original features.

Proposes a new method for feature selection using Bayesian ID with intervention.

problem Feature selection in data with varying importance.
method Probabilistic model for interpolative decomposition with Bayesian inference and Gibbs sampling.
result The proposed Bayesian ID algorithm with intervention selects features with higher priority and comparable reconstructive errors.

The paper explores features from orderbooks to improve intraday electricity price forecasting.

problem Improving probabilistic forecasting of intraday electricity prices.
method Extracted 384 features from orderbooks, selected powerful features, and benchmarked models across two countries and product types.
result Revealed an asymmetric generalization phenomenon in electricity price forecasting models.

Deep learning uses layers of transformations to predict structured data with uncertainty.

problem Predicting structured high-dimensional data efficiently and with uncertainty.
method Applying layers of semi-affine input transformations to find features for probabilistic statistical methods.
result Achieves scalable prediction rules with uncertainty quantification and feature selection.

This paper investigates two feature-scoring criteria that make use of estimated class probabilities: one method proposed by \citet{shen} and a complementary approach proposed below. We develop a theoretical framework to analyze each criterion and show that both estimate the spread (across all values of a given feature)…

2012-06-27abs ↗pdf ↗

New probabilistic approaches offer recourse recommendations even when causal models are imperfect.

problem Limited causal knowledge makes guaranteeing algorithmic recourse impossible.
method Two probabilistic approaches: Bayesian model averaging and average effect computation.
result Probabilistic approaches lead to more reliable recourse recommendations.

New method provides calibrated feature importance explanations for regression models.

problem Lack of uncertainty quantification in existing local explanation methods.
method Extension of Calibrated Explanations method to support regression and probabilistic regression.
result Calibrated Explanations for regression provides quantified uncertainty and robust explanations.

The paper proposes a method to improve prediction accuracy by querying expert knowledge sequentially.

problem Prediction in high-dimensional settings with limited samples and costly expert consultation.
method Formulates knowledge elicitation as a probabilistic inference process, sequentially querying experts to improve predictions.
result The method shows improved prediction accuracy with minimal expert effort.

Proposes a new algorithm for Sparse Bayesian Learning connected to Stepwise Regression.

problem Sparse Bayesian Learning for probabilistic models.
method Coordinate ascent algorithm (RMP) for SBL, showing connection to Stepwise Regression.
result RMP's noise variance parameter limit connects to Stepwise Regression, with derived guarantees.

We present a unifying framework which reduces the construction of probabilistic component analysis techniques to a mere selection of the latent neighbourhood, thus providing an elegant and principled framework for creating novel component analysis models as well as constructing probabilistic equivalents of deterministi…

2013-03-13abs ↗pdf ↗

Bayesian TNKMs automatically infer model complexity and feature relevance.

problem Manual tuning of TN rank and feature dimensions is error-prone and computationally expensive.
method Bayesian approach with hierarchical priors on TN factors for automatic rank and feature selection.
result Superior performance in prediction accuracy, uncertainty quantification, interpretability, and scalability.

Probabilistic Boolean tensor decomposition improves accuracy and scalability.

problem Approximating multi-way binary data with interpretable low-rank factors.
method Scalable sampling-based posterior inference exploiting combinatorial structure.
result Maximum a posteriori decompositions outperform existing techniques.

The problem of joint feature selection across a group of related tasks has applications in many areas including biomedical informatics and computer vision. We consider the l2,1-norm regularized regression model for joint feature selection from multiple tasks, which can be derived in the probabilistic framework by assum…

2012-05-09abs ↗pdf ↗

Study proposes a hybrid method for medium-term load forecasting.

problem Accurate medium-term load forecasting for power system operation and planning.
method Support Vector Regression (SVR) combined with Symbiotic Organism Search Optimization (SOSO) for parameter optimization and feature selection.
result The proposed method outperformed previous methods in the EUNITE competition dataset.

A new probabilistic model detects communities in networks using both structure and node features.

problem Detecting communities in networks with node features for more accurate results.
method Generative probabilistic model considering network structure and node features.
result The model accurately detects communities and determines feature strength.

A method for selecting data for short-term load forecasting using Bayesian probabilistic models.

problem Lack of effective data selection techniques in power load forecasting.
method A fully automatic methodology based on a full Bayesian probabilistic model.
result The method improves the performance of load forecast models using real data.

This work improves understanding of dimension reduction algorithms and their probabilistic embeddings.

problem Improving theoretical understanding of non-linear dimension reduction algorithms.
method Analytical investigation of a generalized multidimensional scaling optimization problem.
result Probabilistic formulation of the problem leads to deterministic embeddings, contrary to standard implementations.

New scoring rules improve probabilistic classification model evaluation.

problem Traditional scoring rules misalign with the preference for correct classifications.
method Introduces Penalized Brier Score (PBS) and Penalized Logarithmic Loss (PLL) to modify proper scoring rules.
result PBS and PLL better identify optimal checkpoints and early stopping points, leading to superior F1 scores.

PSI models and infers feature attributions efficiently and accurately.

problem Modeling and inferring feature attributions in flexible predictive models.
method Probabilistic Shapley inference (PSI) framework using latent random variables and a masking-based neural network architecture.
result PSI learns feature attribution distributions centered at Shapley values, revealing meaningful uncertainty.

PliableBVS extends Bayesian lasso for modeling interactions with modifying variables.

problem Modeling interactions between large and small sets of variables, especially in omics studies.
method Bayesian variable selection with spike-and-slab priors and hierarchical structure.
result PliableBVS outperforms pliable lasso in identifying active main and interaction effects.

The paper automates machine learning pipelines using probabilistic matrix factorization and Bayesian optimization.

problem Automating the selection and tuning of machine learning pipelines.
method Combining collaborative filtering and Bayesian optimization with probabilistic matrix factorization.
result The approach quickly identifies high-performing pipelines across various datasets, significantly outperforming state-of-the-art methods.

Develops an online group feature selection method considering feature stream structure.

problem Online feature selection ignoring feature group structure.
method Formulates online group feature selection problem; develops OGFS method with intra-group and inter-group selection stages.
result Our method outperforms state-of-the-art methods in multiple tasks.

AEFS selects features from high-dimensional data using autoencoders.

problem Feature selection for high-dimensional data in computer vision and machine learning.
method Combines autoencoder regression and group lasso for unsupervised feature selection.
result AEFS selects more important features than traditional methods, including linear and nonlinear information.

Introduces greedy feature selection for classifier-dependent feature ranking.

problem Feature selection for classification tasks.
method Greedy feature selection, identifying the most important feature at each step based on the selected classifier.
result Theoretical and numerical benefits of greedy feature selection.

This work combines recurrent models with diffusion for probabilistic time series forecasting.

problem Scalability and capturing high-dimensional distributions and cross-feature dependencies in time series forecasting.
method Combines recurrent neural networks' efficiency with diffusion models' probabilistic modeling, using stochastic interpolants and conditional generation.
result Offers scalable probabilistic time series forecasting methods.

A new method selects optimal PHMM models for sequence alignment, improving accuracy.

problem Improving sequence alignment accuracy using PHMMs with optimal hidden states.
method Factorized Asymptotic Bayesian algorithm (FIC) for model selection.
result Improved alignment accuracy with more complex models than previous studies.

Paper proposes a novel unsupervised feature selection method using K-means and ADMM.

problem Finding a subset of features for high-dimensional unsupervised learning problems.
method Developed K-means Derived Unsupervised Feature Selection (K-means UFS) using ADMM to solve NP-hard optimization.
result K-means UFS outperforms baselines in feature selection for clustering.

GOLFS selects features for clustering by combining global and local information.

problem Feature selection for high-dimensional clustering without labels.
method Combines global and local information via manifold learning and regularized self-representation.
result Improves feature selection and clustering accuracy.

NGP selects N features from P using neural networks in a greedy, iterative process.

problem Feature selection for non-linear prediction problems.
method Neural Greedy Pursuit (NGP) algorithm, selecting features sequentially in an iterative loss minimization procedure.
result NGP provides better performance than DeepLIFT and Drop-one-out loss methods.