Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

106213319425 · Jun 202019922001200920182026
48 results for feature automation

Python library automates feature engineering and selection for linear models.

problem Difficulties in training and explaining complex machine learning models.
method Automated feature engineering and selection for linear models.
result Improves prediction accuracy of linear models while retaining interpretability.

Automated machine learning simplifies model selection and tuning.

problem Manual tuning of machine learning models by data scientists is time-consuming and requires extensive expertise.
method Review of AutoML techniques including automated feature engineering, model learning, and deep learning.
result Current AutoML techniques can significantly reduce the burden of manual tuning.

Automates feature extraction from JSON data for machine learning.

problem Manual feature engineering for JSON data is laborious, lossy, and prone to bias.
method Automates feature extraction using Mill.jl and JsonGrinder.jl.
result Creates a differentiable machine learning model from raw JSON samples.

Automates feature engineering for predictive models using reinforcement learning.

problem Lack of a well-defined basis for effective feature engineering.
method Performance-driven exploration of a transformation graph using reinforcement learning.
result Automated feature engineering reduces human intervention and costs.

Study automates feature selection and clustering for HFT stock price forecasting.

problem Manual feature selection and clustering for high-frequency trading (HFT) stock price forecasting.
method Dual competitive feature importance mechanism and clustering via shallow neural network topology.
result Enhanced forecasting ability of the RBFNN regressor through automated feature selection and clustering.

AutoFS combines trainers to improve feature selection efficiency and effectiveness.

problem Balancing feature selection efficiency and effectiveness.
method Interactive Reinforced Feature Selection (IRFS) framework with diverse trainers.
result Improved feature selection efficiency and effectiveness compared to existing methods.

Study proposes automated framework for REM Sleep Behaviour Disorder detection.

problem Early detection of REM Sleep Behaviour Disorder (RBD) as a predictor of Parkinson's disease.
method Automated sleep staging followed by RBD identification using a Random Forest classifier and 156 features from EEG, EOG, and EMG channels.
result Automated RBD detection achieved 96% accuracy, surpassing individual established metrics.

Cardea automates machine learning for EHRs, improving model building efficiency.

problem Lack of a trusted, open-source framework for automated machine learning in EHRs.
method Uses FHIR for data structure, AUTOML frameworks for feature engineering, model selection, and tuning, and an adaptive data assembler.
result Demonstrates framework's effectiveness on 5 prediction tasks, highlighting its flexibility and human competitiveness.

The paper predicts brain tumor patient survival using segmentation and features.

problem Predicting overall survival of brain tumor patients.
method Automated brain tumor segmentation and use of age, shape, and volumetric features for prediction.
result Random forest classifier achieves 59% accuracy on test dataset and 67% on gross total resection dataset.

Study finds transparency and model performance metrics increase trust in AutoML systems.

problem Understanding what information influences trust in AutoML systems.
method Three studies: qualitative interviews, controlled experiment, and card-sorting task.
result Transparency and model performance metrics are most important for establishing trust in AutoML systems.

Proposes a framework for automated radiation therapy treatment planning with uncertainty quantification.

problem Quantifying uncertainties in dose-related quantities for automated treatment planning.
method Three-step pipeline: feature extraction, dose statistic prediction, and dose mimicking.
result Probabilistic treatment plans agree better with clinical counterparts than non-probabilistic ones.

A novel method automates quality control of fMRI scans, improving accuracy and generalizability.

problem Lack of automated QC for fMRI scans limits clinical neuroscience research.
method Train machine learning classifiers using runtime log features to predict scan quality.
result Classifiers trained on FLAG-QC features outperform previous methods (AUC=0.79 vs AUC=0.56).

A RL framework selects features to balance bias and accuracy dynamically.

problem Bias in automated feature selection when predictors are correlated.
method Multi-component reward function with policy gradient for dynamic regularization and bias mitigation.
result Model balances fairness and accuracy during training.

Automated feature engineering improves interpretable models without manual work.

problem Lack of interpretability in complex models causes trust and stability issues.
method Use elastic black-box models to create simpler, interpretable glass-box models.
result Extracted features from complex models improve linear model performance.

Study improves radio show segmentation using audio embeddings.

problem Automated segmentation of radio shows.
method Created audio embeddings from multi-class classification tasks on different datasets, evaluated performance against text-only baseline.
result Audio embeddings from non-speech sound event classification significantly outperformed text-only baseline by 32.3% in F1-measure.

Automates feature selection and weighting in molecular systems.

problem Optimal feature selection and alignment in molecular systems.
method Differentiable Information Imbalance (DII) method for automated feature ranking and scaling.
result Automated feature selection and scaling that preserves information content and interpretability.

Automated feature extraction for bearing health monitoring.

problem Predicting mechanical faults in process industries to prevent shutdowns.
method Stacked autoencoder neural network and OSELM for automated feature extraction.
result 100% detection accuracy for bearing health states.

Feature Learning aims to extract relevant information contained in data sets in an automated fashion. It is driving force behind the current deep learning trend, a set of methods that have had widespread empirical success. What is lacking is a theoretical understanding of different feature learning schemes. This work p…

2015-04-01abs ↗pdf ↗

Deep learning matches classical feature-based AS models for TSP.

problem Automated selection of algorithms for the TSP.
method Evolved instances, deep neural network, visual representation.
result Deep learning approach matches classical feature-based models.

TODS automates time series outlier detection with customizable pipelines.

problem Automated detection of outliers in time series data.
method Modular system with 70 primitives for data processing, time series analysis, and detection algorithms. GUI and data-driven searcher for pipeline design.
result Automated discovery and construction of effective outlier detection pipelines.

This paper improves transportation efficiency by teaching automated vehicles to cooperate.

problem Improving efficiency and safety of transportation systems with automated vehicles.
method Multi-agent graph reinforcement learning with attention mechanism.
result Automated vehicles can achieve better performance when learning to cooperate with each other.

Paper proposes a method to predict EL difficulty using consensus-based labels.

problem Challenges in automatically identifying and resolving ambiguous entity mentions.
method Consensus-based method to generate difficulty labels, supervised classification with various features.
result EL difficulty can be accurately predicted with high accuracy.

Painless Activation Steering automates post-training for LMs without manual intervention.

problem Manual post-training methods are time-consuming and labor-intensive.
method Painless Activation Steering (PAS) is a fully automated approach that requires no manual intervention.
result PAS reliably improves performance for behavior tasks but not for intelligence-oriented tasks.

A feature-weighted mean shift algorithm improves clustering in high-dimensional data.

problem Clustering high-dimensional data with traditional mean shift algorithms.
method Feature-weighted mean shift algorithm.
result The algorithm outperforms conventional mean shift and preserves computational simplicity.

This paper addresses practical challenges in portfolio optimisation for automated trading.

problem Implementing optimal portfolio weights into real trades with transaction costs and lot sizes.
method Two-stage framework: optimises portfolio weights first, then generates realistic trades.
result The two-stage approach effectively converts optimal portfolios into actionable trades, mitigating practical difficulties.

A new DRL system using LSTM improves stock trading performance.

problem Adapting DRL to financial data with low signal-to-noise ratios.
method Cascaded LSTM networks for feature extraction and reinforcement learning.
result Our model outperforms previous models in cumulative returns and Sharp ratio.

Deep learning model reconstructs material microstructures from feature representations.

problem Reconstructing complex material microstructures accurately and efficiently.
method Convolutional deep belief network for automated feature learning and dimension reduction.
result Material reconstructions preserve microstructural features and material properties.

AutoBayes automates Bayesian graph exploration for robust machine learning.

problem Learning representations invariant to nuisance variations in machine learning.
method Automated Bayesian inference framework exploring different graphical models.
result Significant performance improvement with nuisance-invariant machine learning pipelines.

Simplified feature selection using a single agent with restructured choice strategy.

problem Efficiency and cost issues in multi-agent reinforced feature selection.
method Single-agent approach with restructured choice strategy, including scanning method, feature prioritization, state representation, and reward scheme.
result Improved efficiency and effectiveness of feature selection.