ParK efficiently solves kernel ridge regression for large datasets.
problem Large-scale kernel ridge regression efficiency and accuracy.
method Partitioning feature space with random projections and iterative optimization.
result Provably maintains statistical accuracy with reduced space and time complexity.
Space-efficient feature maps improve string alignment kernel scalability.
problem String alignment kernels scale poorly with quadratic complexity, limiting large-scale applications.
method Presented SFMEDM, a space-efficient feature map for edit distance with moves using metric embedding and random Fourier features.
result Demonstrated superior performance of SFMEDM in prediction accuracy, scalability, and computation efficiency.
Efficient RL in large POMDPs with latent determinism and embeddings.
problem Efficient reinforcement learning in large-scale POMDPs with latent states and observations.
method Conditional Hilbert space embeddings, linear optimal Q Q Q -function, deterministic latent transitions, gap assumption. result Computationally and statistically efficient algorithm for exact optimal policy.
Gradient-based feature selection for large datasets.
problem Feature selection for large datasets with high-order correlations.
method Iterative mini-batch calculation, discrete-to-continuous relaxation.
result Efficiently finds higher-order feature correlations in both N > D and N < D regimes.
Paper uses RL and DCAE to classify large unstructured data with fewer features.
problem Classifying large unstructured data with high precision using fewer features.
method Deep Convolutional Autoencoder (DCAE) for feature learning and Double DQN/Retrace RL algorithms for policy optimization.
result The approach achieves high classification precision with fewer features than traditional methods.
For supervised and unsupervised learning, positive definite kernels allow to use large and potentially infinite dimensional feature spaces with a computational cost that only depends on the number of observations. This is usually done through the penalization of predictor functions by Euclidean or Hilbertian norms. In …
Quantum-enhanced feature spaces improve machine learning performance.
problem Large feature spaces and computationally expensive kernel functions in machine learning.
method Two novel quantum methods: quantum variational classifier and quantum kernel estimator.
result Quantum-enhanced classifiers achieve better performance on noisy quantum computers.
New method selects variables for nonlinear regression in large datasets.
problem Nonlinear regression with large-scale datasets.
method Kernel-based variable selection with random features.
result Outstanding performance on large-scale synthetic and real datasets.
Improves MKL for multi-class classification with better feature selection and representation.
problem Real-world multi-class classification problems with non-linear separations.
method Large-margin multiple kernel learning (LMMK) with sparsity term for discriminative feature selection.
result Competitive classification accuracy and sparse non-zero kernel weights.
This paper considers the multi-task learning problem and in the setting where some relevant features could be shared across few related tasks. Most of the existing methods assume the extent to which the given tasks are related or share a common feature space to be known apriori. In real-world applications however, it i…
New methods bound uncertainties in large, noisy data for robust SVM classification.
problem Uncertainty in large, noisy data for robust SVM classification.
method Formulate robust optimization problem with bounding schemes for random features using Random Fourier Features and Nyström methods. Solve with stochastic approximation techniques.
result Efficient solutions for large, noisy data classification.
Novel Newton method for large-scale kernel methods using random features.
problem Efficiently solving large-scale finite-sum minimization problems in RKHS.
method Randomized feature-based Newton method for empirical risk minimization.
result Local superlinear and global linear convergence of the method.
Proposes feature conformal prediction for broader application in semantic feature spaces.
problem Establishing valid prediction intervals in semantic feature spaces.
method Extends conformal prediction to semantic feature spaces using deep representation learning.
result Feature conformal prediction outperforms regular conformal prediction under mild assumptions.
This paper presents a general graph representation learning framework called DeepGL for learning deep node and edge representations from large (attributed) graphs. In particular, DeepGL begins by deriving a set of base features (e.g., graphlet features) and automatically learns a multi-layered hierarchical graph repres…
Random feature approximation speeds up spectral methods and improves learning rates.
problem Improving the efficiency and generalization of spectral methods in large-scale algorithms.
method Combining random feature approximation with spectral regularization methods.
result Optimal learning rates for estimators over various regularity classes, including those not in the RKHS.
New approach for learning large Bayesian networks using feature clustering and compression.
problem Learning large Bayesian networks efficiently and accurately.
method Feature space clustering, compression, and Hill-Climbing algorithm with BIC and MI score functions.
result Potential for parallelizable block learning and improved structure accuracy.
Enhances stock return prediction using LLMs and hybrid models.
problem Insufficient use of semantic information and alignment of LLMs with stock features.
method LG model with three strategies for global information modeling and SCRL for embedding alignment.
result Superior performance in Rank Information Coefficient and returns compared to models relying only on stock features.
Study analyzes feedback complexity for sparse feature retrieval in deep networks.
problem Learning sparse superposed features with feedback.
method Analysis of feedback complexity in sparse settings, including triplet comparisons.
result Establishes tight bounds and strong upper bounds for feature recovery.
New method selects features for sequential decision making.
problem Dynamic feature selection for instance-wise decisions.
method Latent variable model trained in a supervised manner; reasoning across stochastic latent space.
result Outperforms existing methods on various datasets.
New geometric interpretation explains over-parameterized models and adversarial perturbations.
problem Geometric understanding of over-parameterized regression and adversarial perturbations.
method Alternative geometric interpretation of regression in feature space.
result Adversarial perturbations are a natural feature of biased models due to underlying geometry.
Gradient descent reshapes the function space of neural networks.
problem Understanding how feature learning affects the function space of neural networks.
method Characterized the evolution of the feature space during training using a two-layer neural network.
result Gradient descent induces a data-adaptive deformation that selectively enhances signal-aligned directions.
This paper uses Reinforcement Learning to select features from a large dataset.
problem Selecting the best features to minimize variance and bias in machine learning models.
method Formulated the feature selection problem as a Markov Decision Process (MDP) and used Temporal Difference (TD) algorithm.
result The approach using Reinforcement Learning outperformed other methods in selecting features.
Graph-Hist classifies social media graphs using feature histograms.
problem Classifying large, sparse social media graphs.
method Extracts latent features, bins nodes, and classifies based on multi-channel histograms.
result Improves bot detection in social media graphs.
Efficiently models complex spatiotemporal patterns with large datasets.
problem Flexible modeling of large-scale spatiotemporal point patterns.
method Cox process inference using Fourier features for scalable inference.
result Approximate Bayesian method fits over 100,000 events in 3D.
Semi-automatic data annotation helps experts label unlabeled samples based on feature space projection.
problem Laborious manual data annotation for machine learning.
method Interactive semi-automatic approach using feature space projection and semi-supervised learning.
result Reduces user annotation effort and improves classification accuracy.
SASE improves attributed graph clustering for large graphs with linear time and space complexity.
problem Challenges in clustering large attributed graphs due to high computational and memory costs.
method SASE combines node features smoothing, scalable spectral clustering, and adaptive order selection.
result SASE achieves a 6.9% improvement in ACC and a 5.87x speedup on the ArXiv dataset.
RECol generates error columns to improve outlier detection.
problem Outlier detection in data with complex relationships.
method Generates reconstruction error columns for leave-one-out feature sets.
result Improves ROC-AUC and PR-AUC values of common outlier detection methods.
The coarse category was established by Roe to distill the salient features of the large-scale approach to metric spaces and groups that was started by Gromov. In this paper, we use the language of coarse spaces to define coarse versions of asymptotic property C and decomposition complexity. We prove that coarse propert…
Random feature models approximate functions in Banach spaces efficiently.
problem Approximating functions in Banach spaces efficiently.
method Randomly initialized feature maps and linear readout training.
result Universal approximation in Bochner spaces for Banach space-valued models.
MISSION selects features for large datasets efficiently and interpretably.
problem Feature selection challenges in ultra large-scale datasets.
method MISSION uses Count-Sketch data structure for stochastic gradient descent.
result MISSION accurately and efficiently selects features on large-scale datasets.
Characterizes RFF regression in large n , p , N n,p,N n , p , N setting, providing precise learning phases and double descent curve.
problem Characterizes RFF regression in large n , p , N n,p,N n , p , N setting. method Characterizes the exact asymptotics of random Fourier feature (RFF) regression in the realistic setting of large n , p , N n,p,N n , p , N . result Characterizes two qualitatively different phases of learning and the corresponding double descent test error curve.
MatrixRL tackles RL in high-dimensional spaces with feature and kernel methods, achieving near-optimal regret bounds.
problem Challenges of exploration in RL with large state-action spaces.
method MatrixRL combines linear bandit techniques with feature and kernel methods to learn a low-dimensional representation of the transition model.
result MatrixRL achieves an O ( H 2 d log T T ) {O}\big(H^2d\log T\sqrt{T}\big) O ( H 2 d log T T ) regret bound, near-optimal in time T T T and dimension d d d . A new method selects features for clustering without labels.
problem Identifying meaningful features in large datasets.
method Differentiable unsupervised feature selection using a gated Laplacian.
result The method improves clustering performance in noisy data.
No-trick kernel adaptive filtering uses deterministic features for scalability and robustness.
problem Scalability issues in kernel methods for large datasets.
method Deterministic feature-map construction using polynomial-exact solutions.
result Deterministic features outperform random Fourier features in performance and scalability.
One of the major advantages in using Deep Learning for Finance is to embed a large collection of information into investment decisions. A way to do that is by means of compression, that lead us to consider a smaller feature space. Several studies are proving that non-linear feature reduction performed by Deep Learning …
A new method reduces the number of features needed for kernel approximation from cubic to logarithmic.
problem Large datasets make kernel methods computationally expensive and impractical.
method Combines random feature maps with data-dependent feature selection to achieve Nystrom-like performance with fewer features.
result Achieves small kernel matrix approximation error and better test set accuracy with fewer features than state-of-the-art methods.
NN-Stacking improves predictive power of regression models by adjusting stacking coefficients with features.
problem Low predictive power of linear stacking methods.
method NN-Stacking uses neural networks to estimate adaptive stacking coefficients.
result NN-Stacking leads to better predictive power, especially in large datasets.
New method calibrates photometric redshift PDFs more accurately.
problem Inaccurate photometric redshift uncertainties lead to systematic errors.
method Local re-calibration using feature-space regression of Probability Integral Transform (PIT) distributions.
result Calibrated PDFs are more accurate at all locations in feature space.
Proposes a new method to approximate kernel functions for large datasets.
problem Limited applicability of kernel methods for large scale datasets.
method Pseudo Random Fourier Features (PRFF) for reducing feature dimensions and improving performance.
result Improves prediction performance and reduces feature dimensions compared to RFF.
The objective of the paper is to study accuracy of multi-class classification in high-dimensional setting, where the number of classes is also large ("large L L L , large p p p , small n n n " model). While this problem arises in many practical applications and many techniques have been recently developed for its solution, to t…
PML-LFC improves PML by estimating label confidence from both feature and label spaces.
problem PML challenges in real-world scenarios where only some labels are relevant.
method PML-LFC estimates label confidence using feature and label space similarities, training a predictor with these values.
result PML-LFC achieves superior performance on synthetic and real-world datasets.
Efficient random binning features improve kernel methods for large datasets.
problem Kernel methods' quadratic complexity limits their scalability to large datasets.
method Proposes and analyzes Random Binning (RB) features, showing faster convergence and parallelizability.
result RB features achieve faster convergence rates and parallelizability advantages compared to other random features.
FSL-Net detects and localizes feature shifts in large, high-dimensional datasets.
problem Feature shifts between data sources lead to erroneous features in various applications.
method FSL-Net is a neural network trained on multiple datasets to localize feature shifts.
result FSL-Net accurately localizes feature shifts from unseen datasets without re-training.
A deep learning method for XML with autoencoder and ranking loss.
problem XML with large label collections, high complexity, inter-label and feature dependencies, and noisy labels.
method Word-vector-based self-attention, ranking-based AutoEncoder architecture.
result Competitive performance on benchmark datasets.
Large learning rates cause oscillations in NN weights that improve generalization.
problem Improving generalization of neural networks trained with large learning rates.
method Theoretical analysis and feature-noise data generation model.
result Oscillating SGD with large learning rates benefits NN generalization by effectively learning weak features.
Paper proposes an efficient RL algorithm for discounted MDPs using feature mapping.
problem Efficient reinforcement learning for large state and action spaces.
method Uses feature mapping to represent states and actions in a low-dimensional space, proposing a novel algorithm with polynomial regret bound.
result Achieves a O ( d T / ( 1 − γ ) 2 ) O(d\sqrt{T}/(1-γ)^2) O ( d T / ( 1 − γ ) 2 ) regret bound, near-optimal up to a ( 1 − γ ) − 0.5 (1-γ)^{-0.5} ( 1 − γ ) − 0.5 factor. A method reduces dimensionality for multi-block data, enhancing feature extraction and classification accuracy.
problem Tractable feature extraction from large-scale, multi-dimensional data.
method Common and individual feature extraction from multi-block data structures using tensor decompositions.
result Significant reduction in dimensionality and enhanced accuracy in feature extraction and classification.
Proposes LM3FE for multi-modal feature extraction in image classification.
problem High-dimensional features and multi-modal data challenges.
method Large margin multi-modal multi-task feature extraction (LM3FE) framework.
result LM3FE outperforms single-task feature extraction and multi-modal feature extraction.