Improved online learning algorithm for easy data with reduced regret.
problem Improving prediction accuracy with limited advice for easy data.
method Second Order Difference Adjustments (SODA) algorithm for online learning.
result Achieves improved regret guarantees for both stochastic and adversarial loss sequences.
Model shows feature learning can improve neural scaling laws for hard tasks.
problem Understanding and improving neural network scaling laws for various task difficulties.
method Developed a solvable model of neural scaling laws, identified three scaling regimes, and demonstrated feature learning's impact on scaling exponents.
result Feature learning can improve scaling with training time and compute for hard tasks, nearly doubling the exponent.
The paper explores when to prioritize easy or hard samples in learning tasks.
problem Determining the optimal order of learning easy or hard samples.
method Theoretical analyses and experiments were conducted to propose and validate four priority modes.
result Four priority modes (easy-first, hard-first, medium-first, two-ends-first) can be flexibly applied.
The paper develops a theory for iterative self-improvement of models, proving conditions for better performance with easy-to-hard curricula.
problem Lack of theoretical foundation for iterative self-improvement in practical settings.
method Modeling self-improvement as maximum-likelihood fine-tuning on reward-filtered distributions and proving finite-sample guarantees.
result Explicit feedback loop and conditions for better performance with easy-to-hard curricula.
We present a new anytime algorithm that achieves near-optimal regret for any instance of finite stochastic partial monitoring. In particular, the new algorithm achieves the minimax regret, within logarithmic factors, for both "easy" and "hard" problems. For easy problems, it additionally achieves logarithmic individual…
LIBS2ML is a fast, scalable library for second order machine learning.
problem Large-scale machine learning problems in big data.
method Scalable second order learning algorithms using MEX files.
result Efficient solution for large-scale learning problems.
We introduce a learning framework called learning using privileged information (LUPI) to the computer vision field. We focus on the prototypical computer vision problem of teaching computers to recognize objects in images. We want the computers to be able to learn faster at the expense of providing extra information du…
DeepDIVA simplifies reproducible deep learning experiments.
problem Difficulty in reproducing deep learning research results.
method A framework for easy experimentation and reproduction.
result Facilitates sharing and easy experimentation of experiments.
The paper defines multivariate confidence intervals that are easy to interpret and retain qualities of one-dimensional counterparts.
problem Applying confidence intervals to multivariate data.
method Defining multivariate confidence intervals that extend one-dimensional definitions and providing efficient approximate algorithms.
result Multivariate confidence intervals retain qualities of one-dimensional counterparts and are easy to interpret.
In Natural Language Processing (NLP) tasks, data often has the following two properties: First, data can be chopped into multi-views which has been successfully used for dimension reduction purposes. For example, in topic classification, every paper can be chopped into the title, the main text and the references. Howev…
ELKI 0.7.5 enhances data mining with R*-tree and open-source algorithms.
problem Improving data mining algorithms and performance.
method R*-tree index structures for high performance and scalability.
result Enhanced performance and scalability in cluster analysis and outlier detection.
Model learns brevity by exposing to easy problems, improving efficiency without explicit length penalties.
problem Excessive verbosity in step-by-step reasoning models trained with RLVR.
method Retaining and up-weighting moderately easy problems as implicit length regularizers.
result Model generates solutions that are, on average, nearly twice as short without explicit length penalties.
Ludwig simplifies deep learning for non-experts.
problem Making deep learning accessible to non-experts.
method Type-based data abstraction and declarative configuration files.
result Ludwig democratizes deep learning, making it accessible to a broader audience.
The results from most machine learning experiments are used for a specific purpose and then discarded. This results in a significant loss of information and requires rerunning experiments to compare learning algorithms. This also requires implementation of another algorithm for comparison, that may not always be correc…
NeuroNER simplifies ANN-based NER for non-experts.
problem Challenging use of ANNs for NER by non-experts.
method Graphical web-based user interface for easy annotation, training, and prediction of entities.
result NeuroNER streamlines NER process for non-expert users.
Randomly biased data makes complex models as easy to learn as simple ones.
problem Learning complex models like multi-index and sparse Boolean functions.
method Introducing a small random shift in the first moment of the data distribution.
result Randomly biased data makes Gaussian single index models and sparse Boolean functions as easy to learn as linear functions.
ARU adapts deep forecasting models in streaming data with efficient updates.
problem Adapting deep globally trained models for streaming data efficiently.
method ARU combines deep global models with closed-form linear models for per-series adaptation.
result ARU outperforms local adaptation methods on various datasets.
Our method upweights easy samples to mitigate forgetting in fine-tuning.
problem Catastrophic forgetting in fine-tuning pre-trained models.
method Sample weighting based on pre-trained model's losses.
result Our method reduces forgetting by up to 0.8% on MetaMathQA while preserving more accuracy on pre-training datasets.
A simple post-learning method boosts deep learning performance.
problem Improving classification performance in deep learning.
method Re-training the final classifier after initial training.
result Enhanced classification performance in deep learning.
The study analyzes neural network predictions of knot invariants and finds that braid representations work best.
problem Understanding and predicting knot invariants using neural networks.
method Investigated different knot representations and invariants, proposed a cosine similarity score.
result Braid representations are best for predicting knot invariants, and some invariants are easier to learn than others.
CEB enhances model resilience through simple entropy bottleneck.
problem Improving model robustness against adversarial attacks.
method Conditional Entropy Bottleneck (CEB) combined with data augmentation.
result CEB significantly boosts adversarial robustness on various benchmarks.
Treeffuser predicts tabular data distributions using gradient-boosted trees.
problem Probabilistic prediction with flexible, non-parametric models.
method Gradient-boosted trees for score estimation in conditional diffusion model.
result Treeffuser outperforms existing methods in probabilistic prediction tasks.
Improves likelihood-free inference for high-dimensional data.
problem Efficient inference for complex models with high-dimensional data.
method Generative Adversarial Networks (GANs) for data-driven summary features.
result Significant improvement in scalability and handling complex distributions.
Algorithm classifies market regimes using time series signatures.
problem Classifying different market conditions from time series data.
method Utilizes path signatures and a metric structure for clustering.
result Established a connection between regime separation and point clustering.
Easy to make models more vulnerable to adversarial perturbations.
problem Designing robust models to adversarial perturbations is hard.
method Inject vulnerabilities into linear layers by increasing sensitivity to low variance components in training data.
result Poisoning attacks can induce vulnerabilities to imperceptible backdoor signals in state-of-the-art networks.
The paper explores when linear system identification is hard or easy, especially for under-actuated systems.
problem Statistical hardness of learning linear systems, especially under-actuated or under-excited systems.
method Using tools from minimax theory and recent statistical tools for finite sample analysis of system identification.
result The controllability index of linear systems affects the sample complexity of identification, making some systems hard to learn.
In this note, we present a new averaging technique for the projected stochastic subgradient method. By using a weighted average with a weight of t+1 for each iterate w_t at iteration t, we obtain the convergence rate of O(1/t) with both an easy proof and an easy implementation. The new scheme is compared empirically to…
For the existence of a branched covering Sigma~ --> Sigma between closed surfaces there are easy necessary conditions in terms of chi(Sigma~), chi(Sigma), orientability, the total degree, and the local degrees at the branching points. A classical problem dating back to Hurwitz asks whether these conditions are also suf…
We aim to design strategies for sequential decision making that adjust to the difficulty of the learning problem. We study this question both in the setting of prediction with expert advice, and for more general combinatorial decision tasks. We are not satisfied with just guaranteeing minimax regret rates, but we want …
We introduce elastic geodesic grids for easy-to-fabricate, deployable structures.
problem Approximating freeform surfaces with deployable structures.
method Geodesic curves on target surfaces, kinematic mechanism, differential geometry.
result Elastic geodesic grids can approximate freeform surfaces easily and deployably.
ECOD detects outliers without parameters, fast and simple.
problem Detecting outliers in large, high-dimensional datasets efficiently and interpretably.
method ECOD estimates empirical cumulative distribution functions per dimension, then computes tail probabilities and outlier scores.
result ECOD outperforms state-of-the-art methods in accuracy, efficiency, and scalability.
Proposes a self-paced multi-label learning method to handle diverse labels efficiently.
problem Learning from multi-label data with a large label space is NP-hard and prone to overfitting.
method Self-paced multi-label learning with diversity (SPMLD) approach, incorporating gradual label inclusion and diversity maintenance.
result The proposed SPMLD framework optimizes a non-convex objective function using block coordinate descent.
Simple method improves uncertainty estimation for distribution shifts.
problem Improving uncertainty estimation in deep image classification under distribution shifts.
method Exposing original model to corrupted images and performing simple statistical calibration.
result Superior performance on various distribution shifts and unsupervised domain adaptation tasks.
Linear regression models are not as interpretable as commonly believed.
problem Interpretability of linear regression models is often overlooked.
method Analysis of common XAI metrics and challenges faced by linear regression models.
result Linear regression models are not inherently interpretable and require careful consideration.
This note provides an easy construction of fake octagons.
problem Existence and construction of fake octagons.
method Elementary cut-and-paste surgery to produce infinitely many distinct fake octagons.
result Any iterate of the surgery produces a fake octagon that is different from others in the family.
A new mechanism for GANs improves text generation by evaluating sub-sequences.
problem Exposure bias and mode collapse in GANs for text generation.
method Segmenting the sequence into sub-sequences and evaluating them individually.
result Significant improvement in text generation models on benchmark data.
Tree Index evaluates cluster quality by creating decision trees from data.
problem Evaluating the quality of cluster results from various techniques.
method Tree Index creates a decision tree from clustered data, combining entropy and depth of leaves.
result Tree Index discriminates between sensible and non-sensible clusters on brain dataset.
Method detects anomalies in small, imbalanced data sets.
problem Anomaly detection in small, imbalanced data sets.
method A novel (1+ε)-class classification method. result Better performance on anomaly detection problems.
Adversarial learning for mixture Hawkes processes improves performance.
problem Learning mixture models of Hawkes processes from event sequences.
method Iterative self-paced learning with adversarial self-paced mechanism.
result The proposed method outperforms traditional methods consistently.
Binary BPS improves sampling for easy mixtures.
problem Sampling from binary distributions efficiently.
method Generalized Bouncy Particle Sampler for binary variables.
result Binary BPS outperforms binary HMC for easy mixtures.
Wide deep neural networks are easy to optimize without constraints.
problem Optimizing wide deep neural networks.
method Analysis of optimization landscapes and empirical-risk minimization.
result Wide neural networks have no confined points, making optimization easier.
New measures quantify how data augmentation improves model performance.
problem Understanding the effectiveness of data augmentation in deep learning.
method Introduced Affinity and Diversity measures to quantify augmentation performance.
result Augmentation performance is best achieved by optimizing both Affinity and Diversity.
This paper is purely expositional. The statement of the Kuratowski graph planarity criterion is simple and well-known. However, its classical proof is not easy. In this paper we present the Makarychev proof (with further simplifications by Prasolov, Telishev, Zaslavski and the author) which is possibly the simplest. In…
New algorithm improves reinforcement learning policies without degrading performance.
problem Policy updates may degrade performance in reinforcement learning with general function approximators.
method Derives a new policy improvement bound with an average divergence instead of sup norm, leading to Easy Monotonic Policy Iteration.
result Generates sequences of policies with guaranteed non-decreasing returns.
MLaut automates machine learning benchmarking.
problem Benchmarking machine learning algorithms on diverse datasets.
method High-level workflow interface, database of datasets, scikit-learn and keras integration.
result Deep neural networks perform poorly on standard supervised learning tasks.
Generative-discriminative method improves label prediction and instance generation.
problem Difficulty in obtaining high-quality labeled instances.
method Generative-discriminative complementary learning method using CC-GAN.
result Improves accuracy in predicting ordinary labels and generating high-quality instances.
Neural model with parameterized algorithms improves graph CO problem solving.
problem Solving NP-hard graph combinatorial optimization problems efficiently and accurately.
method Combining neural models and parameterized algorithms to identify and handle hard and easy parts of CO instances.
result Framework produces superior solution quality and out-of-distribution generalization.
SYNC generates synthetic data from aggregated sources using Gaussian copulas.
problem Creating synthetic datasets from aggregated sources.
method SYNC uses Gaussian copula models to infer high-resolution data from low-resolution sources.
result SYNC successfully merges sampled subsets into a single synthetic dataset.