Self-taught optimizer improves code generation using language models.
problem Improving code generation using language models.
method Recursive self-improvement of a scaffolding program that generates code.
result Improved scaffolding program generates programs with significantly better performance.
Improves algorithmic recourse to guide towards both acceptance and improvement.
problem Algorithmic recourse recommendations may not lead to improvement.
method Improvement-Focused Causal Recourse (ICR) requires recommendations to guide towards improvement and leverages causal knowledge to design accurate decision systems.
result ICR guides towards both acceptance and improvement given correct causal knowledge.
This study evaluates price improvements in order flow auctions on Ethereum.
problem Improving trading outcomes in blockchain-based trading platforms.
method Utilized open-source tools to attribute price improvements to specific system inputs.
result Auction-enhanced interfaces can provide statistically significant improvements in trading outcomes, averaging 4-5 basis points.
Self-improvement refines language models by verifying their own outputs.
problem Improving language models without external feedback.
method Formalizing self-improvement as sharpening, using the model itself as a verifier.
result RLHF-based self-improvement can outperform SFT-based methods.
If ( M , g ) (M,g) ( M , g ) is a compact Riemannian manifold of dimension n ≥ 2 n\ge 2 n ≥ 2 we give necessary and sufficient conditions for improved L p ( M ) L^p(M) L p ( M ) -norms of eigenfunctions for all 2 < p ≠ p c = 2 ( n + 1 ) n − 1 2<p\ne p_c=\tfrac{2(n+1)}{n-1} 2 < p = p c = n − 1 2 ( n + 1 ) , the critical exponent. Since improved L p c ( M ) L^{p_c}(M) L p c ( M ) bounds imply improvement all other exponents, these conditions are nece…
Study improves online learning with adaptable agents in various settings.
problem Learning with improving agents in online settings.
method Extensive analysis of combinatorial dimensions, multiclass setup, bandit feedback, and agent cost.
result Characterization and analysis of online learnability in the model.
We prove new improved endpoint, L p c L^{p_c} L p c , p c = 2 ( n + 1 ) n − 1 p_c=\tfrac{2(n+1)}{n-1} p c = n − 1 2 ( n + 1 ) , estimates (the "kink point") for eigenfunctions on manifolds of nonpositive curvature. We do this by using energy and dispersive estimates for the wave equation as well as new improved L p L^p L p , 2 < p < p c 2<p< p_c 2 < p < p c , bounds of Blair and the author \cite{BSTop}, \…
Improves policies with high certainty, even in small samples.
problem Ensuring new policies are better than the baseline with high probability.
method Leverages powerful safety tests and multiple testing for threshold policies.
result Controls the rate of adopting a worse policy to pre-specified error level.
IVON improves LoRA finetuning with minimal overhead and significant accuracy gains.
problem Improving LoRA finetuning with Bayesian methods.
method IVON, a variational algorithm, with posterior pruning.
result Significant accuracy improvements over AdamW and other Bayesian methods.
Unified framework for policy improvement in RL with benefits in data efficiency and computation.
problem Improving data efficiency and computation in reinforcement learning for continuous control.
method Local, regularized policy improvement with tree search for continuous action spaces.
result Improves data efficiency and reduces wall-clock time in high-dimensional domains.
New scalarizing functions improve multi-objective Bayesian optimisation.
problem Improving multi-objective Bayesian optimisation efficiency.
method Comparing two infill criteria based on hypervolume improvement.
result Effective scalarizing functions enhance hypervolume maximisation.
This paper improves image retrieval accuracy through novel relevance feedback methods.
problem Improving image retrieval accuracy in Content-Based Image Retrieval (CBIR).
method Novel addition to feature re-weighting and classification techniques, focusing on 0-th iteration improvement.
result Significantly improved retrieval accuracy from relevance feedback.
New insights into Bartnik mass from improvability of dominant energy scalar.
problem Characterizing Bartnik mass minimizing initial data sets.
method Introducing improvability concept, proving non-improvability consequences, and analyzing pp-wave counterexamples.
result Bartnik mass minimizing initial data sets are characterized, advancing conjectures.
This paper proposes a new approach to RL by focusing on the value-improvement path.
problem Value prediction problems in RL are sequence-dependent and require holistic approach.
method Characterize and approximate the value-improvement path holistically.
result A representation that spans the value-improvement path provides accurate value approximations for future policy improvements.
New RL algorithms improve control tasks with data reuse.
problem Real-world control requires performance guarantees and data efficiency.
method Generalized Policy Improvement combining on-policy guarantees and sample reuse.
result Extensive experimental analysis shows benefits of new algorithms.
Different technological domains have significantly different rates of performance improvement. Prior theory indicates that such differing rates should influence the relative speed of diffusion of the products embodying the different technologies since improvement in performance during the diffusion process increases th…
Improved uniform convergence bound with fat-shattering dimension reduces sample complexity gap.
problem Gap between upper and lower bounds on sample complexity for fat-shattering dimension.
method Provided an improved uniform convergence bound.
result Closed the gap between existing upper and lower bounds on sample complexity.
This paper investigates a type of instability that is linked to the greedy policy improvement in approximated reinforcement learning. We show empirically that non-deterministic policy improvement can stabilize methods like LSPI by controlling the improvements' stochasticity. Additionally we show that a suitable represe…
Improves online learning with expert demonstrations, quality matters.
problem Improving online learning through offline demonstration data.
method Thompson sampling applied to a multi-armed bandit model, informed by expert demonstrations and Bayes' rule.
result Substantial empirical regret reduction with expert demonstrations, improving online performance.
Ensemble models improve prediction calibration for mismatched distributions.
problem Calibration issues in deep neural networks with mismatched train and test distributions.
method Simple data augmentation and mixing techniques for ensemble models.
result Improves calibration and accuracy on CIFAR10 and CIFAR100 benchmarks.
There is growing evidence that converting targets to soft targets in supervised learning can provide considerable gains in performance. Much of this work has considered classification, converting hard zero-one values to soft labels---such as by adding label noise, incorporating label ambiguity or using distillation. In…
The paper develops a theory for iterative self-improvement of models, proving conditions for better performance with easy-to-hard curricula.
problem Lack of theoretical foundation for iterative self-improvement in practical settings.
method Modeling self-improvement as maximum-likelihood fine-tuning on reward-filtered distributions and proving finite-sample guarantees.
result Explicit feedback loop and conditions for better performance with easy-to-hard curricula.
Recent work has increased the performance of Generative Adversarial Networks (GANs) by enforcing a consistency cost on the discriminator. We improve on this technique in several ways. We first show that consistency regularization can introduce artifacts into the GAN samples and explain how to fix this issue. We then pr…
This paper calculates the exact probability distribution of hypervolume improvement for bi-objective problems.
problem Calculating the exact probability distribution of hypervolume improvement in bi-objective problems.
method Cell partition-based method to derive the probability distribution of hypervolume improvement from a bi-variate Gaussian random variable.
result The proposed ε \varepsilon ε -PoHVI acquisition function outperforms other related functions in Bayesian optimization. We present Chen-Ricci inequality and improved Chen-Ricci inequality for curvature like tensors. Applying our improved Chen-Ricci inequality we study Lagrangian and Kaehlerian slant submanifolds of complex space forms and C-totally real submanifolds of Sasakian space forms.
IMeL turns RL into SL by interpolating improved experiences.
problem Improving reinforcement learning efficiency and scalability.
method IMeL uses a reservoir of experiences and a NN regressor for interpolation.
result IMeL achieves preliminary results and proposes itself as a baseline.
Deep Convolutional Neural Networks (CNNs) are more powerful than Deep Neural Networks (DNN), as they are able to better reduce spectral variation in the input signal. This has also been confirmed experimentally, with CNNs showing improvements in word error rate (WER) between 4-12% relative compared to DNNs across a var…
Improved estimates for singularities in capillary surfaces.
problem Understanding the singularities of minimizing capillary hypersurfaces.
method Improved estimates based on connections to the one-phase Bernoulli problem.
result The singular set is of codimension at least 4, improving for specific angles.
This paper improves prediction rule ensembles using model-based data generation.
problem Improving the sparsity and predictive accuracy of prediction rule ensembles.
method The authors use surrogate models to train Lasso regression with data generated by a boosted decision tree ensemble, improving PRE performance.
result The use of surrogacy models can substantially improve the sparsity of PRE while retaining predictive accuracy.
Improves bounds on surface decompositions.
problem Surface decomposition bounds
method Improved theorem of Bers for surfaces with boundary and closed surfaces.
result Slight improvement on best known bound for closed surfaces.
Improved regret bound for adversarial bandit convex optimisation.
problem Minimizing regret in zeroth-order adversarial bandit convex optimisation.
method Identifying an improved exploratory distribution for convex functions.
result Proved minimax regret bound of O ( d 2.5 n log ( n ) ) O(d^{2.5} \sqrt{n} \log(n)) O ( d 2.5 n log ( n )) . Improves tree model performance by considering future node splits.
problem Improving tree model performance.
method Next-Depth Lookahead Tree (NDLT) model that evaluates future node splits.
result Enhanced tree model performance.
QFIL improves offline RL by filtering data to reduce bias and variance.
problem Improving offline reinforcement learning policies with limited data.
method QFIL uses a filtered dataset to improve policies, trading off bias and variance through quantile selection.
result QFIL provides a safe policy improvement step with function approximation and effectively balances bias and variance.
Improves fairness in machine learning by adding underrepresented group data.
problem Machine learning biases across subgroups due to under-representation or societal biases.
method Data augmentation via pairwise mixup across subgroups to balance subpopulations.
result Achieves fair outcomes with robust if not improved accuracy.
New framework studies policy learning problems under data scarcity.
problem Learning improving policies when data is insufficient.
method Developed a mathematical framework for policy learning problems.
result Reduced policy learning problems to simpler ones in sample complexity.
Pruning improves model generalization in over-parameterized models, contradicting traditional theories.
problem Pruning's effect on generalization in over-parameterized models.
method Empirical study on standard pruning algorithms and additional regularization effects.
result Pruning leads to better training and regularization, improving generalization.
Improved YOLOv5 model detects mask-wearing with enhanced accuracy.
problem Detecting mask-wearing in high traffic public places.
method Improved YOLOv5l with Multi-Head Attentional Self-Convolution, Swin Transformer Block, I-CBAM module, and enhanced feature fusion.
result 1.1% improvement in mAP(0.5) and 1.3% improvement in mAP(0.5:0.95) compared to YOLOv5l.
Paper derives a simplified formula for Expected Improvement using log-transformed data.
problem Challenges in enhancing Bayesian optimization with Expected Improvement.
method Derives a closed form of Expected Improvement for Gaussian process trained on log-transformed objective.
result Provides a simplified formula for Expected Improvement.
ViCE uses superpixels to enhance self-supervised learning for better dense visual embeddings.
problem Lack of high-resolution feature maps from self-supervised models.
method Superpixels for dense representation learning, contrasting over regions.
result Improves unsupervised semantic segmentation on benchmarks like Cityscapes and COCO.
We improve private training accuracy with learning rate schedules and matrix factorizations.
problem Private training with learning rate schedules and correlated noise.
method General upper and lower bounds for learning rate schedules, memory-efficient constructions, and schedule-aware factorizations.
result Schedule-aware factorizations improve accuracy in private training.
Generative AI boosts productivity and improves customer service quality.
problem Productivity and quality of customer support agents.
method Staggered introduction of a generative AI-based conversational assistant in customer support.
result AI increases productivity by 15% on average, with significant heterogeneity across workers.
Proposes a method to improve learning when training data is not representative.
problem Improving supervised learning when training data is not representative (covariate shift).
method Conditioning on propensity scores to balance covariates within strata.
result Significantly improved target prediction and AUC (0.958) on supernovae classification challenge.
MQTransformer improves forecast accuracy with context-aware attention.
problem Improving probabilistic demand prediction accuracy.
method Incorporates Transformer architectures for context alignment and feedback-aware attention.
result Significant improvements in forecast accuracy, reducing excess variability.
DA improves solar wind forecasts by updating model boundary conditions.
problem Improving solar wind forecasting accuracy.
method Variational Data Assimilation with solar wind model and in-situ observations.
result DA forecasts are more accurate than non-DA forecasts, especially when STEREO-B's latitude is offset from Earth.
Recent developments in the field of robot grasping have shown great improvements in the grasp success rates when dealing with unknown objects. In this work we improve on one of the most promising approaches, the Grasp Quality Convolutional Neural Network (GQ-CNN) trained on the DexNet 2.0 dataset. We propose a new arch…
Improved iterative methods for risk parity portfolio weights.
problem Solving for portfolio weights in risk parity allocation.
method Enhanced CCD and Newton methods, including a rescaling step and improved initial guess.
result Improved CCD method is the best, three times faster with 40% fewer iterations.
Develops framework for estimating and improving DTRs with time-varying IV in the presence of unmeasured confounding.
problem Estimating DTRs from observational data with unmeasured confounding.
method Time-varying instrumental variable (IV) framework for estimating and improving DTRs.
result IV-optimal and IV-improved DTRs perform better than DTRs assuming no unmeasured confounding.
Improvement-aware algorithms can achieve zero error in learning tasks.
problem Learning with agents who can improve their performance.
method Develops algorithms that account for agent improvement to achieve zero error.
result Improvement can reduce sample complexity or make learning harder.