INFUSER improves reasoning by co-evolving a generator and solver with adaptive curriculum.
problem Improving reasoning in language models with minimal external supervision.
method INFUSER uses a Generator and Solver to co-evolve iteratively, rewarding the Generator with an influence score and the Solver with correctness rewards.
result INFUSER outperforms self-evolution baselines by over 20% on Olympiad and SuperGPQA benchmarks.
INFUSER improves reasoning by self-evolving with a generator and solver that co-learn from unstructured documents.
problem Improving reasoning through self-evolution
method INFUSER uses a generator and solver co-evolving in a document pool to improve reasoning.
result INFUSER outperforms strong self-evolution baselines on Olympiad and SuperGPQA benchmarks.
This work trains a model to generate high-quality samples from random noise.
problem Generating high-quality samples from random noise.
method Infusion training to learn a Markov chain transition operator.
result The method produces high-quality samples in a small number of steps.
CHEER boosts poor models using rich model knowledge.
problem Improving performance of models trained on limited data.
method Develops CHEER framework to infuse rich model knowledge into poor models.
result CHEER significantly improves model performance on physiological datasets.
Deep RL controls anesthesia more accurately than traditional methods.
problem Controlling the level of unconsciousness during anesthesia.
method Deep Reinforcement Learning (DRL) to map patient state to propofol dosage.
result Deep RL model outperformed traditional controllers (1.7% vs 3.4% median absolute performance error).
Syntax-enhanced models boost machine translation and NLP performance.
problem Limited training data and complex models struggle in NLP tasks.
method Syntax information was explicitly fed into Transformer and BERT models.
result Syntax-infused models achieved significant BLEU improvements.
SIVAE integrates sentences and their syntactic trees for improved text generation.
problem Improving the grammar of generated text.
method SIVAE uses two separate latent spaces for sentences and syntactic trees, optimizing a joint distribution with two encoders and two decoders.
result SIVAE generates sentences with better grammar compared to existing models.
Propagation-regularization improves GNN performance by infusing extra graph information.
problem The effectiveness of graph Laplacian regularization in GNNs is questioned and improved upon.
method Introducing Propagation-regularization (P-reg) to enhance GNN performance.
result P-reg boosts GNN performance on various tasks across multiple datasets.
The paper improves itemset quality assessment by incorporating background knowledge.
problem Assessing the quality of discovered itemsets is challenging due to many patterns being explainable by background knowledge.
method The authors introduce a maximum entropy approach to efficiently infuse additional background knowledge such as row margins, lazarus counts, and bounds of ones.
result More sophisticated models that incorporate background knowledge fit the data better and improve frequency prediction of itemsets.
Proposes KTAN for better training of student networks with both intermediate representations and probability distributions.
problem Reduces large computation and storage cost of deep networks by transferring generalization ability.
method Holistically considers intermediate representations and probability distributions; uses a Teacher-to-Student layer and adversarial learning.
result Significantly improves performance of student networks on image classification and object detection tasks.
Paper introduces CoD for few-shot task-aware knowledge distillation using counterfactual explanations.
problem Lack of data for task-aware distillation in resource-constrained scenarios.
method Counterfactual-explanation-infused Distillation CoD for few-shot task-aware knowledge distillation.
result CoD achieves superior performance with significantly fewer samples than baseline methods.
Noise analysis detects backdoors in DNNs quickly.
problem Detecting backdoors in DNNs trained on compromised data.
method Noise-infused image titration curves to quantify robustness and detect backdoors.
result DNNs with backdoors are more sensitive to noise and reveal their targets.
Mod-DeepESN improves echo state networks for complex, multi-scale tasks.
problem Efficiency in solving complex, multi-scale temporal tasks.
method Incorporates intrinsic plasticity into a modular deep echo state network architecture.
result Significantly outperforms state-of-the-art for time series prediction tasks.
Paper introduces a trading agent using LLMs for risk assessment and trading recommendations.
problem Developing a trading agent that can handle financial risks effectively.
method Extending CPPO algorithm with LLM-generated risk assessment and trading signals from financial news.
result Backtesting shows improved performance of the trading agent compared to benchmarks.
Method generates visual explanations for similarity models without classification.
problem Lack of visual explanations for similarity models trained without classification loss.
method Gradient-based visual attention using learned feature embeddings.
result Attention maps improve model performance and can be used as constraints.
Deep RL model optimizes dynamic portfolio allocation.
problem Automating dynamic portfolio optimization with machine learning.
method Model-based deep reinforcement learning with prediction, data augmentation, and behavior cloning modules.
result Robust, profitable, and risk-sensitive trading strategy compared to baselines.
NESYM combines AI and Earth models for new climate insights.
problem Replacing traditional Earth models with AI.
method Neural Earth System Modelling (NESYM) integrating AI and climate models.
result Artificial intelligence may render traditional models obsolete.
Adaptively preconditions SGLD for faster convergence and better generalization.
problem Pathological curvature in deep network loss landscapes.
method Adaptive estimation of noise parameters to precondition isotropic gradient noise.
result Adaptively preconditioned SGLD achieves faster convergence and generalization equivalent of SGD.
New method improves text classification without labeled target data.
problem Improving text classification under domain shift without labeled target data.
method Diversity-based generalization using multi-head attention with diversity constraints.
result Method matches state-of-the-art performance without labeled target data.
Research shows that information asymmetry affects how quickly companies adjust their capital structure and expected returns.
problem The relationship between capital structure adjustment speed and expected returns is influenced by information asymmetry.
method A hybrid data regression model was used to test the hypotheses based on data from 120 companies in the Tehran Stock Exchange.
result Information asymmetry positively affects the relationship between capital structure adjustment speed and expected returns.
Paper proposes AI for stock market forecasting using external knowledge.
problem Forecasting stock prices influenced by external factors.
method Learning from historical data and external temporal knowledge graphs modeled as Hawkes processes.
result Dynamic representations effectively rank stocks based on returns.
BelMan uses Bayesian methods to optimize decisions in multi-armed bandit problems.
problem Optimizing decisions in multi-armed bandit problems with varying rewards and beliefs.
method BelMan uses a geometric approach with information projection and reverse projection to balance exploration and exploitation.
result BelMan outperforms other algorithms in specific scenarios involving many arms and continuous rewards.
Generative framework learns effective, lower-dimensional models from high-dimensional data.
problem Predicting long-term behavior of complex, multiscale systems with limited data.
method Physics-aware probabilistic model order reduction with latent variables.
result Guaranteed long-term stability and predictive accuracy in multiscale physical systems.
Study bridges GARCH and NN models for volatility forecasting.
problem Lack of interaction between GARCH and NN approaches for volatility forecasting.
method Established equivalence between GARCH and NN models, introduced GARCH-NN approach.
result GARCH-NN approach enhances volatility forecasting compared to standalone models.
Self-taught optimizer improves code generation using language models.
problem Improving code generation using language models.
method Recursive self-improvement of a scaffolding program that generates code.
result Improved scaffolding program generates programs with significantly better performance.
A statistical description and model of individual healthcare expenditures in the US has been developed for measuring value in healthcare. We find evidence that healthcare expenditures are quantifiable as an infusion-diffusion process, which can be thought of intuitively as a steady change in the intensity of treatment …
LatFormer improves geometric reasoning by incorporating lattice symmetry priors in attention mechanisms.
problem State-of-the-art models struggle with geometric reasoning tasks in the ARC and LARC datasets.
method Introduced LatFormer, a model that uses lattice symmetry priors in attention masks.
result LatFormer requires 2 orders of magnitude fewer data than standard attention mechanisms.
RAGIC predicts stock intervals with risk considerations, improving prediction accuracy and coverage.
problem Limited success in predicting stock market outcomes due to stochastic nature and risk oversight.
method RAGIC uses a GAN with a risk module and temporal module to generate risk-sensitive stock intervals.
result RAGIC achieves a consistent 95% coverage with narrow interval widths, balancing accuracy and risk.
RETAIN model improves glucose forecasting for diabetics, offering both accuracy and interpretability.
problem Inability of deep learning models to interpret their predictions in healthcare.
method Two-level attention mechanism in a recurrent neural network (RETAIN) architecture.
result RETAIN model achieves comparable accuracy to LSTM and FCN models while being highly interpretable.
Taylorized training improves neural network training at finite width.
problem Understanding and improving neural network training at finite width.
method Training the k-th order Taylor expansion of the neural network at initialization.
result Taylorized training agrees with full neural network training better as k increases and can significantly close the performance gap.
Self-training outperforms pre-training on COCO object detection and segmentation datasets.
problem The effectiveness of pre-training in improving object detection and segmentation models is limited.
method Investigated self-training as an alternative method to utilize additional data.
result Self-training consistently improves model performance across various dataset sizes and data augmentation levels.
Un-trained neural networks outperform trained methods in MRI reconstruction.
problem Accelerated MRI reconstruction with minimal training data.
method Variation of Deep Decoder without training data.
result Un-trained approach significantly outperforms other methods in reconstruction accuracy.
Meta-learning performance depends on train-validation split type.
problem Understanding the importance of train-validation split in meta-learning.
method Theoretical and experimental study comparing train-val and train-train methods.
result Train-train method can achieve strictly better excess loss in realizable cases.
New method reveals how training data influence diffusion model outputs.
problem Difficulty in assessing training data impact on diffusion model outputs.
method Use of ensembles trained on carefully engineered splits of training data to identify influential training examples.
result Demonstrated the viability of ensembles as generative models and validity of assessing influence.
Paper shows adversarial training can be fooled by new type of noise.
problem Adversarial training can be fooled by new types of noise.
method Designing ADVIN, a new type of inducing noise.
result ADVIN can degrade adversarial training robustness by 99.9%.
New MIP methods improve training of integer-valued neural networks.
problem Training integer-valued neural networks with limited data and resources.
method Formulated new MIP models to optimize training efficiency and handle more data.
result Significantly outperforms previous state-of-the-art methods in accuracy, training time, and data usage.
Free adversarial training reduces the generalization gap compared to vanilla method.
problem Improving generalization in adversarial training.
method Analysis of algorithmic stability in free adversarial training.
result Free adversarial training shows a lower generalization gap.
ADASS selects adaptive subsets for SGD training acceleration.
problem Fixed sample size in SGD limits training efficiency.
method ADASS selects adaptive subsets based on Lipschitz constants.
result ADASS achieves comparable accuracy with full training set.
Solve-training trains neural nets to map physical solutions efficiently.
problem Representing complex physical solutions with neural networks.
method Variational training using loss functions from physical models.
result Effective neural network representation of solution maps without expensive labels.
MixTrain improves verifiable robustness of neural networks without sacrificing efficiency.
problem Efficiently making neural networks robust against adversarial attacks.
method Stochastic robust approximation and dynamic mixed training techniques.
result Achieves up to 95.2% verified robust accuracy with significantly reduced training time.
learn2mix trains neural nets faster by adjusting class proportions dynamically.
problem Training neural nets efficiently with limited resources and imbalanced classes.
method Adaptive class proportion adjustment during training.
result Neural nets trained with learn2mix converge faster than static methods.
Loss-guided training accelerates node embedding methods on graphs.
problem Training efficiency in graph learning methods with implicit positive examples.
method Dynamic adjustment of training distribution based on loss values.
result Significant acceleration in training and computation over static methods.
Paper explores fast adversarial training to improve robustness with less computation.
problem Efficiently defending against adversarial examples.
method Integrates simple self-attacks for faster training, focusing on overfitting recovery.
result Shows superior robust accuracy with reduced training time compared to strong adversarial training.
Three LF training criteria improve neural network acoustic models without cross-entropy pre-training.
problem Improving purely sequence-trained neural network acoustic models.
method Comparison of three lattice-free discriminative training criteria (MMI, bMMI, sMBR) on LVCSR tasks.
result LF-bMMI models outperform plain LF-MMI models by 5% WER on Switchboard datasets.
Free adversarial training improves robustness without generating adversarial examples.
problem Training robust models against adversarial attacks is costly and impractical for large-scale datasets.
method Recycles gradient information from parameter updates to generate adversarial examples.
result Free adversarial training achieves comparable robustness to PGD training at negligible cost.
Detects backdoors in trained models without poisoned training data.
problem Detecting backdoors in DNNs trained without access to the poisoned training set.
method Proposes a novel detector using the maximum achievable misclassification fraction (MAMF) statistic.
result Detects backdoors and infers source and target classes.
Pre-training improves model robustness and uncertainty.
problem Improving model robustness and uncertainty in machine learning.
method Adversarial pre-training and task-specific methods.
result Approximately a 10% absolute improvement in adversarial robustness.
Self-training improves GANs for semi-supervised learning.
problem Training GANs with limited labeled data.
method Combining self-training with GANs' infinite data generation.
result Self-training improves GANs' performance in semi-supervised learning.