The paper shows how to learn near-optimal behavior in reinforcement learning with rich observations.
problem Learning near-optimal behavior in reinforcement learning with rich observations and function approximation.
method Introduces a new model called contextual decision processes and a new algorithm that engages in systematic exploration to learn these processes with low Bellman rank.
result The algorithm provably learns near-optimal behavior with a number of samples that is polynomial in all relevant parameters.
VASE uses Bayesian neural networks to improve exploration in sparse reward environments.
problem Exploration in environments with continuous control and sparse rewards.
method VASE uses a Bayesian neural network model of the environment dynamics and variational inference to alternately update the model's accuracy and policy.
result VASE outperforms other surprise-based exploration techniques in continuous control sparse reward environments.
Study on adversarial examples from data size, task, and model factors.
problem Understanding adversarial examples from data size, task, and model perspectives.
method Systematic study on adversarial examples from three aspects: data size, task-dependent, and model-specific factors.
result Adversarial generalization requires more data than standard generalization.
The classification procedure of streaming data usually requires various ad hoc methods or particular heuristic models. We explore a novel non-parametric and systematic approach to analysis of heterogeneous sequential data. We demonstrate an application of this method to classification of the delays in responding to the…
Study explores how to efficiently explore communities with limited budget.
problem Maximizing the number of members met with limited budget in community exploration.
method Systematic study from offline optimization to online learning, including greedy methods and upper confidence algorithms.
result Achieved logarithmic and constant regret bounds in online learning setting.
A geometric analysis of the time series of returns has been performed in the past and it implied that the most of the systematic information of the market is contained in a space of small dimension. Here we have explored subspaces of this space to find out the relative performance of portfolios formed from the companie…
This review examines deep learning in financial fraud detection over 5 years.
problem Improving deep learning techniques for financial fraud detection.
method Systematic literature review of 57 studies using performance metrics.
result Deep learning models enhance fraud detection across various financial domains.
New L 1 L_1 L 1 -Coverage objective simplifies exploration in reinforcement learning.
problem Challenges in exploration for high-dimensional domains.
method Introduces L 1 L_1 L 1 -Coverage objective to enable efficient exploration and planning. result First computationally efficient algorithms for online reinforcement learning with low coverability.
Machine learning reduces workload in healthcare systematic reviews by 70%.
problem Efficiently screening abstracts for systematic reviews in healthcare.
method Training SVM classifiers on labeled abstracts to classify RCTs.
result SVM classifier achieved 90% accuracy and 0.84 F1 score.
Paper explores zero-shot cross-lingual reading comprehension using pre-trained multi-lingual model.
problem Lack of training data for every language in reading comprehension tasks.
method Systematic exploration of zero-shot cross-lingual transfer learning with a multi-lingual language representation model.
result Zero-shot cross-lingual transfer learning is feasible and translating source data into target language is not necessary.
Complex network theory models China's credit system to control systemic risk.
problem Insufficient understanding of China's credit network structure during financial crises.
method Constructed bipartite financial institution-firm network and analyzed its typological properties.
result Credit network structure can amplify local risks to the whole economy.
Deep learning improves handwriting style transfer and extraction.
problem Improving handwriting style transfer and extraction using deep neural networks.
method Used a deep conditioned autoencoder on IRON-OFF handwriting data-set to explore style transfer and extraction.
result Improved metrics of state-of-the-art methods by a large margin in style transfer and extraction experiments.
We develop a systematic method for classifying supersymmetric orbifold compactifications of M-theory. By restricting our attention to abelian orbifolds with low order, in the special cases where elements do not include coordinate shifts, we construct a "periodic table" of such compactifications, organized according to …
Paper predicts stock volatility using ESG news, showing deep learning's effectiveness.
problem Predicting stock volatility using ESG news.
method ESG news extraction, news representations, and Bayesian inference of deep learning models.
result Deep learning models predict stock volatility better than traditional methods.
The paper explores sampling problems and shows minimal exploration is needed.
problem Exploration-exploitation trade-off in sampling.
method Systematic definition of regret, proposal of a simple algorithm.
result Near-optimal regret bounds achieved with minimal exploration.
Statistical modeling of nuclear data provides a novel approach to nuclear systematics complementary to established theoretical and phenomenological approaches based on quantum theory. Continuing previous studies in which global statistical modeling is pursued within the general framework of machine learning theory, we …
New technique learns STL formulas for classifying time-series data.
problem Lack of interpretability in traditional machine learning for time-series data.
method Systematically explores the space of all STL formulas, prunes the space, and learns semantically equivalent formulas.
result Automatically learns STL formulas for clustering and classifying real-valued time-series data.
This paper surveys scalable automated alignment methods for LLMs.
problem Scalability issues in traditional human-annotated alignment methods for LLMs.
method Categorizes and discusses various automated alignment methods.
result Emerging automated alignment methods are effective and scalable.
Improves RLHF sample efficiency by scaling reward complexity polynomially.
problem Exponential sample complexity in RLHF algorithms for skewed preferences.
method SE-POPO, an online RLHF algorithm that achieves polynomial sample complexity.
result SE-POPO outperforms existing algorithms in sample efficiency.
Study evaluates three position sizing methods for put-writing on S&P 500 Index options.
problem Underdeveloped practical implementation of short-dated volatility-selling strategies.
method Kelly criterion, VIX-based volatility scaling, hybrid method.
result Ultra-short-dated, out-of-the-money options deliver superior risk-adjusted returns.
Automates feature engineering for predictive models using reinforcement learning.
problem Lack of a well-defined basis for effective feature engineering.
method Performance-driven exploration of a transformation graph using reinforcement learning.
result Automated feature engineering reduces human intervention and costs.
Generative model calibrates 3D battery cathode morphologies from 2D images.
problem Calibrate 3D morphologies of all-solid-state battery cathodes from 2D microscopy images.
method Combining GANs with excursion sets of Gaussian random fields.
result Calibrated digital twins enable systematic exploration of morphological scenarios.
Bayesian quadrature uses probabilistic models for estimating intractable integrals.
problem Estimating intractable integrals in complex models.
method Probabilistic, model-based approach using Gaussian processes.
result Comprehensive review and systematic taxonomy of Bayesian quadrature methods.
Improved regret bounds for structured linear contextual bandits with Gaussian noise.
problem Optimizing bandit learning algorithms for structured contexts with Gaussian perturbations.
method Proposed simple greedy algorithms for structured linear contextual bandits with Gaussian noise.
result Unified regret analysis for structured parameters with geometric quantities as bounds.
This paper classifies stablecoin designs to mitigate volatility risks.
problem Mitigating the volatility of stablecoins.
method Systematic design classification, component analysis, and future direction identification.
result Identified strengths and drawbacks of existing stablecoin designs.
Quantum codes on hyperbolic lattices outperform Euclidean ones with higher rates and lower overhead.
problem Improving quantum error correction performance with hyperbolic lattices.
method Unified framework using Hyperbolic Cycle Basis algorithm for CSS codes construction and benchmarking.
result Achieved higher encoding rates and lower qubit overhead in hyperbolic quantum error correction codes.
Study uses supercomputers to improve financial predictions.
problem Improving financial predictions through better exploration of data.
method Refactored and ran algorithm on Fugaku supercomputer, exploring more rules.
result Increasing the number of explored rules improves predictive performance.
Systematically compares methods for explaining RNN predictions.
problem Understanding the decisions of RNNs, particularly LSTMs, through relevance assignments.
method Systematic comparison of methods in various settings.
result Best method reveals linguistic phenomena in sentiment analysis.
New formulas forecast fractional Brownian motion for financial trading.
problem Forecasting financial log-prices following fractional Brownian motion.
method Theoretical formulas for accuracy metrics in fBm forecasting.
result Optimal trading strategies in fBm framework identified.
Paper explores using 0-surgery to find exotic 4-manifolds.
problem Distinguishing smooth structures on 4-manifolds.
method Uses 0-surgery homeomorphisms to generate exotic pairs.
result Found 5 topologically slice knots potentially leading to exotic spheres.
Unified framework improves NLP tasks by converting diverse problems into text-to-text format.
problem Improving natural language processing tasks through transfer learning.
method Unified text-to-text transformer framework, comparing various pre-training objectives and architectures.
result Achieved state-of-the-art results on multiple NLP benchmarks.
Safe linear stochastic bandits ensure safe exploration with optimal regret.
problem Ensuring safe exploration in stochastic bandits while minimizing regret.
method Combining known safe arms with exploratory arms to safely expand the set of safe arms over time.
result The algorithm achieves an expected regret of O ( T log ( T ) ) O(\sqrt{T}\log (T)) O ( T log ( T )) . Agent learns to trade currency pairs with improved risk management.
problem Improving systematic FX trading performance with online transfer learning.
method Online inductive transfer learning using feature representation from Gaussian mixture model to a reinforcement learning agent.
result Annualized portfolio information ratio of 0.52, compound return of 9.3%.
Paper explores anchoring for vision models, improving generalization and safety.
problem Anchoring can lead to undesirable shortcuts, limiting generalization.
method Introduced a new anchored training protocol with a regularizer to mitigate undesirable shortcuts.
result Significant performance gains in generalization and safety metrics.
We study the effects of non-systematic and systematic mortality risks on the required initial capital in a pension plan, in the presence of financial risks. We discover that for a pension plan with few members the impact of pooling on the required capital per person is strong, but non-systematic risk diminishes rapidly…
This paper collects a large dataset for conversational recommendation research.
problem Creating dialogue systems for conversational recommendation.
method Collects a large dataset (ReDial) and explores neural architectures, mechanisms, and methods for conversational recommendation systems.
result Demonstrates the utility of the collected dataset for systematic probing of model sub-components in conversational recommendation.
Deep learning reconstructs ultra-short laser pulses.
problem Characterization of ultra-short laser pulses (femtosecond to attosecond).
method Deep neural network technique.
result First deep learning method to reconstruct ultra-short optical pulses.
This review analyzes recent advances in solving index tracking problems.
problem Creating a portfolio that closely follows a specific index with lower costs.
method Systematic review of mathematical approaches and metaheuristics.
result Metaheuristics have been extensively applied and improved in solving index tracking problems.
Investigates how neural network capacity impacts the learning of compositional languages.
problem The importance of communicative bandwidth in emergent language learning.
method Exploration of how neural network capacity affects the learning of compositional languages.
result There is a specific range of model capacity and channel bandwidth that induces compositional structure.
NetHack Learning Environment (NLE) tests RL algorithms, offering scalable, complex, and challenging gameplay.
problem Challenging environments for testing RL algorithms.
method Procedurally generated, stochastic, rich, and complex NetHack environment.
result Demonstrates empirical success for early stages of NetHack using RL.
MCLNN improves music genre classification with automated feature exploration.
problem Music genre classification using neural networks.
method MCLNN uses a mask to enforce sparseness and learn time-frequency representations.
result MCLNN achieves competitive accuracy compared to state-of-the-art methods.
Bayesian DOE accelerates experimental design with improved efficiency.
problem Enhancing experimental design efficiency and reliability.
method Bayesian framework, conditional density estimation, informative data selection.
result Significantly improved computational efficiency of experimental design.
Detects systematic anomalies in consumer complaints using NLP.
problem Detecting small, frequent anomalies in consumer complaints.
method NLP conversion of narratives, followed by anomaly detection algorithm.
result Demonstrates effectiveness of NLP for detecting systematic anomalies.
This study explores complex structures on Lie algebras from graph perspectives.
problem Existence and characterization of complex structures on 2-step nilpotent Lie algebras.
method Introducing adapted complex structures and analyzing integrability conditions.
result Characterization of graphs that admit abelian adapted complex structures and unique invariant subgraphs.
Neural nets analyze crypto markets for multi-timeframe trading.
problem High-frequency trading in cryptocurrency markets.
method Multi-timeframe trend analysis and high-frequency direction prediction networks.
result Positive risk-adjusted returns through machine learning.
Neural network learns its size and structure during training.
problem Adapting neural network architecture to specific datasets.
method Flexible setup allowing neural network to learn size and topology during training.
result Trained networks achieve virtually identical performance and have learned optimal structure.
This paper analyzes MORL and proposes efficient algorithms to learn Pareto optimal policies.
problem Understanding and efficiently learning Pareto optimal policies in multi-objective reinforcement learning.
method Systematic analysis of optimization targets, reformulation of Tchebycheff scalarization, online UCB-based algorithm, preference-free framework.
result Identification of Tchebycheff scalarization as a favorable method and efficient algorithms for learning Pareto optimal policies.
Survey explores how transfer learning improves deep reinforcement learning.
problem Challenges in reinforcement learning efficiency and effectiveness.
method Categorizes and analyzes transfer learning approaches.
result Transfer learning enhances reinforcement learning performance.