Logit dynamics formula reveals self-regulation in softmax policy gradient methods.
problem Understanding the stability and convergence of softmax policy gradient methods.
method Deriving the exact formula for the L2 norm of the logit update vector.
result Logit update magnitudes are modulated by action probability and policy concentration.
Policy gradient methods converge globally and efficiently for linear quadratic regulators.
problem Policy gradient methods struggle with non-convex optimization in linear quadratic regulators.
method Model-free policy gradient methods are shown to globally converge and have polynomial complexity in sample and computational terms.
result Policy gradient methods globally converge to the optimal solution for linear quadratic regulators with polynomial efficiency.
Study shows a linear quadratic regulator's imitation learning converges globally.
problem Global convergence of imitation learning for linear quadratic regulators.
method Analyzed alternating gradient algorithm and established Q-linear rate of convergence.
result Established a unique saddle point for globally optimal policy and reward function.
New method makes neural networks more resistant to adversarial attacks.
problem Neural networks are vulnerable to adversarial examples.
method Regulates adversarial gradients to increase robustness.
result Generated networks are near-immune to various adversarial attacks.
Policy gradient methods converge for LQR problems with noisy state dynamics.
problem Finding optimal policies in noisy LQR problems over finite time horizons.
method Policy gradient methods with convergence guarantees for finite time and stochastic state dynamics.
result Global linear convergence for policy gradient methods in LQR problems with weak assumptions.
New model-free algorithm achieves similar LQR regret guarantees.
problem Model-free control of linear dynamical systems under quadratic costs.
method Online policy gradient scheme with policy space cost analysis.
result Achieves regret scaling with √T, matching model-based methods.
Policy gradient converges to globally optimal policy in nearly linear-quadratic systems.
problem Finding optimal policies in nonlinear control systems with partial information.
method Policy gradient algorithm designed for nearly linear-quadratic regulators with small Lipschitz nonlinear components.
result Policy gradient algorithm converges to globally optimal policy with linear rate.
URN neural network dynamically generates various neural structures during training.
problem Creating neural networks with flexible, dynamic structures during training.
method Introduced Unstructured Recursive Network (URN) and used gradient descent on a single loss function.
result Different neural structures can emerge from a single URN during training.
The Conant-Ashby theorem is verified for hypergraph observers, leading to unique learning rules.
problem Verifying conditions for hypergraph observers to maintain internal models.
method Formalizing persistent observers, applying the Conant-Ashby theorem, and using natural gradient descent.
result Natural gradient descent is the unique admissible learning rule for hypergraph observers.
Study on variance of policy gradient in simple RL environments.
problem Understanding variance of policy gradient estimators in continuous RL.
method Analyzes REINFORCE estimator in linear-quadratic environments with Gaussian noise.
result Derives and validates bounds on estimator variance empirically.
EAGC boosts GCD by regulating gradient entanglement, improving known and novel category separability.
problem Gradient entanglement distorts supervised gradients and overlaps known and novel class representations.
method EAGC uses AGA and EEP to align and project gradients, reducing entanglement and overlap.
result EAGC consistently boosts GCD performance, setting new state-of-the-art results.
This paper studies how gradient descent in control systems can perform well on unseen data.
problem The extent of a learned controller's ability to extrapolate to unseen initial states.
method Theoretical study of policy gradient in Linear Quadratic Regulator (LQR) problems, focusing on the role of exploration.
result The performance of a learned controller on unseen initial states depends on the degree of exploration induced by the system.
IEBN normalizes noise by enhancing instance-specific information, improving deep learning performance.
problem Improving deep learning performance by regulating noise in batch normalization.
method Integrates self-attention mechanism to recalibrate channel information in BN.
result IEBN outperforms BN with improved generalization and stability.
Perceptor Gradients learns symbolic representations from raw data.
problem Learning transferable symbolic representations from raw data.
method Decomposes policy into perceptor network and task encoding program.
result Efficiently learns symbolic representations for control tasks.
Attention augments forest for tabular data accuracy.
problem Training tabular data models with high accuracy and efficiency.
method Tree Attention Block (TAB) in differentiable forest framework.
result Attention augmented differentiable forest achieves comparable and sometimes higher accuracy than GBDT models.
Optimal control methods achieve significantly smaller regret than previously thought.
problem Optimal control in linear dynamical systems with adversarial changes.
method Online gradient descent and online natural gradient methods.
result Achieves logarithmic regret scaling as O(poly(log T)) instead of O(sqrt(T)).
Paper develops risk statistics for portfolios considering regulator-based risk.
problem Traditional risk statistics fail to describe regulator-based risk.
method Develop dual representation for regulator-based risk statistics.
result Derived dual representation for regulator-based risk statistics.
Proposes a method to find traffic regulations efficiently.
problem Finding appropriate traffic regulations in congested events.
method Graph Convolutional Networks for modeling regulation effects.
result The method can find a road to close that reduces travel time.
Model proposes how regulators should oversee complex algorithms in high-stakes applications.
problem Regulating complex algorithms used in high-stakes applications like lending, testing, and hiring.
method Proposes a model where regulators are limited in learning about complex algorithms with misaligned preferences, and explores different regulatory approaches.
result Complex algorithms can improve welfare, but regulation should focus on the source of incentive misalignment for optimal results.
Regulated curves on Banach manifolds with continuous projections and regulated derivatives are studied.
problem Regulated curves on Banach manifolds with continuous projections and regulated derivatives.
method Building a Banach manifold structure on the set of such curves.
result Existence of a 'local addition' on such a manifold for any Banach manifold.
We show that any objective risk measurement algorithm mandated by central banks for regulated financial entities will result in more risk being taken on by those financial entities than would otherwise be the case. Furthermore, the risks taken on by the regulated financial entities are far more systemically concentrate…
New mechanism designs regulate herding in financial markets.
problem Herding causes irrational market decisions and volatility.
method A trilateral game framework based on optimal control theory.
result Effective mechanisms improve social welfare.
In a market system, regulations are designed to prevent or rectify market failures that inhibit fair exchange, such as monopoly or transactions with hidden costs. Because regulations reduce profits to those possessing unfair advantage, these advantaged corporations (whether individuals, companies, or other collective o…
Hyperboost uses gradient boosting for hyperparameter optimization, outperforming state-of-the-art methods.
problem Hyperparameter tuning for machine learning algorithms
method Gradient boosting surrogate model with quantile regression and distance metric
result Hyperboost outperforms state-of-the-art techniques in empirical tests
MiCA regulation led to a shift in stablecoin dominance.
problem Impact of MiCA regulation on stablecoin trading.
method Comparative analysis of regulated and non-regulated exchanges.
result USDC gained market share and trading volume post-MiCA regulation.
New risk statistics for loss-based regulation.
problem Regulatory focus on losses over gains.
method Developed new risk statistics using scenario analysis.
result New risk statistics extend existing measures.
Optimal insurance investment under VaR regulation improves policyholders' utility.
problem Optimal investment for participating insurance contracts under VaR-regulation.
method Martingale approach for constrained non-concave optimization problems.
result VaR constraints lead to more prudent investment, improving policyholders' utility.
We use ellipsoids to solve power system voltage regulation problems.
problem Voltage regulation in power systems under uncertainty.
method Tractable ellipsoidal approximation for chance constrained optimizations.
result Efficiently trained machine learning model approximates uncertainty region.
Study shows model-based methods require fewer samples than model-free methods for LQR tasks.
problem Comparing model-based and model-free methods in reinforcement learning for continuous control tasks.
method An asymptotic analysis of sample complexity for policy evaluation in LQR tasks.
result Model-based methods require asymptotically less samples than model-free methods for policy evaluation in LQR tasks.
Proposes a game-theoretic framework for ML trust regulation.
problem Lack of coordination between ML model builders and regulators.
method Formulates trustworthy ML as a multi-objective multi-agent optimization problem and introduces regulation games and ParetoPlay.
result Enables efficient enforcement of ML model specifications without discouraging participation.
The FCA improved insider trading regulation after 2012, reducing abnormal returns.
problem Regulation of insider trading before and after the UK Financial Services Act 2012.
method Event study methodology using abnormal returns analysis.
result Abnormal returns were reduced after the FCA took over from the FSA.
A deterministic trading strategy by a representative investor on a single market asset, which generates complex and realistic returns with its first four moments similar to the empirical values of European stock indices, is used to simulate the effects of financial regulation that either pricks bubbles, props up crashe…
Proposes guidelines for developing medical AI products.
problem Lack of clear pathways for regulating medical AI.
method Statistical risk perspective and deep understanding of machine learning methodologies.
result Enhanced development of medical AI products and regulations.
Biophysical models explain deep learning in gene regulation.
problem Difficulty in interpreting deep learning models in gene regulation.
method Expressed biophysical models as neural networks with explicit interpretations.
result Biophysical networks can be inferred from MPRAs.
An asset network systemic risk (ANWSER) model is presented to investigate the impact of how shadow banks are intermingled in a financial system on the severity of financial contagion. Particularly, the focus of this study is the impact of the following three representative topologies of an interbank loan network betwee…
Modeling pollution from competing firms using mean-field games.
problem Pollution regulation of competitive firms producing similar goods.
method Developed a mean-field game model with cap-and-trade regulation.
result Explicit solutions found through Riccati differential equations.
Introduces gradient decay in Softmax for better generalization.
problem Improving generalization performance in neural networks.
method Gradient decay hyperparameter in Softmax for varying gradient rates based on probability.
result Gradient decay rate affects generalization performance and can be tuned for better optimization.
Develops new methods for isospectral orbifolds and regulator quotients.
problem Isospectral orbifolds and regulator quotients in Vignéras constructions.
method New sufficient criteria for isospectrality and regulator quotients, linking torsion homology and Galois representations.
result Produces small exotic isospectral orbifolds and sufficient criteria for regulator quotients.
Modern physics has demonstrated that matter behaves very differently as it approaches the speed of light. This paper explores the implications of modern physics to the operation and regulation of financial markets. Information cannot move faster than the speed of light. The geographic separation of market centers means…
A neural network method tackles high-dimensional diffeomorphic mapping problems.
problem High-dimensional diffeomorphic mapping struggles with the curse of dimensionality.
method Combines variational principles with quasi-conformal theory for accurate, bijective mappings.
result Validated accuracy, robustness, and effectiveness in complex registration scenarios.
This study examines how ChiNext IPOs' initial returns are influenced by regulation regime changes.
problem Investors' behavior and pricing of ChiNext IPOs under different regulation regimes.
method Analysis of three time periods with two different regulation regimes and three sets of listing day trading restrictions.
result Regulation regime changes significantly impact ChiNext IPO pricing and overreaction.
Regulated Bitcoin futures led to higher volatility and trading volume.
problem Estimating the impact of regulated Bitcoin futures on volatility and volume.
method Employed a new causal approach, C-ARIMA.
result Regulated Bitcoin futures increased Bitcoin volatility by more than double.
This paper tackles robust control of LQR systems with multiplicative noise using policy gradient methods.
problem Robustness in reinforcement learning control of complex systems with multiplicative noise.
method Policy gradient algorithms with gradient domination property for non-convex cost functions.
result Global convergence of policy gradient algorithms to the globally optimum control policy.
We show that the regulator, which is the difference between the homology torsion and the combinatorial Ray-Singer torsion, of fnite abelian coverings of a fixed complex has sub-exponential growth rate.
We investigate a randomization procedure undertaken in real option games which can serve as a basic model of regulation in a duopoly model of preemptive investment. We recall the rigorous framework of [M. Grasselli, V. Leclère and M. Ludkovsky, Priority Option: the value of being a leader, International Journal of Theo…
ReF-ER improves RL by selectively updating experiences based on policy similarity.
problem Deterioration of experience replay accuracy when policies diverge.
method Enforces policy similarity in replay memory by skipping unlikely experiences and regulating policy changes.
result ReF-ER consistently improves RL performance across various benchmarks.
Study optimal liquidation strategies in lit and dark pools with and without regulation.
problem Optimal liquidation strategies in dark and lit pools with execution uncertainty.
method Design optimal make-take fee policies, solve HJB-Fokker-Planck systems, use BSDEs.
result Explicit solutions for optimal strategies in both competitive and regulated markets.
Improved logistic models for better interpretability in regulated industries.
problem Limited interpretability in machine learning models for regulated industries.
method Blend logistic regression with machine learning techniques for enhanced interpretability.
result Solid performance with minimal analyst effort and enhanced interpretability.