Subcartesian spaces follow Leibniz' rule, simplifying differential calculus.
problem Simplifying differential calculus in subcartesian spaces.
method Showed derivations satisfy the chain rule and have maximal integral curves.
result Subcartesian spaces follow Leibniz' rule, simplifying differential calculus.
Paper explores subdifferential chain rules for matrix factorization and related machine learning models.
problem Clarke subdifferential chain rules for matrix factorization and factorization machines.
method Analyzes conditions for subdifferential chain rules to hold, especially for overparameterized models.
result Subdifferential chain rules hold for matrix factorization and factorization machines under certain conditions.
Develops a chain rule for ReLU networks and extends approximation theory to global error estimates.
problem Applying standard chain rule to ReLU networks and extending approximation results globally.
method Introduces a derivative for ReLU networks and converts bounded domain results to global estimates.
result Extends neural network approximation theory to include regularity properties for ReLU networks.
New method calculates barrier option Greeks using Wiener path integrals.
problem Computing first-order Greeks for barrier options efficiently.
method Developed chain rules for Wiener path integrals.
result Effectiveness demonstrated through numerical examples.
New training method for neural nets using multilevel entropic regularization.
problem Training efficiency and generalization bounds for neural nets.
method Multilevel relative entropy, chaining mutual information, Gibbs posterior distribution.
result Proves the Gibbs posterior achieves the unique minimum of the empirical risk minimization problem.
In many healthcare settings, intuitive decision rules for risk stratification can help effective hospital resource allocation. This paper introduces a novel variant of decision tree algorithms that produces a chain of decisions, not a general tree. Our algorithm, α-Carving Decision Chain (ACDC), sequentially carves o…
Predict and explain service failures in supply-chain networks using data models.
problem Predict and explain service failures in supply-chain networks, particularly last-mile pickup and delivery.
method Used supervised classification with Random Forests and Association Rules on a dataset of 500,000 services.
result Classifier reaches an average sensitivity of 0.7 and specificity of 0.7 for 5 types of failure.
CT compares two distributions using Bayes' theorem and chain rule.
problem Measuring the difference between two probability distributions.
method Conditional transport (CT) using chain rule and Bayes' theorem.
result CT strikes a good balance between mode-covering and mode-seeking behaviors.
Lie systems method simplifies Riccati hierarchy study.
problem Simplifying study of Riccati hierarchy equations.
method Lie systems approach to projective Riccati equations.
result Characterization of Riccati chain equations geometrically.
Method measures weight similarity in neural networks using normalization and statistical inference.
problem Quantifying weight similarity in non-convex neural networks.
method Chain normalization rule and hypothesis-training-testing statistical inference.
result Weights of identical neural networks converge to similar local solutions.
The Widrow-Hoff rule simplifies language data simulation.
problem Simulating language phenomena computationally.
method Implementation and application of the Widrow-Hoff rule.
result The Widrow-Hoff rule offers new perspectives on language simulation.
A conformal procedure improves CoT reasoning by aggregating reasoning paths and calibrating abstention rules.
problem Aggregation uncertainty in chain-of-thought reasoning makes correct answers less reliable.
method Introduces a conformal procedure for CoT reasoning that uses weighted score aggregation and abstention rules.
result Achieves higher selective accuracy with abstention, reducing confident-error rate.
Random Intersection Chains selects important interactions from categorical features.
problem Heavy computational burden in considering all interactions for categorical features.
method Randomly generates chains of intersections, estimates and selects frequent patterns.
result Selected patterns are the most frequent in the data set.
Paper finds optimal selling rule for pairs trading with stock constraints.
problem Identifying the best time to sell in pairs trading of stocks.
method Optimal pairs-trading selling rule with constraints on trading.
result Closed-form solution for optimal policy determined by a threshold curve.
A new stopping rule based on E-values helps efficiently use sampling in Bayesian Deep Ensembles.
problem How long should sampling continue in Bayesian Deep Ensembles to yield significant improvements?
method Formulated as a sequential anytime-valid hypothesis test, using E-values to decide when to stop sampling.
result Only a fraction of the full-chain budget is often required for significant improvements.
This paper is concerned with an optimal stock selling rule under a Markov chain model. The objective is to find an optimal stopping time to sell the stock so as to maximize an expected return. Solutions to the associated variational inequalities are obtained. Closed-form solutions are given in terms of a set of thresho…
Bayesian learning with fuzzy inference for data analysis.
problem Data analysis and knowledge extraction from complex data.
method Rule-based fuzzy inference systems with Bayesian estimation and MCMC.
result Effective for regression and classification tasks, especially in financial services.
A one-to-one correspondence is drawn between law invariant risk measures and divergences, which we define as functionals of pairs of probability measures on arbitrary standard Borel spaces satisfying a few natural properties. Divergences include many classical information divergence measures, such as relative entropy a…
This paper justifies the use of straight-through estimator in training quantized neural nets.
problem Minimizing loss in quantized neural nets with vanishing gradients.
method Introduced straight-through estimator (STE) and proved its effectiveness in two-linear-layer network with binarized ReLU activations.
result Proved that the coarse gradient derived from STE is a descent direction for minimizing population loss.
We propose a new statistical model for computational linguistics. Rather than trying to estimate directly the probability distribution of a random sentence of the language, we define a Markov chain on finite sets of sentences with many finite recurrent communicating classes and define our language model as the invarian…
Study optimal adaptive allocation for multi-armed bandits with Markovian rewards.
problem Optimal adaptive allocation for multi-armed bandits with Markovian rewards.
method Round-robin Kullback-Leibler upper confidence bounds for optimal adaptive allocation.
result Logarithmic dependence of regret on time horizon, asymptotically optimal.
New gradient estimators simplify policy gradient algorithms in reinforcement learning.
problem Challenges in creating effective gradient estimators for reinforcement learning.
method Total derivative rule and graphical models for gradient estimation.
result New gradient estimators lead to improved performance in reinforcement learning.
Enhances ECC for multi-label learning under class imbalance.
problem Class imbalance in multi-label datasets.
method Coupling ECC with random undersampling, and extending ECC with varying binary models per label and chains of different sizes.
result Improves ECC's performance on multi-label datasets using various evaluation metrics.
New framework improves stochastic optimization for variational inference.
problem Improving variational posterior approximations in high-dimensional models.
method Developed a robust stochastic optimization framework using Markov chains.
result Demonstrated improved accuracy and robustness across diverse models.
New Upsilon invariants rule out stable equivalence of knot complexes.
problem Stable equivalence of knot complexes and its invariants.
method Secondary Upsilon invariants defined by Kim and Livingston.
result Relations between Upsilon invariants do not extend to stable equivalence.
The paper develops a stationary-distribution theory for Random Forest ensemble size selection.
problem Determining the optimal number of trees in Random Forests.
method Modeling the ensemble size as a birth-death Markov chain and deriving its stationary distribution.
result The stationary ensemble size B∗ scales as O(ε−2) as ε↓0. This paper analyzes voter coalitions in MakerDAO's decentralized governance.
problem Understanding the governance structure and influence of voter coalitions in DAOs.
method Applied clustering algorithm to voting history of MakerDAO to identify voter coalitions.
result The emergence of a dominant voter coalition signals governance centralization in DAOs.
New Sasaki-Einstein 7-spheres found via Berglund-Hübsch transpose.
problem Finding new Sasaki-Einstein 7-spheres.
method Berglund-Hübsch transpose rule applied to Kähler-Einstein 3-folds.
result 75 new Sasaki-Einstein rational homology 7-spheres discovered.
New method simplifies causal inference with tiered background knowledge.
problem Large equivalence classes of DAGs limit causal information.
method Integrates tiered background knowledge to create 'tiered MPDAGs' with simplified structure.
result Tiered MPDAGs are chain graphs with chordal components, simplifying causal effect estimation.
Knot theory applied to proteins, distinguishing folded linear chains.
problem Classifying proteins as unknots when intra-chain interactions are ignored.
method Developing knot theory for folded linear molecular chains, considering self-bonding, and using Gauss codes and quandles.
result Extended knot theory to distinguish topologies of proteins with intra-chain bonds.
A new Gibbs sampler method speeds up Bayesian inference.
problem Efficient sampling from complex posterior distributions.
method Recycling auxiliary samples within Gibbs estimators.
result Significant improvement in accuracy and computational efficiency.
Optimal sample complexity for autoregressive chain-of-thought learning proven.
problem Determining the minimum number of samples needed for accurate autoregressive chain-of-thought learning.
method Proved upper bound on sample complexity using Daniely-Shalev-Shwartz dimension and roll-out stable parity dimension.
result The sample complexity is bounded by the local next-token class rate, with no dependence on rollout length.
Improves sampling quality in model composition using MH-like acceptance rule for score-based diffusion models.
problem Inability to apply MH corrections in score-based diffusion models for model composition.
method Introduces a novel MH-like acceptance rule based on line integration of the score function.
result Relative improvements similar to energy-based models without explicit energy parameterization.
Upper bound on expected supremum of Bernoulli process.
problem Bounding the supremum of Bernoulli processes.
method Using properties of the index set and function class, extending earlier results on Gaussian processes.
result An upper bound on the expected supremum of a Bernoulli process.
New calculus on spacetimes for nonlinear differential equations.
problem Nonlinear differential equations on metric measure spacetimes.
method Introduces maximal weak subslope and variational calculus.
result Establishes a comparison theorem for nonlinear p-d'Alembertian. Method learns CTMC models from steady-state data, predicting unseen states.
problem Learning CTMC models from aggregate steady-state statistics without sequence examples.
method ∞-SGD, a stochastic gradient descent method that avoids infinite sums.
result Successfully learns CTMC models and predicts unseen states.
LLMs translate natural language trading intents into correct option strategies using a domain-specific language.
problem Challenges in translating natural language trading intents into correct option strategies due to the complexity of option chain data.
method Introduce Option Query Language (OQL) as a domain-specific intermediate representation to abstract option markets into high-level primitives under grammatical rules. Use LLMs as semantic parsers and validate queries by an engine.
result Significantly improves execution accuracy and logical consistency over direct baselines.
ToolChain-CRC addresses the risk-control problem for retrieval-augmented and tool-using agents under drift.
problem Risk-control problem for retrieval-augmented and tool-using agents under drift.
method ToolChain-CRC uses conformal risk-control under exchangeable calibration runs.
result Trajectory-level risk control keeps accepted-trajectory risk below the target.
Paper bridges matching rules and height functions in aperiodic tilings.
problem Relationship between matching rules and height functions in aperiodic tilings.
method Cochain-first framework to establish equivalence between matching rules, Ammann bar continuity, cycle closure of 1-cochains, and height-function existence.
result Unified framework for aperiodic tilings including Penrose and canonical projection tilings.
New framework compares two stochastic learning dynamics in games.
problem Inability to distinguish between different learning rules leading to the same steady-state behavior.
method Developed a framework for comparative analysis of stochastic learning dynamics with different update rules.
result Identified distinct behaviors in the paths to stochastically stable states for LLL and ML.
ISOMORPH creates a digital twin for supply chain logistics, advancing time-series forecasting benchmarks.
problem Lack of public benchmarks for supply chain logistics time-series forecasting.
method Developed a digital twin simulator with interpretable parameters and modular topology, generating datasets and verifying conservation laws.
result Foundation models achieve MASE values exceeding public benchmarks at low-to-moderate horizons, supporting UQ.
This paper constructs an algebra on a 3-torus with specific properties for fluid dynamics.
problem Constructing an algebraic structure on a 3-torus with specific properties.
method Combining combinatorial graded intersection algebra with Sullivan's and Lawrence-Sullivan-Ranade's subcomplexes.
result The construction of an algebra with specific properties on the 3-torus.
Dynamic abstention improves LLM accuracy by selectively terminating unpromising reasoning.
problem LLMs waste compute on incorrect responses, leading to inefficiency.
method Formal reinforcement learning framework with abstention reward parameter.
result Dynamic abstention outperforms natural baselines in selective accuracy.
New method uses diffusions to measure sample quality in multivariate targets.
problem Measuring convergence to multivariate continuous targets.
method Ito diffusions and explicit multivariate Stein factor bounds.
result Established near-linear relationship between diffusion Stein discrepancies and Wasserstein distances.
New GAN formulation addresses mode collapse issue.
problem Mode collapse in GANs.
method Randomized decision rules, empirical Bayes, stochastic gradient MCMC.
result Proposed method converges to Nash equilibrium.
Improved Bayesian inference for neuronal ensemble inference reduces computational cost.
problem Efficient inference of neuronal ensembles from activity data.
method Modified MCMC algorithm with simulated annealing for hyperparameter control.
result Our method reduces computational cost while maintaining or improving inference accuracy.
A framework for differentiability on persistence barcodes.
problem Differentiability on the space of persistence barcodes.
method Lifts to the space of ordered barcodes to compute derivatives.
result A chain rule enabling gradient descent for objective functions.
We define an algebraic/combinatorial object on the front projection Σ of a Legendrian knot called a Morse complex sequence, abbreviated MCS. This object is motivated by the theory of generating families and provides new connections between generating families, normal rulings, and augmentations of the Chekanov-Eliashb…