Improved modeling of persistence diagrams for data analysis.
problem Determining significant outliers in persistence diagrams.
method Modification of the RST (Replicating Statistical Topology) model using MCMC Metropolis-Hastings algorithm.
result The modified RST model improves the goodness of fit in persistence diagram analysis.
The paper classifies self-replicating 3D shapes using algebraic models.
problem Understanding self-replicating 3D shapes.
method Using idempotents in the (2+1)-cobordism category to classify 3-manifolds.
result A classification theorem for self-replicating 3-manifolds.
Study replicability in high-dimensional statistics, resolving open problems.
problem Ensuring consistent results in high-dimensional statistical tasks.
method Introduced replicable learning algorithms and established computational and statistical equivalence with high-dimensional isoperimetric tilings.
result Matching sample complexity upper and lower bounds for replicable mean estimation and coin problem.
Study reveals statistical bias in dataset replication, reducing accuracy drop from 11-14% to 3.6%.
problem Statistical bias in dataset replication affects model generalization accuracy.
method Analyzed ImageNet-v2, identified and corrected for bias, and compared results.
result Correcting bias reduces accuracy drop from 11-14% to 3.6%.
Replicable clustering algorithms for k-medians, k-means, and k-centers are proposed.
problem Designing clustering algorithms that produce the same partition on repeated runs under the same distribution.
method Utilizing approximation routines for combinatorial clustering problems in a black-box manner.
result Replicable algorithms for statistical k-medians, k-means, and k-centers with specified approximation and sample complexities. New study on replicability and stability in machine learning algorithms.
problem Ensuring consistent results in machine learning models without fixing randomness.
method Introduced global stability and list replicability concepts, proving their equivalence and boosting list replicability.
result Global stability can only be achieved weakly, while list replicability can be boosted to achieve high probability of consistent results.
Study on computational aspects of replicable learning, bridging statistical and algorithmic perspectives.
problem Understanding the computational connections between replicability and various learning paradigms.
method Design of replicable learners, lifting framework, and transformation techniques.
result Efficient replicable learners for specific learning problems under various distributions.
The Ising model replicates financial asset statistical features.
problem Replicating statistical characteristics of financial markets.
method Employed the Ising model and Monte Carlo simulations.
result The Ising model can replicate most financial asset statistical features.
Researchers analyze a new neural network training method.
problem Training robust configurations in discrete weight neural networks.
method Replicated simulated annealing combining physics and classical simulated annealing.
result Explicit criteria for algorithm convergence and successful sampling.
Study detects P-type bifurcations in single system realizations using unreliable kernel density estimates.
problem Detecting P-type bifurcations in signals with unreliable kernel density estimates.
method Create persistence diagrams from single system realization, statistically analyze resulting set, compare point process modeling methods.
result Subsampling outperforms other point process modeling methods in predicting P-type bifurcations.
ERICA assesses replicability of cluster analysis results.
problem Lack of quantitative scrutiny for clustering results.
method ERICA: a framework to assess replicability of cluster analysis.
result Clusters are found to be replicable in synthetic data but not in real-world datasets.
Recent advances in smart cities applications enforce security threads such as node replication attacks. Such attack is take place when the attacker plants a replicated network node within the network. Vehicular Ad hoc networks are connecting sensors that have limited resources and required the response time to be as lo…
Study on replicability in reinforcement learning algorithms.
problem Ensuring consistent policy outputs in reinforcement learning.
method Mathematical study focusing on replicability in discounted tabular MDPs with a generative model.
result Design of efficient replicable and TV indistinguishable algorithms for policy estimation.
A family of replicator-like dynamics, called the escort replicator equation, is constructed using information-geometric concepts and generalized information entropies and diverenges from statistical thermodynamics. Lyapunov functions and escort generalizations of basic concepts and constructions in evolutionary game th…
New forecasting framework sktime replicates and improves M4 study results.
problem Improving univariate forecasting performance using simple machine learning approaches.
method Designing and implementing a new forecasting API in sktime, using it to replicate and extend M4 study results.
result Simple hybrid and pure approaches can boost statistical model performance and achieve competitive results on hourly data.
ARF synthesizes epidemiological data to match original findings.
problem Synthetic data quality and privacy in epidemiology.
method Adversarial Random Forests (ARF) for efficient data synthesis.
result ARF-generated synthetic data consistently matches original epidemiological findings.
Corporate bond factor research is flawed due to measurement errors and ex-post filtering.
problem Replication crisis in corporate bond factor research.
method Analysis of 108 signals across nine thematic clusters, correction of transaction prices and return filtering.
result Majority of previously documented factors do not produce statistically significant alphas after correction.
The paper improves generative models to avoid replicating observed examples.
problem Improving generative models to avoid replicating observed examples.
method Theoretical insights into the Wasserstein GAN, constrained to left-invertible push-forward maps, generating distributions that avoid replication and significantly deviate from the empirical distribution.
result Left-invertibility achieves this without compromising statistical optimality.
ERICA assesses reproducibility in cluster analysis.
problem Lack of a unified framework for evaluating cluster analysis replicability.
method ERICA (iterative clustering assignments) method to quantify replicability.
result Demonstrates ERICA's ability to identify reproducible cluster structure.
We present an agent based model of a single asset financial market that is capable of replicating several non-trivial statistical properties observed in real financial markets, generically referred to as stylized facts. While previous models reported in the literature are also capable of replicating some of these stati…
New algorithm ensures replicable results in multi-armed bandits with minimal extra regret.
problem Ensuring consistent results in multi-armed bandit studies.
method Incorporates randomness into decision-making to ensure replicability while maintaining minimal extra regret.
result For large time horizons, proposed algorithm suffers only K2/ρ2 times smaller amount of exploration than existing algorithms. Formulates superhedging under costs and uncertainty for continuous assets.
problem Superhedging with transaction costs and model uncertainty for continuous processes.
method New topological framework for continuous asset prices with parametric model uncertainty.
result Formulates a superhedging theorem in the presence of transaction costs and model uncertainty.
This paper investigates whether the gravity model (GM) can explain the statistical properties of the International Trade Network (ITN). We fit data on international-trade flows with a GM specification using alternative fitting techniques and we employ GM estimates to build a weighted predicted ITN, whose topological pr…
Resampling techniques are widely used in statistical inference and ensemble learning, in which estimators' statistical properties are essential. However, existing methods are computationally demanding, because repetitions of estimation/learning via numerical optimization/integral for each resampled data are required. I…
A semi-static approach efficiently replicates and prices callable interest rate derivatives.
problem Efficiently replicating and pricing callable interest rate derivatives under dynamic market conditions.
method Proposes a semi-static hedging algorithm that updates the replication portfolio on a finite number of instances, rather than continuously.
result The hedging error can be made arbitrarily small with a sufficiently large replication portfolio, and closed-form error margins are determined.
Study uses deep learning for pairs trading in Polish equities, achieving profits in 2017-2019.
problem Statistical arbitrage in Polish equities market using traditional methods.
method Deep learning (LSTMs) for asset replication, PCA for risk factor analysis, Ornstein Uhlenbeck process for residual modeling.
result Deep learning methods, especially LSTMs, show promise for profitable trading in Polish equities.
Machine learning helps estimate risk premiums of stocks without knowing their factors.
problem Estimate risk premiums of stocks without knowing their underlying factors.
method Used elastic-net machine learning to project stock returns onto peers and construct replicate portfolios.
result Unique stocks have higher SARP and excess returns than ubiquitous stocks.
Consider a financial market in which an agent trades with utility-induced restrictions on wealth. For a utility function which satisfies the condition of reasonable asymptotic elasticity at −∞ we prove that the utility-based super-replication price of an unbounded (but sufficiently integrable) contingent claim i…
Randomness is crucial for stability in learning and statistics, especially for differential privacy.
problem Quantifying the amount of randomness needed for algorithmic stability.
method Weak-to-strong boosting theorem for stability, characterizing randomness complexity of PAC Learning.
result Randomness complexity is tightly controlled by the best replication probability of any deterministic algorithm solving the task.
Replicates deep learning strategy for trading factor residuals, finds strong performance.
problem Exploiting mis-pricing from unexplained cross-sectional variation in factor models.
method Adhering to PIT principles, used CNNs and Transformers on recent data.
result Out-of-sample Sharpe ratios exceeding 10 in certain tests.
In this study, a novel topology optimization approach based on conditional Wasserstein generative adversarial networks (CWGAN) is developed to replicate the conventional topology optimization algorithms in an extremely computationally inexpensive way. CWGAN consists of a generator and a discriminator, both of which are…
The paper examines how to test if two learning algorithms produce similar outcomes.
problem Testing if two learning algorithms produce similar outcomes when trained on different data sets.
method Using Total Variation (TV) distance to measure similarity of posterior distributions.
result TV indistinguishable learning rules are equivalent to existing stability notions and can be statistically amplified.
Assessing generative models is not an easy task. Generative models should synthesize graphs which are not replicates of real networks but show topological features similar to real graphs. We introduce an approach for assessing graph generative models using graph classifiers. The inability of an established graph classi…
Replication study shows Deep-SE still not as effective as previously thought for agile effort estimation.
problem Improving accuracy in estimating agile software development effort.
method Close replication of Deep-SE using additional data and comparison with multiple baselines.
result Deep-SE outperforms only a few cases, suggesting more work is needed.
The paper introduces a new method to improve model generalization by routing model copies through permutations.
problem Improving model generalization in machine learning.
method The method replicates a model \(M\) times and rewire the contexts in which local learning messages are computed using permutations.
result The method improves generalization by structured message sharing rather than coupling parameters.
New algorithm ensures consistent results in constrained MAB problems.
problem Achieving consistent results in constrained MAB problems.
method Developed replicable algorithms for constrained MAB problems using the optimism principle.
result Regret and constraint violation of replicable algorithms match those of non-replicable ones.
Unified framework for fixed-income pricing and liability replication.
problem Static arbitrage and discount curve construction.
method Model-free framework for static fixed-income pricing and liability replication.
result Existence of strictly positive discount curves reproducing market prices and least-cost super-replicating portfolios.
Characterizes super-replication prices in a financial market model.
problem Characterizing prices in a financial market model.
method Characterizes prices as the supremum of mono-prior super-replication prices through extreme priors and martingale measures.
result Super-replication prices are the supremum of mono-prior super-replication prices.
Bayesian models can be tricked into believing false data.
problem Vulnerability of Bayesian inference to data poisoning attacks.
method Developed attacks to manipulate Bayesian posterior through deletion and replication of data.
result Demonstrated that Bayesian inference can be steered to target distributions.
New algorithm prevents strategic replication in multi-armed bandit problems.
problem Strategic replication by agents can exploit bandit algorithms' balance.
method Designs Hierarchical UCB (H-UCB) and Robust Hierarchical UCB (RH-UCB) algorithms.
result Achieves O(lnT)-regret and sublinear regret in realistic scenarios. Extends super-replication theorem with dynamic strategies and transaction costs.
problem Dynamic super-replication under proportional transaction costs.
method Generalizes admissible strategies and defines a well-defined super-replication price process.
result Well-defined super-replication price process in dynamic setting.
New uniformity tester ensures consistent results across different samples.
problem Non-replicable behavior of uniformity testing algorithms.
method Develops a replicable uniformity tester with improved sample complexity.
result Achieves nearly linear dependence on replicability factor ρ. I review few conceptual steps in analytic description of topological interactions, which constitute the basis of a new interdisciplinary branch in mathematical physics, "Statistical Topology", emerged at the edge of topology and statistical physics of fluctuating non-phantom rope-like objects. This new branch is called…
In this work we introduce the notion of fully incomplete markets. We prove that for these markets the super-replication price coincide with the model free super-replication price. Namely, the knowledge of the model does not reduce the super-replication price. We provide two families of fully incomplete models: stochast…
By the classical Martingale Representation Theorem, replication of random vectors can be achieved via stochastic integrals or solutions of stochastic differential equations. We introduce a new approach to replication of random vectors via adapted differentiable processes generated by a controlled ordinary differential …
We study super--replication of contingent claims in markets with fixed transaction costs. This can be viewed as a stochastic impulse control problem with a terminal state constraint. The first result in this paper reveals that in reasonable continuous time financial market models the super--replication price is prohibi…
Optimizing expensive black-box systems with limited data is an extremely challenging problem. As a resolution, we present a new surrogate optimization approach by addressing two gaps in prior research -- unimportant input variables and inefficient treatment of uncertainty associated with the black-box output. We first …
Adaptive replication improves stochastic function optimization.
problem Challenges in accurately estimating functions with high variance.
method Trust-region-based Bayesian optimization with adaptive replication.
result Adaptive replication substantially improves solution accuracy and efficiency.