This paper examines how institutional liquidity affects prediction markets.
problem How institutional liquidity impacts prediction markets and their quality.
method Defines a market-quality lens, separates channels, and uses synthetic microstructure lab.
result Institutional liquidity does not necessarily translate to equal gains for all traders.
We give a complete proof of a propagation theorem of multiplicity-free property from fibers to spaces of global sections for holomorphic vector bundles. The propagation theorem is formalised in three ways, aiming for producing various multiplicity-free theorems in representation theory for both finite and infinite dime…
We propose a data-driven framework for optimizing privacy-preserving data release mechanisms to attain the information-theoretically optimal tradeoff between minimizing distortion of useful data and concealing specific sensitive information. Our approach employs adversarially-trained neural networks to implement random…
This research generates synthetic data streams for handling concept drifts and novel classes.
problem Handling concept drifts and novel classes in dynamic data streams.
method Synthetic data stream generation for both concept drifts and novel classes.
result Demonstrates the effectiveness of unsupervised drift detectors in open set recognition.
Optimal transport framework for density estimation with constraints.
problem Density estimation under expectation constraints.
method Minimizes Wasserstein distance subject to expected value constraints and regularization.
result Framework effectively addresses non-smooth constraints through annealing-like algorithm.
New mass inequalities and proofs for causal variational principles.
problem Proving new mass inequalities for causal variational principles.
method Proved a new inequality for minimizers of causal variational principles and applied it to prove the positive mass theorem.
result Introduced a positive quasilocal mass and proved new mass inequalities.
In this paper, we will describe a concept of a cryptocurrency issuance protocol which supports digital currencies in a Proof-of-Work (< PoW >) like manner. However, the methods assume alternative utilization of assets used for cryptocurrency creation (rather than purchasing electricity necessary for < mining >).
A deep learning framework for survival analysis combining piecewise exponential models.
problem Survival analysis with competing risks, multi-state modeling, and time-varying effects.
method Piecewise exponential models embedded in a neural network.
result Predicted Alzheimer's disease progression using tabular and 3D point cloud data.
Formalizes synthetic differential geometry in Lean.
problem Formalizing synthetic differential geometry in a proof assistant.
method Formalization of synthetic differential geometry with Lean and mathlib.
result Proves a Taylor theorem for functions of several variables.
This review covers learning under concept drift, including detection, understanding, and adaptation.
problem Unforeseeable changes in data distribution over time impact machine learning performance.
method Reviews and analyzes methodologies and techniques for concept drift detection, understanding, and adaptation.
result Establishes a framework for learning under concept drift with three main components.
We address the problem of causal discovery from data, making use of the recently proposed causal modeling framework of modular structural causal models (mSCM) to handle cycles, latent confounders and non-linearities. We introduce σ-connection graphs (σ-CG), a new class of mixed graphs (containing undirected, bidirected…
Neural symbolic processing aims to combine the generalization of logical learning approaches and the performance of neural networks. The Neural Theorem Proving (NTP) model by Rocktaschel et al (2017) learns embeddings for concepts and performs logical unification. While NTP is promising and effective in predicting fact…
A tool predicts BN performance for real-world datasets.
problem Lack of validation for BN results in real-world applications.
method Synthetic datasets and structure learning algorithms to estimate BN performance.
result Automatic recommendations for BN performance based on synthetic data.
New method detects when models influence their own drift in real-time data streams.
problem Models can induce concept drift in real-time data streams.
method CheckerBoard Performative Drift Detection (CB-PDD)
result CB-PDD effectively detects performative drift in real-time data streams.
Cycle-StarNet bridges theory and data by adapting synthetic spectra to observational data.
problem Lack of consistency between theoretical stellar models and observational data.
method Hybrid generative domain adaptation using unsupervised learning on large spectroscopic surveys.
result Improved spectral fitting and reduced gap between synthetic and observational data.
Study extends null distance concept to Lorentzian length spaces for spacetime analysis.
problem Understanding spacetime convergence and topology in Lorentzian geometry.
method Extend null distance concept to Lorentzian length spaces, study Gromov-Hausdorff convergence.
result First results on compatibility of null distance with synthetic curvature bounds in warped product Lorentzian length spaces.
Proof confirms preservation of projective limits in synthetic differential geometry.
problem Prove preservation of projective limits in synthetic differential geometry.
method Detailed proof using synthetic differential geometry and Cahiers topos.
result Projective limits preserved in synthetic differential geometry.
We use the theory of large deviations to study the pricing of investment-grade tranches of synthetic CDO's. In this paper, we consider a simplified model which will allow us to introduce some of the concepts and calculations.
Study uses simulation-based inference to decode brain activity from synthetic stimuli.
problem Reversing the process of brain activity emulation to recover stimuli or their properties.
method Pairing brain emulator with LLMs to learn a probabilistic mapping from brain maps to stimulus parameters.
result LLMs can serve as controllable stimulus generators and parameters can be recovered from brain maps.
Concept-driven OPE reduces variance in off-policy decision evaluation.
problem High variance in off-policy decision evaluation due to limited sample sizes.
method Integrating human-explainable concepts into OPE to reduce variance.
result Concept-based OPE estimators remain unbiased and reduce variance when concepts are known and predefined.
Adaptive sampling detects local concept drift with limited labels.
problem Detecting local concept drift in dynamic environments with scarce labels.
method Combines residual-based exploration and exploitation with EWMA monitoring.
result Superior performance in label efficiency and drift detection accuracy.
Human explanations of high-level decisions are often expressed in terms of key concepts the decisions are based on. In this paper, we study such concept-based explainability for Deep Neural Networks (DNNs). First, we define the notion of completeness, which quantifies how sufficient a particular set of concepts is in e…
Debias concept-based explanations by removing confounding information.
problem Correlation between concepts and confounding features.
method Causal prior graph and two-stage regression technique.
result Success in removing biases and improving concept ranking.
ProteuS generates synthetic financial data with regime changes for testing drift detection.
problem Simulating concept drift in financial markets for model evaluation.
method ARMA-GARCH models fitted to ETF data, generating synthetic time series with predefined regime changes.
result Generated datasets reveal the complexity of detecting and adapting to market regime changes.
We present a formal proof in Lean of probably approximately correct (PAC) learnability of the concept class of decision stumps. This classic result in machine learning theory derives a bound on error probabilities for a simple type of classifier. Though such a proof appears simple on paper, analytic and measure-theoret…
Unified approach to learn interpretable concepts from data.
problem Building interpretable machine learning models and highly-performing foundation models.
method Relating causal representation learning and foundation models, defining concepts and proving their recoverability.
result Provable recovery of human-interpretable concepts from diverse data.
The study identifies latent concepts from diverse observations without assuming specific models.
problem Lack of general theoretical support for concept learning.
method Develops a nonparametric framework for identifying latent concepts from multiple classes of observations.
result Correctness guarantees for concept identification without parametric assumptions.
In many real-world applications, data are often collected in the form of stream, and thus the distribution usually changes in nature, which is referred as concept drift in literature. We propose a novel and effective approach to handle concept drift via model reuse, leveraging previous knowledge by reusing models. Each…
Extends neural network training framework to handle noise and uncertainty.
problem Handling noise and uncertainty in neural network training.
method Integrates non-zero aleatoric noise and derives posterior covariance for epistemic uncertainty.
result Derives an estimator for posterior covariance, providing a handle on epistemic uncertainty.
Humans reason with concepts and metaconcepts: we recognize red and green from visual input; we also understand that they describe the same property of objects (i.e., the color). In this paper, we propose the visual concept-metaconcept learner (VCML) for joint learning of concepts and metaconcepts from images and associ…
In data stream mining, predictive models typically suffer drops in predictive performance due to concept drift. As enough data representing the new concept must be collected for the new concept to be well learnt, the predictive performance of existing models usually takes some time to recover from concept drift. To spe…
A new framework detects concept drift in streaming data.
problem Detecting distributional changes in non-stationary data streams.
method Treating model parameters as random variables, ERICS uses information theory measures to identify concept drift.
result ERICS effectively detects concept drift compared to existing methods.
This research adapts scoring rules for training survival models, improving predictive performance.
problem Training survival models with traditional methods struggles with censoring.
method Adapting scoring rules for survival analysis, creating a flexible framework for model training.
result Scoring rules can be successfully incorporated into model training, yielding competitive performance.
LLMs help automate extraction of actuarial variables from unstructured claims data.
problem Manual processing of unstructured claims data is time-consuming and inconsistent.
method Two-stage processing architecture using LLMs, modular Python pipeline.
result LLM-based extraction achieved high accuracy and practical actuarial value.
The famous Švarc-Milnor Lemma says that a group G acting properly and cocompactly via isometries on a length space X is finitely generated and induces a quasi-isometry equivalence g→g⋅x0 for any x0∈X. We redefine the concept of coarseness so that the proof of the Lemma is automatic.
Geometric framework detects concept frustration between human concepts and machine representations.
problem Aligning human concepts with machine learning representations.
method Geometric framework and similarity measures for detecting concept frustration.
result Concept frustration affects machine learning model performance and reorganizes learned concept representations.
Synthetic tabular data improves privacy while maintaining model performance.
problem Protecting privacy in synthetic data generation for machine learning.
method Deep generative models for tabular data, emphasizing privacy and model performance.
result Deep generative models enhance synthetic data generation for tabular datasets.
Energy-based models can generate complex images by combining simpler concepts.
problem Generating natural images that satisfy complex logical combinations of concepts.
method Energy-based models combine probability distributions of simpler concepts to generate compositions.
result Energy-based models can generate images that satisfy conjunctions, disjunctions, and negations of concepts.
This work proves DP learnability implies online learnability for general classification tasks.
problem Link between differential privacy and online learning for general classification tasks.
method Establishes Ramsey-type theorems for trees to prove DP learnability implies online learnability.
result DP learnability implies online learnability for general classification tasks.
Synthetic proof of Gannon-Lee theorem for spacetimes.
problem Proving incompleteness in globally hyperbolic spacetimes.
method Synthetic null energy condition and synthetically asymptotically regular trappedness condition.
result Generalized classical incompleteness theorem to weighted spacetimes.
Bayesian non-parametric model adapts to concept drifts in streaming data.
problem Inference under concept drift phenomenon for non-stationary data streams.
method Variational inference algorithm for Dirichlet process mixture models with exponential forgetting.
result The proposed model outperforms state-of-the-art algorithms in clustering problems.
Framework learns interpretable concepts from data without interventions.
problem Learning spurious correlations between concepts in CBMs.
method Causal representation learning (CRL) to align latent variables with interpretable concepts using few labels.
result Framework provides theoretical guarantees on correctness and number of required labels without interventions.
Develops new synthetic Ricci flow concepts for metric measure spaces.
problem No specific problem stated; focuses on new mathematical concepts.
method Formulated in terms of dynamic convexity and local concavity of entropy, and global/short-time asymptotic transport cost estimates.
result Shows these properties characterise smooth (weighted) Ricci flows.
A new method for conditional sampling using paired Wasserstein Autoencoders.
problem Conditional sampling from complex data distributions.
method Derive a novel loss function for Wasserstein Autoencoders to enable sampling from OT-type couplings.
result Learned cost-optimal transport maps and conditional sampling from an OT-type coupling.
The paper proves stability of the positive mass theorem using intrinsic flat convergence.
problem Stability of the positive mass theorem in mathematical relativity.
method Intrinsic flat convergence of points and applications to stability.
result Revisits and strengthens the stability results for graphical hypersurfaces of Euclidean space.
AI generates theorems and proofs for training theorem provers.
problem Limited human-written theorems and proofs for supervised learning.
method Proposes a neural generator to automatically synthesize theorems and proofs.
result Synthetic data improves automated theorem proving in Metamath.
Conditional Generative Models are now acknowledged an essential tool in Machine Learning. This paper focuses on their control. While many approaches aim at disentangling the data through the coordinate-wise control of their latent representations, another direction is explored in this paper. The proposed CompVAE handle…
AI techniques explain synthetic tabular data weaknesses.
problem Challenges in evaluating synthetic tabular data quality.
method Apply explainable AI to a binary detection classifier.
result Reveals inconsistencies, unrealistic dependencies, or missing patterns in synthetic data.