A framework schedules hyperparameters for model-based reinforcement learning, improving performance.
problem Inadequate scheduling of hyperparameters in model-based reinforcement learning.
method Theoretical analysis and AutoMBPO framework to automatically schedule real data ratio and other hyperparameters.
result Training with hyperparameters scheduled by AutoMBPO significantly improves performance.
This paper uses alternative data to forecast Japanese real estate performance.
problem Accurate rent and price forecasting in Japanese real estate markets.
method Created a comprehensive house price index using over 5 million transactions and economic factors.
result Alternative data variables can forecast real estate performance effectively.
Novel neural likelihood ratio estimation for negative data in particle physics.
problem Estimating likelihood ratios with negative probability densities and weights.
method Introducing a novel loss function and a new model architecture based on signed mixture models.
result Demonstrated improved estimation on a real-world example from particle physics.
The paper proposes using density ratio estimation to evaluate synthetic data quality.
problem Improving the quality and utility of synthetic data for analysis.
method Density ratio estimation to measure synthetic data quality.
result Density ratio estimation yields more accurate global utility estimates than existing methods.
Paper detects changes in graph-based data streams using likelihood-ratios.
problem Detecting changes in synchronized graph-based data streams.
method Kernel-based likelihood-ratio estimation over graph nodes.
result Effective detection and localization of change-points.
Study analyzes factors affecting capital adequacy in Bangladesh's banks.
problem Factors influencing capital adequacy in commercial banks in Bangladesh.
method Fixed Effect, Random Effect, and Pooled Ordinary Least Square (POLS) methods.
result Several independent variables significantly affect capital adequacy, with specific relationships between leverage, liquidity risk, and other factors.
New method corrects bias in density ratio estimation for missing data.
problem Missing data bias in density ratio estimation.
method Adapted KLIEP method (M-KLIEP) for MNAR data.
result M-KLIEP restores consistency and minimax optimality.
The paper presents a framework to quantify the trade-off between synthetic and real data.
problem Improving generalization with synthetic data when real data is scarce.
method Learning-theoretic framework leveraging algorithmic stability to derive generalization error bounds.
result Optimal synthetic-to-real data ratio minimizing expected test error as a function of Wasserstein distance.
In real-world classification problems, the class balance in the training dataset does not necessarily reflect that of the test dataset, which can cause significant estimation bias. If the class ratio of the test dataset is known, instance re-weighting or resampling allows systematical bias correction. However, learning…
Logistic regression is a widely used method in several fields. When applying logistic regression to imbalanced data, for which majority classes dominate over minority classes, all class labels are estimated as `majority class.' In this article, we use an F-measure optimization method to improve the performance of logis…
How does missing data affect our ability to learn signal structures? It has been shown that learning signal structure in terms of principal components is dependent on the ratio of sample size and dimensionality and that a critical number of observations is needed before learning starts (Biehl and Mietzner, 1993). Here …
New method estimates hazard ratios without bias in observational studies.
problem Uninterpretable hazard ratios due to unspecified baseline hazard.
method Kernel-based machine learning to model risk set changes.
result Debiased maximum-likelihood estimators identify true hazard ratios.
Paper tackles unbounded density ratio estimation for covariate shift adaptation.
problem Understudied challenge in statistical learning: unbounded density ratios.
method Three-step estimation method: relative density ratio, truncation, and transformation.
result Established rigorous convergence guarantees for density ratio and regression estimators.
Smart Bayes integrates generative and discriminative features for improved classification.
problem Improving classification performance by combining generative and discriminative modeling.
method Integrates generative likelihood-ratio features into a logistic-regression-style classifier.
result Often outperforms logistic regression and Naive Bayes in simulations and real data.
As one of the most popular linear subspace learning methods, the Linear Discriminant Analysis (LDA) method has been widely studied in machine learning community and applied to many scientific applications. Traditional LDA minimizes the ratio of squared L2-norms, which is sensitive to outliers. In recent research, many …
The study prevents model collapse in overparameterized linear regression by mixing real and synthetic labels.
problem Preventing model collapse in overparameterized linear regression.
method Iterative mixing of real and synthetic labels, deriving generalization error formulae.
result Optimal mixing ratio converges to the reciprocal of the golden ratio for isotropic features.
DRCD identifies causal direction between continuous and discrete variables using density ratio monotonicity.
problem Inferring causal direction between continuous and discrete variables from observational data.
method Density Ratio-based Causal Discovery (DRCD) method.
result DRCD identifies causal direction between continuous and discrete variables using density ratio monotonicity.
The objective of change-point detection is to discover abrupt property changes lying behind time-series data. In this paper, we present a novel statistical change-point detection algorithm based on non-parametric divergence estimation between time-series samples from two retrospective segments. Our method uses the rela…
Investments with best performance are not associated with best Sharpe ratios.
problem The relationship between performance and risk-adjusted return (Sharpe ratio) is counterintuitive for heavy-tailed distributions.
method Synthetic and real data analysis of returns distributions.
result The best-performing investments are not the best in terms of Sharpe ratio, and vice versa.
Improved online changepoint detection for autocorrelated data.
problem Changepoint detection in autocorrelated data with false positives or delays.
method Generalized Likelihood Ratio (GLR) statistic for AR(p) processes, online focus algorithm.
result AR(p)-focus algorithm achieves high detection power in correlated data.
Sequential hypothesis testing is a desirable decision making strategy in any time sensitive scenario. Compared with fixed sample-size testing, sequential testing is capable of achieving identical probability of error requirements using less samples in average. For a binary detection problem, it is well known that for k…
MBORE optimizes multi-objective problems using density-ratio estimation.
problem Optimizing complex, multi-objective functions with expensive evaluations.
method Extends BORE to multi-objective Bayesian optimisation, using density-ratio estimation.
result MBORE outperforms BO on high-dimensional and real-world problems.
SPRT-TANDEM improves sequential classification accuracy with fewer samples.
problem Efficiently classifying sequential data with high accuracy and low sampling cost.
method Deep neural network-based SPRT algorithm that estimates log-likelihood ratio of two hypotheses.
result SPRT-TANDEM achieves statistically significantly better classification accuracy than other classifiers with fewer samples.
A method detects changes in heterogeneous data streams over graph nodes.
problem Detecting changes in data streams from nodes of a graph.
method Online non-parametric method using likelihood-ratio estimation.
result The method accurately identifies change-points in real-world applications.
Let S be a closed orientable surface of genus at least 2 and let G be a semisimple real algebraic group of non-compact type. We consider a class of representations from the fundamental group of S to G called positively ratioed representations. These are Anosov representations with the additional condition that certain …
We use cross ratios to describe second real continuous bounded cohomology for locally compact topological groups. We also derive a rigidity result for cocycles with values in the isometry group of a proper hyperbolic geodesic metric space.
While sample sizes in randomized clinical trials are large enough to estimate the average treatment effect well, they are often insufficient for estimation of treatment-covariate interactions critical to studying data-driven precision medicine. Observational data from real world practice may play an important role in a…
Study on projective structures on a hyperbolic 3-orbifold using tetrahedra.
problem Analyzing projective structures on hyperbolic 3-orbifolds.
method Parameterization using classical invariants, traces, and geometric cross ratios.
result Computed and analyzed the moduli space of projective structures.
Many manifold embedding algorithms fail apparently when the data manifold has a large aspect ratio (such as a long, thin strip). Here, we formulate success and failure in terms of finding a smooth embedding, showing also that the problem is pervasive and more complex than previously recognized. Mathematically, success …
Proposes a deep neural network for multi-dimensional functional data classification.
problem Classifying multi-dimensional functional data with non-Gaussian distributions.
method Trains a deep neural network on the principle components of the training data.
result FDNN achieves minimax optimality when log density ratio has a locally connected modular structure.
Alignment of neural network representations is influenced by SNR and sample size.
problem Understanding how neural network representations align across different conditions.
method Controlled training of neural networks on perturbed datasets, analyzing alignment and generalization.
result Alignment varies monotonically with SNR but non-monotonically with sample size, with minimal alignment near the interpolation threshold.
A new metric evaluates generative models by comparing real and generated samples.
problem Evaluating the quality of generative models.
method Relative Density Ratio (RDR) function, optimization on variational form of φ-divergence.
result The RDR function provides a clear, interpretable, and numerically stable evaluation metric.
Direct neural ratio estimator for likelihood-free inference.
problem Efficient likelihood estimation for complex models.
method Amortized likelihood ratio estimation using neural networks.
result DNRE often outperforms previous ratio estimators.
Unified framework for transfer learning regression without extra cost.
problem Transfer learning for regression models without additional implementation cost.
method Introduces a density-ratio reweighting function estimated via Bayesian framework.
result Unified integration of three TL methods with no extra cost.
DACE estimates covariance from compressed data, improving accuracy.
problem Estimating covariance from large, distributed data.
method Data-aware weighted sampling for unbiased estimation.
result DACE provides more accurate covariance estimation under compression.
This is a survey article on two topics. The Energy E of knots can be obtained by generalizing an electrostatic energy of charged knots in order to produce optimal knots. It turns out to be invariant under Moebius transformations. We show that it can be expressed in terms of the infinitesimal cross ratio, which is a con…
This work simplifies SVM parameter selection using S&S ratio.
problem SVM parameter tuning for optimal performance.
method S&S ratio to model SVM performance; automatic RP, kernel, and parameter selection.
result Optimized SVM parameters with reduced computational complexity.
A flexible nonparametric online changepoint detection algorithm for high-frequency data.
problem Detecting changes in real-time in high-frequency data streams with limited computational resources.
method NP-FOCuS, a sequential likelihood ratio test for a change in the empirical cumulative density function, using functional pruning.
result NP-FOCuS outperforms current nonparametric online changepoint techniques in various settings.
The paper proposes an asset allocation strategy using the Sortino ratio for better performance.
problem Traditional asset allocation methods like the Sharpe ratio do not penalize negative returns adequately.
method The Sortino ratio is used to maximize asset allocation, penalizing only negative return variances.
result The Sortino ratio-based strategy outperforms traditional methods like the Kelly criterion.
Spectral algorithms improve under covariate shift with novel weighted techniques.
problem Improving spectral algorithms' performance under covariate shift.
method Analysis of spectral algorithms in non-parametric regression over RKHS, proposing a weighted spectral algorithm with clipped weights.
result Normalized weighted spectral algorithm achieves optimal capacity-independent convergence rates, and clipped weights can approach optimal capacity-dependent rates.
We develop a new loss function for estimating quasiprobabilistic density ratios.
problem Discontinuous or non-surjective relationships between optimal classifiers and target densities.
method Introduce a convex loss function compatible with both probabilistic and quasiprobabilistic densities.
result Achieve state-of-the-art results in estimating di-Higgs production in particle physics.
SVR-Tree improves classification trees for imbalanced and sparse data.
problem Classification difficulties in imbalanced and sparse data.
method Proposes SVR-Tree, penalizing the Surface-to-Volume Ratio of decision sets.
result SVR-Tree improves generalization error compared to other imbalance algorithms.
Featurization improves density ratio estimation for complex data.
problem Difficulty in estimating density ratios for high-dimensional, different distributions.
method Invertible generative model to map distributions into a common feature space.
result Improved accuracy in density ratio estimation through feature space.
We present a discriminative clustering approach in which the feature representation can be learned from data and moreover leverage labeled data. Representation learning can give a similarity-based clustering method the ability to automatically adapt to an underlying, yet hidden, geometric structure of the data. The pro…
The VIX is used to enhance quantitative trading strategies.
problem Improving Sharpe ratio and reducing trading risks in quantitative strategies.
method Postprocessing quantitative strategies with VIX signals.
result Increased Sharpe ratio and reduced trading risks.
Uniform proof reconstructs spaces using cross ratio on boundary.
problem Reconstructing spaces using cross ratio on boundary.
method Using CAT(-1) spaces and cross ratio on visual boundary.
result Spaces can be reconstructed using cross ratio on boundary.
Meta-learning method improves PU classification performance.
problem Improving binary classifiers from PU data in unseen tasks.
method Adapts model to PU data using related tasks and neural networks.
result Proposed method outperforms existing methods on synthetic and real-world datasets.
Optimal data split ratio is sqrt(p):1 for linear regression.
problem Lack of clear guidance on optimal training/testing data split ratio.
method Showed that optimal ratio is sqrt(p):1 for linear regression.
result Optimal ratio for training/testing split is sqrt(p):1.