Paper uses Super-App data to improve income estimation models.
problem Improving accuracy of income estimation models.
method TreeSHAP method for Stochastic Gradient Boosting Interpretation.
result Alternative data from Super-Apps capture more information than traditional financial data.
Unified approach to estimating class-prior and classifier from PU data.
problem Learning binary classifier from positive and unlabeled data with class-prior estimation.
method Alternately estimating class-prior and training classifier.
result Unified approach improves classifier performance by accounting for class-prior estimation error.
We show that the cusp volume of a hyperbolic alternating knot can be bounded above and below in terms of the twist number of an alternating diagram of the knot. This leads to diagrammatic estimates on lengths of slopes, and has some applications to Dehn surgery. Another consequence is that there is a universal lower bo…
Automatic computation speeds up crosscap number calculation for alternating knots.
problem Computing crosscap numbers for alternating knots efficiently.
method Introduced an automatic computation with complexity O(E3). result Crosscap numbers of alternating knots can be computed in O(E3) time. We consider the problem of sparse coding, where each sample consists of a sparse linear combination of a set of dictionary atoms, and the task is to learn both the dictionary elements and the mixing coefficients. Alternating minimization is a popular heuristic for sparse coding, where the dictionary and the coefficient…
We give sharp two-sided linear bounds of the crosscap number (non-orientable genus) of alternating links in terms of their Jones polynomial. Our estimates are often exact and we use them to calculate the crosscap numbers for several infinite families of alternating links and for several alternating knots with up to twe…
New theory shows how learning algorithms can create a bias towards negative outcomes.
problem Negativity bias in adaptive learning algorithms.
method Generalization of the Hot Stove Effect to settings with negative estimates leading to smaller sample sizes.
result Negativity bias persists even when negative estimates do not lead to avoidance.
We consider learning high-dimensional multi-response linear models with structured parameters. By exploiting the noise correlations among responses, we propose an alternating estimation (AltEst) procedure to estimate the model parameters based on the generalized Dantzig selector. Under suitable sample size and resampli…
Study bounds on cusp volumes of alternating knots on surfaces.
problem Bounding cusp volumes of knots on surfaces.
method Analyzing hyperbolic knots with alternating projections on embedded surfaces.
result Two-sided bounds on cusp area in terms of twist number and surface genus.
Proposes PG-DA for Bayesian MMNL estimation to handle non-conjugacy.
problem Non-conjugacy in the Bayesian estimation of MMNL models.
method Pólygamma data augmentation technique applied to MMNL estimation.
result Similar posterior estimates for binary choice scenarios, but empirical identification issues for J≥3 alternatives. Alternative approach to rigidity of high-dimensional isometric immersions.
problem Rigidity of high-dimensional isometric immersions between compact manifolds.
method Quantitative rigidity estimates, reducing to Euclidean setting and applying Friesecke-James-Müller rigidity estimate.
result Quantitative results showing close proximity to isometric immersions for small stretching and bending energy.
A new covariance estimator reduces dimensionality and improves portfolio forecasting.
problem Estimating high-dimensional covariance matrices with weak factors.
method Sparse Approximate Factor (SAF) model with l1-regularization. result SAF estimator outperforms other methods in portfolio forecasting.
Improved spatial prediction for massive datasets using SME model.
problem Efficiently estimating parameters in massive spatial datasets.
method Spatial Mixed Effects (SME) model with AECM algorithm for flexibility.
result Improved estimation without sacrificing prediction accuracy.
Paper introduces a new robust method for estimating Pareto tail index from grouped data.
problem Limited robust methods for estimating Pareto tail index from grouped data.
method Method of Truncated Moments (MTuM)
result Inferential justification and validation of MTuM through simulation study.
Proposes unbiased estimators for training mixture of experts models.
problem Efficiently training large-scale mixture of experts models on modern hardware.
method Two unbiased estimators based on principled stochastic assignment procedures.
result Both estimators are more effective and robust than biased alternatives.
Simplified proof and new C0 estimate for Kähler-Einstein metrics.
problem Existence of Kähler-Einstein metrics on Calabi-Yau manifolds.
method Alternative C0 a priori estimate for the Monge-Ampère equation. result Established a new uniform bound for the solution of the Monge-Ampère equation.
New method uses zeroth-order queries to approximate proximal sampling efficiently.
problem Approximating proximal sampling with zeroth-order information.
method Direct simulation of heat flow dynamics, treating intermediate distribution as Gaussian mixture.
result Inherits exponential convergence under isoperimetric conditions, avoids rejection sampling.
New estimators reduce variance in training variational autoencoders with discrete latent variables.
problem Training variational autoencoders with discrete latent variables requires efficient gradient estimation.
method Introduce ReinMax-Rao and ReinMax-CV estimators using Rao-Blackwellisation and control variates.
result Demonstrate superior performance on training variational autoencoders with discrete latent spaces.
The telegraph process models a random motion with finite velocity and it is usually proposed as an alternative to diffusion models. The process describes the position of a particle moving on the real line, alternatively with constant velocity +v or −v. The changes of direction are governed by an homogeneous Poisso…
A new R package for high-dimensional regression and precision matrix estimation.
problem High-dimensional linear regression and precision matrix estimation challenges.
method flare package implements various regression methods and extensions for sparse precision matrix estimation.
result The flare package is efficient and scalable for large problems.
The estimation of probabilities of default (PDs) for low default portfolios by means of upper confidence bounds is a well established procedure in many financial institutions. However, there are often discussions within the institutions or between institutions and supervisors about which confidence level to use for the…
ES optimization improved by structured control variates.
problem Improving accuracy of Evolution Strategies in RL.
method RL-specific variance reduction through structured control variates.
result Structured control variates outperform general variance reduction methods.
PCMC-Net uses neural networks to estimate transition rates in choice models, improving accuracy over traditional methods.
problem Inference limitations of traditional PCMC models when examples are scarce or new alternatives are observed.
method Amortized inference approach embedding PCMC definition into a neural network.
result Neural network outperforms feature engineered and machine learning models in airline booking prediction.
Novel unsupervised random forests improve density estimation and data synthesis.
problem Density estimation and data synthesis for complex tabular data.
method Recursive unsupervised random forests with alternating generation and discrimination rounds.
result Provable consistency and smooth densities with fast execution.
New method stabilizes machine learning predictions across random seeds.
problem Machine learning predictions vary across random seeds, causing instability.
method Introduces adaptive cross-bagging to eliminate seed dependence.
result Adaptive cross-bagging achieves targeted stability in debiased machine learning.
Estimates heterogeneous treatment effects in panel data with a new method.
problem Estimating heterogeneous treatment effects in panel data with general treatment patterns.
method Partition observations into clusters with similar treatment effects using a regression tree, then estimate average treatment effects for each cluster.
result Our method achieves superior accuracy compared to alternative approaches.
Paper proposes an alternative to anchor points for learning with noisy labels.
problem Learning with noisy labels is challenging due to inaccurate labels.
method Estimates transition matrix using clusterability condition and noisy labels.
result Estimation of transition matrix is more accurate and efficient than anchor points.
Model analyzes cooccurrence data for recommender systems and item relevance.
problem High-dimensional cooccurrence data from online platforms.
method Shared parameter Alternating Tweedie (SA-Tweedie) model with Fisher scoring and learning rate adjustment.
result SA-Tweedie model outperforms other methods in optimizing parameters.
The paper develops methods for conditional inference on the asset with the highest Sharpe ratio.
problem Performing inference on the asset with the highest Sharpe ratio among correlated assets.
method Conditional inference procedure using multivariate Sharpe ratio standard error, alternative tests, and asymptotic adjustments.
result The conditional inference procedure achieves nominal type I rate and maintains near-nominal rejection rates under the conditional null.
A faster method for estimating effects in large data using fixed-point trees.
problem Estimating heterogeneous effects in large dimensions with computational efficiency.
method Fixed-point approximation to eliminate Jacobian estimation and speed up GRFs.
result Significant computational efficiency improvement without sacrificing statistical accuracy.
VB methods improve MMNL estimation speed and accuracy.
problem Scalable Bayesian estimation of MMNL models.
method Extending VB methods to include both fixed and random utility parameters, and conducting extensive simulations.
result VB methods, especially VB-NCVMP-Delta, are up to 16 times faster than MCMC and MSLE while maintaining similar accuracy.
If a hyperbolic link has a prime alternating diagram D, then we show that the link complement's volume can be estimated directly from D. We define a very elementary invariant of the diagram D, its twist number t(D), and show that the volume lies between v_3(t(D) - 2)/2 and v_3(16t(D) - 16), where v_3 is the volume of a…
In the field of structural reliability, the Monte-Carlo estimator is considered as the reference probability estimator. However, it is still untractable for real engineering cases since it requires a high number of runs of the model. In order to reduce the number of computer experiments, many other approaches known as …
We investigate an existing distributed algorithm for learning sparse signals or data over networks. The algorithm is iterative and exchanges intermediate estimates of a sparse signal over a network. This learning strategy using exchange of intermediate estimates over the network requires a limited communication overhea…
Estimates network structure and interaction rules from multiple agent trajectories.
problem Modeling multi-agent systems on networks from data.
method Jointly infers network topology and interaction kernels using non-convex optimization.
result ORALS estimator is consistent and asymptotically normal under coercivity conditions.
A new method reduces sample complexity for meta-learning.
problem Efficiently learn new tasks with minimal data.
method Alternating minimization method (MLLAM) for linear regression tasks.
result MLLAM achieves nearly-optimal estimation error with Ω(logd) samples per task. Bayesian neural network approach improves estimation of complex economic models.
problem Difficulty in estimating complex economic simulation models due to intractable likelihood functions.
method Bayesian estimation using deep neural networks to approximate likelihood functions.
result Proposed methodology yields more accurate estimates across various economic models.
We discuss the coherence properties of Expected Shortfall (ES) as a financial risk measure. This statistic arises in a natural way from the estimation of the "average of the 100p % worst losses" in a sample of returns to a portfolio. Here p is some fixed confidence level. We also compare several alternative representat…
Study tests uniformity of categorical data against missing-ball alternatives, finding chi-squared test outperforms.
problem Testing uniformity of categorical data against missing-ball alternatives.
method Characterizes minimax risk, uses collisions and chi-squared test, reduces to structured subset of alternatives.
result Minimax test outperforms chi-squared test under least favorable alternative.
Develops ADMM for estimating precision matrices from noisy, missing data.
problem Estimating precision matrices from noisy and missing data.
method Alternating Direction Method of Multipliers (ADMM) for non-positive semidefinite inputs.
result Empirically compares ADMM with existing methods and characterizes tradeoffs.
We consider sequential decision making problems for binary classification scenario in which the learner takes an active role in repeatedly selecting samples from the action pool and receives the binary label of the selected alternatives. Our problem is motivated by applications where observations are time consuming and…
Unified method for MMD variance estimation improves accuracy and computational efficiency.
problem Variance estimation for MMD in nonparametric testing.
method Unified finite-sample characterization of MMD variance through U-statistic and Hoeffding decomposition; exact acceleration method for univariate case.
result Unified estimators improve accuracy and computational efficiency for MMD variance.
Investigates optimal consumption and investment using alternative data sources.
problem Optimal consumption and investment decisions under hidden economic regimes.
method Develops a novel duality theory for a jump-diffusion process with alternative data.
result Provides conditions for using control approach based on dynamic programming.
Neural networks approximate CDFs for efficient likelihood estimation.
problem Efficiently estimating likelihoods for complex distributions.
method Parameterizing conditional CDFs with neural networks and using automatic differentiation.
result A range of neural network architectures for CDF estimation, from simple to flexible.
For a given knot, we study the minimal number of positive eigenvalues of the double branched cover over spanning surfaces for the knot. The value gives a lower bound for various genera, the dealternating number and the alternation number of knots, and we prove that Batson's bound for the non-orientable 4-genus gives an…
New algorithm improves treatment effect estimation from observational data.
problem Estimating the benefits and harms of interventions from observational data.
method Develops a deep kernel regression algorithm and posterior regularization framework.
result Substantially outperforms state-of-the-art on various benchmarks datasets.
New estimators improve sparse semiparametric additive modeling.
problem Sparse semiparametric additive modeling with structured sparsity.
method Combines group subset selection with shrinkage for nonconvex optimization.
result New estimators outperform alternatives in synthetic and real-world data.
We use simple properties of the Rasmussen invariant of knots to study its asymptotic behaviour on the orbits of a smooth volume preserving vector field on a compact domain in the 3-space. A comparison with the asymptotic signature allows us to prove that asymptotic knots are non-alternating, in general. Further we show…