Two privacy-preserving rating collection methods for recommender systems.
problem Collecting user ratings while maintaining privacy.
method Modified Laplace mechanism and randomized response.
result Both mechanisms are differentially private and preserve data utility.
Active data collection improves convergence rates in operator learning.
problem Improving convergence rates in operator learning with linear target and stochastic input.
method Active data collection strategies with mean-zero stochastic process and continuous covariance kernels.
result Achieves arbitrarily fast error convergence rates with eigenvalue decay of covariance kernels.
New guarantees for ERM with adaptively collected data.
problem Failure of ERM guarantees with adaptively collected data.
method Importance sampling weighted ERM algorithm with maximal inequality.
result First generalization guarantees and fast convergence rates for adaptively collected data.
We present a simple model of firm rating evolution. We consider two sources of defaults: individual dynamics of economic development and Potts-like interactions between firms. We show that such a defined model leads to phase transition, which results in collective defaults. The existence of the collective phase depends…
Auto-Ensemble automates deep learning model ensembling with adaptive learning rate scheduling.
problem Difficulty in collecting diverse and accurate deep learning models through single training.
method Auto-Ensemble collects model checkpoints and uses adaptive learning rate scheduling to ensemble them.
result Ensembled models converge to various local optima, improving performance on few-shot learning.
Large speech dataset for commercial use with 9.98% word error rate.
problem Creating a diverse speech recognition dataset for commercial purposes.
method Internet search for licensed audio data with transcriptions, training model on the dataset.
result Model trained on dataset achieves 9.98% word error rate on Librispeech's test-clean test set.
Optimizes binary rating systems for item ranking.
problem Designing efficient feedback systems for item ranking.
method Formalizes performance, provides algorithm, empirically designs and validates.
result Empirically designed and validated approximately optimal rating system.
Master algorithm selects best contextual bandit from a collection.
problem Model selection in stochastic contextual bandit setting.
method Random selection with probability adjustment based on comparison of cumulative rewards.
result Achieves the same regret rate as the best candidate in a collection of black-box algorithms.
Rating prediction is an important application, and a popular research topic in collaborative filtering. However, both the validity of learning algorithms, and the validity of standard testing procedures rest on the assumption that missing ratings are missing at random (MAR). In this paper we present the results of a us…
New CVs preserve transition rates in molecular dynamics.
problem Designing CVs that accurately capture rare events in high-dimensional systems.
method Integrating manifold learning and group-invariant featurization to construct neural network-based CVs that satisfy orthogonality conditions.
result Achieved a CV for butane that reproduces the anti-gauche transition rate with less than ten percent relative error.
Paper finds exact exponent in optimal error rates for crowdsourcing.
problem Determining the optimal number of workers for accurate label aggregation.
method Using the Dawid-Skene model, the paper establishes matching upper and lower bounds with an exact exponent.
result The exact exponent mI(π) is found, allowing precise sample size requirements. This paper modifies the Ait-Sahalia model to better describe interest rate behaviors.
problem Inadequate specifications of the original Ait-Sahalia model to explain various interest rate phenomena.
method Proposes a modified hybrid Poisson-jump Ait-Sahalia model and uses truncated EM techniques for numerical approximation.
result Validates the modified model using Monte Carlo simulations for bond and barrier option payoffs.
CoRAS adapts image acquisition rates for accurate reconstruction.
problem Determining when enough measurements are collected for accurate image reconstruction.
method Adaptive acquisition rate selection based on reconstruction error probability.
result CoRAS achieves target stopping-time coverage with fewer measurements.
This work tackles collective matrix completion with multiple and heterogeneous data sources.
problem Reconstructing data from multiple heterogeneous matrices.
method Estimation based on minimizing goodness-of-fit and nuclear norm penalization of the whole collective matrix.
result Proposed estimators achieve fast rates of convergence under two settings.
Modeling business cycles via collective risk fluctuations in economic agents' risk space.
problem Understanding and predicting business cycles through economic agents' risk dynamics.
method Continuous numerical risk grades for economic agents, modeling collective economic variables and flows as functions of risk coordinates, deriving equations for their evolution.
result Business and credit cycles are explained as fluctuations of collective economic variables and their mean risks in the risk space of economic agents.
Model shows who pays higher taxes affects wealth distribution.
problem Determining who should pay higher taxes to prevent wealth concentration.
method Dynamic agent model with random wealth multiplicative process and linear tax rate.
result Tax rate structure affects long-term wealth distribution.
ISOKANN learns collective variables and effective dynamics for metastable transitions.
problem Understanding metastable transitions in complex molecular systems.
method Integrates Koopman operators with neural networks to extract CVs and effective dynamics.
result Reconstructs coarse-grained kinetics and reproduces transition times across barriers.
Improved disability insurance model with collective health claims.
problem Enhance disability insurance model with collective health claims.
method Expand classic semi-Markov model with collective health claims, solve many-body problem using mean-field approach.
result Mean-field approach simplifies complex model into a transparent pricing method.
Model estimates loan cure rate using Markov chains.
problem Estimating the cure rate of non-performing loans.
method Developed a Markov-chain model for portfolios.
result Efficient and accessible for smaller institutions.
Paper anonymizes user ratings to protect privacy while improving recommendation accuracy.
problem Protecting user privacy while maintaining recommendation accuracy with anonymized ratings.
method Exhaustively lists recommender models using anonymized ratings and presents item-based collaborative filtering algorithms.
result Item-based collaborative filtering based on anonymized ratings outperforms non-anonymized ratings in some settings.
An empirical analysis of interest rates in money and capital markets is performed. We investigate a set of 34 different weekly interest rate time series during a time period of 16 years between 1982 and 1997. Our study is focused on the collective behavior of the stochastic fluctuations of these time-series which is in…
Paper studies federated nonparametric testing with privacy constraints, achieving optimal rates and adaptive testing.
problem Federated nonparametric goodness-of-fit testing under distributed differential privacy constraints.
method Establishes matching lower and upper bounds on minimax separation rate, constructs adaptive testing procedure.
result Achieves optimal rates and demonstrates phase transition phenomena in federated testing.
Optimizes data collection for ranking and selection problems.
problem Identifying the best system from multiple solutions with limited data.
method Sequential sampling algorithm with MPB estimator and kernel ridge regression.
result OSAR achieves optimal sampling ratios almost surely in the limit.
Improves NMT with user feedback from eBay ratings and search tasks.
problem Improving neural machine translation quality with user feedback.
method Offline bandit learning of NMT parameters using real user feedback from eBay.
result Implicit task-based feedback from cross-lingual search tasks improves NMT quality.
Study growth rates of subgroups in groups with a constricting element.
problem Understanding growth rates of subgroups in groups with a constricting element.
method Examining the spectrum of relative and quotient exponential growth rates of quasi-convex subgroups.
result Determine when growth rates of subgroups are strictly smaller or coincide with the group's growth rate.
Unified Bayesian framework for CAT bond pricing.
problem Uncertainty in catastrophe occurrences and interest rates in CAT bond markets.
method Bayesian framework based on uncertainty quantification of catastrophes and interest rates.
result Unified asset pricing approach with informative expected risk premia.
SVGD algorithm converges at rate 1/sqrt(log log n) for sub-Gaussian distributions.
problem Approximating a probability distribution with particles.
method Stein variational gradient descent (SVGD) with finite particles and sub-Gaussian target distribution.
result SVGD achieves a convergence rate of 1/sqrt(log log n) for sub-Gaussian distributions.
Adapting policy learning for data collected from evolving systems.
problem Challenges in learning optimal policies from adaptively collected data.
method Proposes an algorithm based on generalized augmented inverse propensity weighted (AIPW) estimators to control worst-case estimation variance.
result Achieves minimax rate optimal regret guarantees even with diminishing exploration.
The paper argues that fairness in predictions should be evaluated in context and addressed through data collection.
problem Fairness in predictive models in sensitive applications like healthcare and criminal justice.
method Decompose cost-based metrics of discrimination into bias, variance, and noise; propose actions to estimate and reduce each term; perform case-studies.
result Data collection is often a means to reduce discrimination without sacrificing accuracy.
New models extrapolate false alarms in ASV without new data.
problem Reliable extrapolation of false alarm rates in ASV without new speaker data.
method Generative models in ASV score space for arbitrary systems.
result Models accurately extrapolate false alarm rates for large speaker populations.
This research improves debt collection strategies using advanced machine learning.
problem Accurate estimation of propensity to pay and cashflow for optimal debt collection.
method Developed a machine learning framework with pre-processing and model selection.
result The proposed model outperforms current industry strategies.
Bayesian model predicts emotion from fitness tracker heartbeat data.
problem Predicting emotional valence from consumer fitness tracker heartbeat data.
method End-to-end Bayesian deep learning model using PPG data.
result Peak F1 score of 0.7 for emotional valence classification.
Study shows Bitcoin dominates global crypto-market, leading to self-contained trading.
problem Understanding the dominance and influence of cryptocurrencies in global trading.
method Analysis of daily exchange rates and cross-correlations in 100 highest-capitalization cryptocurrencies.
result The dominance of Bitcoin leads to a self-contained cryptocurrency market.
KSDAgg combines multiple KSD tests to improve goodness-of-fit testing without splitting data.
problem Improving goodness-of-fit testing without data splitting.
method KSDAgg aggregates multiple KSD tests with different kernels to maximize power.
result KSDAgg achieves the smallest uniform separation rate of the collection, up to a logarithmic term.
Quantile regression with ReLU networks achieves minimax rates for various function types.
problem Estimating quantiles from covariates with neural networks.
method Quantile regression with rectified linear unit (ReLU) neural networks.
result ReLU networks achieve minimax rates for broad collections of function types.
Develops a method to efficiently use offline data for RL policy optimization.
problem Lack of online data for offline RL in mobile health applications.
method Advantage learning framework using optimal Q-estimators.
result New policy converges faster than existing methods.
This paper reviews methods for constructing confidence intervals for error rates in 1:1 matching tasks.
problem Challenges in assessing uncertainty of error rates in matching algorithms, especially when data are dependent and error rates are low.
method Derives and examines statistical properties of methods for constructing confidence intervals for error rates in 1:1 matching tasks.
result Coverage and interval width vary with sample size, error rates, and data dependence.
Crowdsourcing has become a primary means for label collection in many real-world machine learning applications. A classical method for inferring the true labels from the noisy labels provided by crowdsourcing workers is Dawid-Skene estimator. In this paper, we prove convergence rates of a projected EM algorithm for the…
DL-FUMI learns heartbeat patterns from BCG signals for precise heart rate estimation.
problem Estimating precise heart rates from ballistocardiogram signals with uncertainty.
method Multiple instance dictionary learning to learn heartbeat concepts from BCG signals.
result DL-FUMI's heartbeat concept achieves superior performance over comparison algorithms.
We analyse four consecutive cycles observed in the USA for employment and inflation. They are driven by three oil price shocks and an intended interest rate shock. Non-linear coupling between the rate equations for consumer products as prey and consumers as predators provides the required instability, but its natural d…
The paper learns particle swarming models from data using Gaussian processes.
problem Understanding the link between individual interaction rules and swarming behavior.
method Proposes a learning approach using Gaussian processes to model latent radial interaction functions and scalar parameters in non-collective friction forces.
result Establishes that a coercivity condition is sufficient for recoverability and provides a finite-sample analysis showing optimal convergence rates.
Inspired by the unsupervised learning or self-organization in the machine learning context, here we attempt to draw `learning curve' for the collective behavior of job-seeking `zero-intelligence' labors in successive job-hunting processes. Our labor market is supposed to be opened especially for university graduates in…
Study proposes GRU-D networks for missing value handling in road surface friction prediction.
problem Missing values in road surface friction data affect prediction accuracy.
method Gated Recurrent Unit (GRU) network with decay mechanism.
result GRU-D networks outperform baseline models in road surface friction prediction.
Ensembles of neural networks learn better by sharing information.
problem Improving performance of neural networks through collective learning.
method Modeling neural networks as socially interacting agents aiming to maximize their own performance and functional relations to others.
result Optimal collective performance emerges from local interactions between networks, leading to specialization and higher confidence.
Bayesian PINNs learn elliptic PDEs with near-minimax posterior contraction rate.
problem Learning elliptic PDEs with noisy data and non-homogeneous boundary conditions.
method Bayesian approach with a Hölder space prior on neural network weights.
result Posterior contracts at near-minimax rate without prior knowledge of solution smoothness.
Paper proposes a pre-conditioning technique to speed up gradient-descent convergence in distributed linear least-squares problems.
problem Expediting convergence of gradient-descent method for ill-conditioned distributed linear least-squares problems.
method Iterative pre-conditioning technique to improve convergence rate of gradient-descent method.
result Pre-conditioned gradient-descent achieves superlinear convergence for unique solutions and improved linear convergence otherwise.
Near-optimal rates for multi-task learning with shared representations.
problem Approximation and statistical complexity of learning multiple operators.
method Multiple Neural Operators (MNO) architecture and comparison with DeepONet.
result Near-optimal upper and lower bounds for approximation and generalization.
Given a graph where vertices represent alternatives and arcs represent pairwise comparison data, the statistical ranking problem is to find a potential function, defined on the vertices, such that the gradient of the potential function agrees with the pairwise comparisons. Our goal in this paper is to develop a method …