Proposes DGBR for stable predictions across unknown environments.
problem Stable predictions across unknown environments in machine learning.
method Joint optimization of a deep auto-encoder for feature selection and a global balancing model for stable prediction.
result Demonstrates stable predictions across unknown environments with empirical experiments.
Gradient descent balances layer magnitudes in deep neural networks without explicit regularization.
problem Balancing magnitudes across layers in deep neural networks.
method Gradient descent with infinitesimal step size enforces layer magnitude balance.
result Gradient descent automatically balances layer magnitudes without explicit regularization.
Regularization leads to balancedness in deep linear networks.
problem Balancedness in deep linear networks.
method Geometric invariant theory and Riemannian geometry of fibers.
result Balancing flows converge to the balanced manifold at a uniform exponential rate.
Gradient descent converges linearly for deep linear networks under specific conditions.
problem Speed of convergence in gradient descent for deep linear neural networks.
method Analysis of gradient descent training for deep linear neural networks minimizing ℓ2 loss. result Gradient descent converges linearly under specific conditions on layer dimensions, initialization, and initial loss.
Proposes a tool to contrast global vs personalized models in clinical prediction.
problem Balancing global vs personalized models in clinical prediction.
method Localized regression approach using autoencoder for dimension reduction.
result Identification of patient subgroups where global models fall short.
The study identifies assets with local balance deviating from global balance to mitigate financial risk.
problem Selecting outperforming assets during financial crises.
method Investigates deviations of local balance from global balance as a criterion for asset selection.
result Assets with local balance deviating from global balance can mitigate financial risk.
Gradient descent proves global convergence for 4-layer matrix factorization.
problem Global convergence of gradient descent on four-layer matrix factorization under random initialization.
method New techniques to show saddle-avoidance properties and extend eigenvalue theories.
result Polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization.
Global balance index measures systemic risk in financial networks.
problem Measuring systemic risk in financial networks.
method Defined global balance index based on a diffusive process and linear system.
result Global balance index correlates with systemic risk measures.
Deeper models have a more favorable optimization landscape, making them more robust to noise.
problem Characterizing the effect of depth on the optimization landscape of linear regression models.
method Robust and over-parameterized setting, simple sub-gradient method.
result A simple sub-gradient method converges to a balanced solution that is close to the ground truth and enjoys a flat local landscape.
Deep learning models forecast multiple yield curves with improved accuracy.
problem Globalization of financial markets affects yield curves.
method Combines self-attention mechanism and nonparametric quantile regression.
result Effective point and interval forecasts of future yields.
Proposes CRLR to handle agnostic selection bias in machine learning.
problem Agnostic selection bias in real-world applications.
method CRLR algorithm that optimizes global confounder balancing and weighted logistic regression.
result CRLR outperforms state-of-the-art methods in robust predictive modeling.
Balanced Activation improves object detection performance on long-tailed datasets.
problem Mismatch between training and testing label distributions in object detection.
method Introduces Balanced Activation (Balanced Softmax and Balanced Sigmoid) to address label distribution shift.
result Balanced Activation provides ~3% gain in mAP on LVIS-1.0 compared to state-of-the-art methods.
Unified theory for causal inference using various methods.
problem Estimating causal effects in ATE estimation.
method Riesz regression, covariate balancing, DRE, TMLE, matching estimator.
result Unified theory integrating multiple methods for ATE estimation.
The paper argues for using Neyman orthogonal score for balancing in debiased machine learning.
problem Debiased machine learning requires a proper approach to balance covariates.
method The paper advocates for using Riesz regression with basis functions of X for balancing.
result Covariate balancing is only valid when the score-relevant regression error is a function of covariates alone.
Proposes a DRL-based MLB for UDNs to balance large-scale traffic.
problem Large-scale load balancing in ultra-dense networks (UDNs).
method Two-layer architecture with DRL for intra-cluster load balancing.
result Empirical results show superior load balancing performance.
Tensor regression networks improve neural network compression and regularization.
problem Improving neural network compression and regularization with low-rank tensor approximations.
method Investigating various low-rank tensor approximations in tensor regression networks.
result Tensor regression networks with Global Average Pooling layer outperformed in deep CNNs, while shallow CNNs with tensor regression and dropout achieved lower test error.
LDAO addresses imbalanced regression by learning local distribution structures.
problem Imbalanced regression with sparse target regions difficult for models.
method LDAO learns local distribution structures, models and samples from each, then merges.
result LDAO outperforms state-of-the-art methods on 45 imbalanced datasets.
Deep learning model detects landmarks in X-ray cephalograms.
problem Automatically detecting landmarks in cephalometric X-ray images.
method 2-stage U-Net with attention mechanism and Expansive Exploration strategy.
result State-of-the-art results in landmark detection.
Deep learning detects cloud changes due to human aerosols.
problem Uncertainty in the effect of anthropogenic aerosols on cloud properties and Earth's energy balance.
method Deep convolutional neural networks to analyze cloud images.
result Identified and characterized specific cloud perturbations due to human aerosols.
Paper introduces Balanced Meta-Softmax for better long-tailed visual recognition.
problem Long-tailed distribution mismatch between training and testing data.
method Balanced Meta-Softmax, an unbiased extension of Softmax, using a Meta Sampler.
result Balanced Meta-Softmax outperforms state-of-the-art solutions on visual recognition and instance segmentation.
Deep linear networks exhibit collapsing features and classifiers across datasets.
problem Understanding the collapse of features and classifiers in deep linear networks.
method Theoretical and empirical analysis of deep linear networks with MSE and CE losses.
result Deep linear networks exhibit NC properties, collapsing features and classifiers to orthogonal vectors.
Less frequent retraining improves forecast accuracy in retail demand forecasting.
problem Balancing forecast accuracy and computational efficiency in global models.
method Analysis of ten machine learning and deep learning models across two large retail datasets with various retraining scenarios.
result Less frequent retraining strategies maintain forecast accuracy while reducing computational costs.
A new deep learning model improves robustness and efficiency in predicting continuous variables.
problem Limited applicability of deep learning in domains with small sample sizes.
method Autoencoder-based residual deep network with shortcut connections.
result Achieves cutting-edge accuracy and efficiency in multiple datasets.
A new XVA strategy rooted in balance sheet perspective improves equity process for bank shareholders.
problem Counterparty risk valuation adjustments (XVAs) in financial derivatives.
method Develops a cost-of-capital XVA strategy in a balance sheet perspective, solving explicitly in static setup and dynamically in trade context.
result Ensures a submartingale equity process corresponding to a target hurdle rate on capital at risk.
This study introduces balanced DRPS and OrderedLogitNN for better QDE of discrete-level questions.
problem Lack of ordinal regression methods and fair evaluation metrics for discrete-level QDE.
method Introduces balanced DRPS and OrderedLogitNN, fine-tunes BERT on RACE++ and ARC datasets.
result OrderedLogitNN outperforms other models on complex QDE tasks.
Novel characterization of augmented balancing weights combining outcome and weighting models.
problem Improving estimation accuracy in machine learning models with balancing weights.
method Characterization of augmented balancing weights as linear models, extending to ridge and lasso regression.
result Equivalence and closed-form expressions for specific model choices, providing insights into performance.
CFR-Pro enhances treatment effect estimation by incorporating local proximity.
problem Treatment selection bias in HTE estimation from observational data.
method Proximity-enhanced CounterFactual Regression (CFR-Pro) with pair-wise proximity regularizer and subspace projector.
result Significantly outperforms competitors in HTE estimation accuracy.
Kernel balancing weights are generalized as KRRR, providing better confidence intervals for treatment effects.
problem Lack of generalization error, correct feature specification, and limited to average effects.
method Interpreting kernel balancing weights as KRRR, relaxing feature specification, and extending Gaussian approximation.
result KRRR provides strong generalization properties and justifies confidence sets for causal functions.
Regression Trees analyze stock returns, revealing market excess return as the most informative factor.
problem Understanding informational content of three factors in stock returns.
method Joint regression tree analysis of daily stock return data for 5 major US corporations.
result The market excess return factor is always the most informative in all cases (solo and joint).
A new beta-VAE based regression model accelerates oilfield optimization studies.
problem Computational expense of full-physics reservoir simulations.
method beta-VAE for interpretable latent space representation, probabilistic dense layers for uncertainty quantification.
result Interpretable latent representation and quantified uncertainty for optimization decisions.
MapLUR uses deep learning on map images to estimate NO2 pollution, outperforming traditional methods.
problem Limited availability of data for traditional LUR models makes them hard to adapt to new areas.
method Data-driven, open-source approach using convolutional neural networks trained on map data.
result MapLUR significantly outperforms traditional LUR models, including those with manually engineered features.
Develops a direct debiased machine learning framework using Bregman divergence.
problem Reduces bias in machine learning estimates of causal effects or structural models.
method Neyman targeted estimation and generalized Riesz regression using Bregman divergence.
result Improves estimation of parameters of interest in causal models.
Proposes a new method to initialize neural networks by estimating global curvature of weights.
problem Improving the initialization of neural networks for better training and convergence.
method Estimates the global curvature of weights across layers using the Hessian matrix norm.
result The proposed method helps in more rigorously initializing weights, leading to better performance.
This study proposes a deep learning framework using ResNeXt for efficient financial data mining.
problem Complex financial data with high dimensionality, nonlinearity, and task correlations.
method Introduces ResNeXt into multi-task learning framework for efficient feature extraction and task collaboration.
result Significantly improved performance in classification and regression tasks on S&P 500 data.
Deep equilibrium models converge globally without explicit computation.
problem Global convergence of deep learning models with implicit layers.
method Analysis of gradient dynamics and proof of convergence rate.
result Deep equilibrium models converge to global optimum at a linear rate.
Global inducing points improve Bayesian neural network performance.
problem Improving Bayesian neural network performance.
method Adapting correlated approximate posterior to all layers in a Bayesian neural network and deep Gaussian processes using learned global inducing points.
result State-of-the-art performance on CIFAR-10 (86.7%) without data augmentation or tempering.
A new method improves Bayesian deep learning by balancing scalability and accuracy.
problem Scalability issues in Bayesian neural networks.
method Collapsed inference scheme that performs Bayesian model averaging using collapsed samples.
result Significant improvements over existing methods in predictive performance and uncertainty estimation.
Astraea improves federated learning accuracy on imbalanced data.
problem Accuracy degradation in federated learning due to imbalanced data distribution.
method Self-balancing federated learning framework with data augmentation and client rescheduling.
result Astraea shows +5.59% and +5.89% improvement in top-1 accuracy on imbalanced datasets.
Hybrid machine learning improves gallstone risk prediction.
problem Complex gallstone disease risk factors and interactions.
method Adaptive LASSO for variable selection, BART for interactions, differential equations for interpretation.
result Enhanced prediction accuracy and actionable insights.
Deep state space model generates text without autoregressive feedback, avoiding biases.
problem Text generation biases and forgetting local nuances.
method Non-autoregressive deep state space model with independent noise and deterministic transition.
result Generative model on par with auto-regressive models, interpretable and without biases.
Inter-domain Deep Gaussian Processes improve inference for non-stationary data.
problem Inference limitations in Gaussian processes for non-stationary data.
method Combines inter-domain and deep Gaussian processes for scalable approximate inference.
result Outperforms inter-domain shallow GPs and conventional DGPs on non-stationary data.
The paper uses deep neural networks to estimate economic models without separability restrictions.
problem Estimating economic models with complex interaction effects and non-separable restrictions.
method Uses deep neural networks as a nonparametric sieve to approximate regression functions from nonlinear latent variable models.
result Economic shape, sparsity, or separability restrictions are imposed more straightforwardly when a flexible latent variable model is used.
Model shows worldwide trade crises can be localized or global, depending on trade balance.
problem Understanding and predicting worldwide trade crises.
method Modeling worldwide trade network using Google matrix analysis and bankruptcy threshold.
result Crisis contagion is localized for high trade balance, global for low trade balance.
Covariance-Driven Regression Trees reduce overfitting in CART.
problem Overfitting in CART decision trees, especially with small sample sizes.
method Covariance-driven splitting criterion for regression trees (CovRT).
result CovRT achieves superior prediction accuracy compared to CART in simulations and real-world tasks.
Novel framework predicts brain biomarker trajectories with superior performance.
problem Challenges in estimating longitudinal brain biomarker trajectories due to variability, inconsistencies, and irregular measurements.
method Personalized deep kernel regression with Adaptive Shrinkage Estimation.
result Superior predictive performance compared to state-of-the-art models.
A new Federated Learning approach balances personalization and global training.
problem Breaking the curse of data heterogeneity in Federated Learning.
method Splitting variables into global and local parameters, using a simple algorithm.
result The approach allows each client to fit their data perfectly, breaking the curse of data heterogeneity.
Deep normative modeling of clinical neuroimaging data improves diagnostic performance.
problem Modeling variation of neuroimaging measures across individuals for psychiatric disorders.
method Proposes a deep normative modeling framework based on neural processes (NPs) for spatially structured mixed-effect modeling of neuroimaging data.
result Substantial improvements in novelty detection performance for certain diagnostic problems.
Meta-GLAR combines global deep representations with local adaptation for improved forecasting accuracy.
problem Joint learning from related time series boosts accuracy but fails for out-of-sample forecasting.
method Meta-GLAR uses a meta-learning approach to adapt RNN representations for each time series.
result Meta-GLAR outperforms state-of-the-art methods in out-of-sample forecasting accuracy.