We investigate Fano schemes of conditionally generic intersections, i.e. of hypersurfaces in projective space chosen generically up to additional conditions. Via a correspondence between generic properties of algebraic varieties and events in probability spaces that occur with probability one, we use the obtained resul…
Paper examines vulnerabilities in data-driven pricing schemes.
problem Vulnerability of clustering-oriented pricing schemes to malicious user behavior.
method Defined a notion of disguising to identify strategic behaviors of malicious users, characterized sensitivity zones to evaluate malicious user percentages, conducted cost benefit analysis.
result Concluded with a vulnerability analysis of data-driven pricing schemes.
New sampling scheme improves ML accuracy in physics simulations.
problem Improving accuracy of ML models in physics simulations.
method Taylor-based data sampling scheme for DNNs.
result Reduces error in DNN solutions of ODE systems.
Automates standards design through machine learning.
problem Manual, time-consuming standards design process.
method Reinforcement learning to optimize proposals.
result Streamlines standards design and innovation.
A new privacy-preserving deep learning scheme for asymmetrically collaborative machine learning.
problem Privacy and efficiency in collaborative machine learning across different data owners.
method Decomposes neural network steps for privacy-preserving training; novel protocol for information leakage.
result Efficient training with stable performance and significant speedup.
Human inertial thinking schemes can be formed through learning, which are then applied to quickly solve similar problems later. However, when problems are significantly different, inertial thinking generally presents the solutions that are definitely imperfect. In such cases, people will apply creative thinking, such a…
New method corrects bias in event study analysis using machine learning.
problem Bias in conventional event study analysis mechanisms.
method Topological machine-learning approach with self-organizing map (SOM).
result Identifies factors of abnormal stock returns and depicts event clusters.
Resource allocation improved using machine learning from terminal positions.
problem Optimizing resource allocation in next-gen wireless systems with fast-changing channel conditions.
method Supervised machine learning using position information of mobile terminals.
result Coordinates-based resource allocation performs similarly to traditional CSI-based methods.
Novel parallelization simplifies machine learning algorithms.
problem Adapting machine learning algorithms to growing data and needs.
method A novel parallelization scheme that applies to broad learning algorithms.
result Reduces runtime to polylogarithmic time on quasi-polynomially many units.
Machine learning aids in discovering and optimizing quantum protocols for long-distance communication.
problem Designing and optimizing quantum communication protocols for long distances.
method Projective simulation, a learning agent combining reinforcement learning and decision making.
result Machine learning identifies and improves quantum protocols like teleportation and entanglement purification.
A new machine learning method solves high-dimensional Kolmogorov PDEs efficiently.
problem Solving high-dimensional Kolmogorov PDEs and SDEs.
method Stochastic weighted minimization and stochastic gradient descent with Malliavin weights.
result Accurate approximation of high-dimensional Kolmogorov PDEs and SDEs without curse of dimensionality.
The paper tackles the trade-off between fairness and accuracy in machine learning models.
problem Ensuring fairness in machine learning often reduces model accuracy.
method The paper introduces formal tools for reconciling the fairness-accuracy tension using Pareto optimality from multi-objective optimization.
result The Chebyshev scalarization scheme is superior for finding Pareto optimal solutions compared to the linear scalarization scheme.
New learning scheme solves high-dimensional semi-linear PDEs using sparse grids and Picard approximations.
problem Solving high-dimensional semi-linear parabolic PDEs.
method Probabilistic learning scheme based on Picard iteration with SGD, employing sparse grid approximation.
result Convergence proof and polynomial complexity in ε−1 for high-dimensional PDEs. New ML-based detection improves PMH signal detection in load-modulated MIMO systems.
problem Detecting PMH signals without prior CSI is challenging and computationally expensive.
method Proposes HEM-ML and HEM-KD schemes using EM and KD-tree for efficient detection.
result Achieves comparable detection results to optimal ML detector with reduced complexity.
Many modern machine learning models are trained to achieve zero or near-zero training error in order to obtain near-optimal (but non-zero) test error. This phenomenon of strong generalization performance for "overfitted" / interpolated classifiers appears to be ubiquitous in high-dimensional data, having been observed …
Unified theory linking atom-centered and message-passing models for molecular properties.
problem Combining atom-centered and message-passing models for accurate molecular property prediction.
method Generalizing ACDC framework to include multi-centered information, providing a complete linear basis for regression.
result Unified understanding of atom-centered and message-passing models, providing a coherent foundation.
New trends explore quantum machine learning to speed up computations and analyze data.
problem Speeding up machine learning computations and analyzing large quantum data.
method Interplay between quantum physics and machine learning, including new algorithms and hardware.
result Breakthroughs in quantum machine learning can provide advantages over classical methods.
Recent advances in cryptography promise to enable secure statistical computation on encrypted data, whereby a limited set of operations can be carried out without the need to first decrypt. We review these homomorphic encryption schemes in a manner accessible to statisticians and machine learners, focusing on pertinent…
Introduces REVE, a regularization scheme that compresses class conditioned entropy.
problem Improving generalization performance of deep learning models.
method Identifies a variable responsible for final prediction, compresses class conditioned entropy, introduces a variational upper bound, and integrates a tractable loss into training.
result Demonstrates the efficiency of REVE on various neural networks and datasets.
New approach connects stochastic gradient descent to ODE splitting schemes.
problem Improving convergence in stochastic optimization.
method Connection between stochastic gradient descent and ODE splitting schemes.
result Derive a new upper bound on global splitting error.
The paper proposes a model reward scheme for collaborative ML based on Shapley value and information gain.
problem Designing fair incentives for collaborative machine learning.
method The paper proposes a reward scheme based on Shapley value and information gain, with properties like fairness and stability.
result The proposed reward scheme satisfies fairness and trade-offs between desirable properties via an adjustable parameter.
PILAE learns DNNs without gradient descent, achieving better performance.
problem Training deep feedforward neural networks efficiently and accurately.
method PILAE uses a pseudoinverse learning algorithm for autoencoder building blocks of MLP DNNs.
result PILAE achieves better performance on tradeoff between training efficiency and accuracy.
We propose a distributed approach to train deep neural networks (DNNs), which has guaranteed convergence theoretically and great scalability empirically: close to 6 times faster on instance of ImageNet data set when run with 6 machines. The proposed scheme is close to optimally scalable in terms of number of machines, …
HAR model outperforms ML in stock forecasting with correct fitting schemes.
problem Realized volatility forecasting using machine learning techniques.
method Investigated the role of fitting schemes in HAR model performance, focusing on training window and re-estimation frequency.
result HAR model consistently outperforms ML models when using a correctly specified fitting approach.
A new method uses nearest neighbors for importance weighting.
problem Data covariate shift problems in machine learning.
method Nearest neighbor classification scheme for determining importance weights.
result Demonstrated effectiveness through comparative experiments on various classification tasks.
New framework ties learning algorithms to data recognition using CMI.
problem Understanding how well machine learning models generalize from training data.
method Information-theoretic framework using Conditional Mutual Information (CMI).
result Bounds on CMI imply various forms of generalization.
New SVM margin bound improves generalization in machine learning.
problem Improving SVM margin bounds for better generalization.
method Stable sample compression schemes to derive new data-dependent generalization bounds.
result Proves a new optimal SVM margin bound with a log factor improvement.
Deep learning schemes solve high-dimensional nonlinear PDEs and variational inequalities.
problem Solving high-dimensional nonlinear PDEs and variational inequalities.
method Machine learning using backward stochastic differential equations and deep neural networks.
result Deep learning schemes converge and give good results up to dimension 50.
Machine learning improves probing of quantum systems via synchronization.
problem Characterizing the dissipation features of quantum systems.
method Machine learning applied to a probing scheme with quantum synchronization.
result Machine learning significantly improves the inference of dissipation features from probe observables.
This paper studies parallelization schemes for stochastic Vector Quantization algorithms in order to obtain time speed-ups using distributed resources. We show that the most intuitive parallelization scheme does not lead to better performances than the sequential algorithm. Another distributed scheme is therefore intro…
It is of fundamental importance to find algorithms obtaining optimal performance for learning of statistical models in distributed and communication limited systems. Aiming at characterizing the optimal strategies, we consider learning of Gaussian Processes (GPs) in distributed systems as a pivotal example. We first ad…
Defense against adversarial examples using k-NN on neural network activations.
problem Adversarial examples that fool machine learning models.
method k-Nearest Neighbor (kNN) on intermediate activations of neural networks.
result Significantly outperforms state-of-the-art defenses on MNIST and CIFAR-10.
Two new coding schemes improve the efficient communication of noisy data.
problem Efficient communication of noisy data in machine learning.
method Ordered Random Coding (ORC) and Hybrid Coding Scheme.
result Improved coding schemes over existing approaches.
Machine learning analysis of galaxy catalogues finds only one class separable.
problem Difficulty in classifying galaxies using visual inspection schemes.
method Generalized Relevance Matrix Learning Vector Quantization and Random Forests.
result Only one class, Little Blue Spheroids, is consistently separable.
Machine learning predicts synchronization transitions in unknown systems.
problem Predicting synchronization transitions in systems with unknown equations.
method Developed a 'parameter-aware' machine learning scheme using reservoir computing or echo state networks.
result Machine learning accurately predicts synchronization transitions, including hysteresis loops.
Deep learning scheme identifies and reconstructs chaotic and stochastic systems from noisy data.
problem Challenging identification of governing equations from noisy and partial observations.
method Jointly learns inference model and governing laws using variational deep learning.
result Framework generalizes state-of-the-art methods and accounts for stochastic variabilities.
Language models improve clinical prediction models using EHR data.
problem Limited patient data for training clinical prediction models.
method Using patient representation schemes from natural language processing.
result 3.5% mean improvement in AUROC on five prediction tasks.
Study reveals patterns in crypto-markets and predicts pump-and-dump schemes.
problem Understanding and predicting pump-and-dump activities in cryptocurrency markets.
method Empirical case study, data analysis, machine learning model building.
result Highly precise and robust machine learning model for predicting pump-and-dump events.
Distributed-OMP recovers sparse vectors with low communication costs.
problem High-dimensional sparse linear regression with limited computation and communication.
method Distributed orthogonal matching pursuit (OMP) scheme.
result Support of the regression vector can be recovered with linear communication per machine and logarithmic in dimension.
A machine learning approach for dynamic stock recommendation outperforms traditional strategies.
problem Lack of time for analysts to check all S&P 500 stocks and the need for a reliable stock selection strategy.
method Selecting representative stock indicators, using five machine learning methods, and choosing the model with the lowest Mean Square Error to rank stocks.
result The proposed scheme outperforms the long-only strategy on the S&P 500 index in terms of Sharpe ratio and cumulative returns.
Gradient codes adapt to varying straggler counts in distributed learning.
problem Mitigating slow worker nodes (stragglers) in distributed machine learning.
method Proposes a flexible gradient coding scheme that concatenates codes for different straggler tolerances, adapting to actual straggler counts.
result Significantly lower latency compared to fixed-tolerance gradient codes.
Scheduling and power allocation improve federated learning efficiency in NOMA networks.
problem Efficiently scheduling and allocating power for federated learning in bandwidth-limited wireless networks.
method Proposed a scheduling policy and power allocation scheme using NOMA to maximize data rate and convergence speed.
result Simulation results show improved federated learning accuracy in NOMA networks.
Game theory models incentivizes honesty in collaborative learning among competitors.
problem Incentivizing honest updates among competitors in collaborative learning schemes.
method Formulated a game to model interactions, studied two learning tasks, proposed mechanisms to incentivize honest communication.
result Rational clients are incentivized to manipulate their updates, preventing learning; proposed mechanisms ensure comparable learning quality to full cooperation.
A framework for partially encrypted machine learning using functional encryption.
problem Performing machine learning on encrypted data without revealing sensitive information.
method Combining adversarial training and functional encryption to efficiently compute quadratic functions and prevent feature leakage.
result The proposed framework maintains high model accuracy while significantly improving data privacy.
Paper explores weighted averaging schemes for SGD, achieving asymptotic normality and optimality.
problem Improving convergence of SGD in various settings.
method Develops a general weighted averaging scheme for SGD and establishes asymptotic normality.
result Establishes asymptotic normality and optimality of weighted averaged SGD solutions.
Neural Turing Machines (NTMs) are an instance of Memory Augmented Neural Networks, a new class of recurrent neural networks which decouple computation from memory by introducing an external memory unit. NTMs have demonstrated superior performance over Long Short-Term Memory Cells in several sequence learning tasks. A n…
Extends JKO scheme for iterative algorithms with unknown parameters.
problem Computational and statistical analysis of iterative algorithms with unknown parameters.
method Develops statistical methods to estimate unknown parameters and adapts JKO scheme.
result Establishes asymptotic theory for the statistical JKO scheme.
Data coarse graining improves model performance by filtering out less relevant features.
problem Lossy data transformations lose information but can improve model generalization.
method Data coarse graining schemes that systematically discard features based on relevance to the learning task.
result A 'high-pass' scheme helps models generalize better by filtering out less relevant features.