Paper proposes a framework for privacy-preserving machine learning.
problem Extracted traces from pervasive systems violate user privacy.
method Interpretable machine learning framework for privacy preservation.
result Users can understand how their privacy is compromised by machine learning.
Designs a single-sensor system for accurate lying posture detection.
problem Designing efficient in-bed lying posture tracking systems.
method Single accelerometer with machine learning algorithms (deep learning and traditional classification).
result Single accelerometer can accurately detect lying postures with high F-Score.
Work addresses long-term accuracy issues in IoT air quality sensors.
problem Limited accuracy of IoT air quality sensors in long-term field deployments.
method Adaptive machine learning strategies for network calibration.
result Prolongs the validity of multisensor calibration models for continuous learning.
Confounding biases recommender systems, even when data seems fully observed.
problem Unmeasured features influencing both treatment and outcome in recommender systems.
method Illustrations and simulation studies showing how common practices introduce confounding.
result Standard recommender system practices can introduce confounding, reducing performance.
AI attacks threaten insurance systems, requiring new defenses.
problem Adversarial attacks on AI in insurance.
method Categorize and discuss various types of attacks and defense methods.
result Need for improved AI systems to resist attacks.
System recommends workouts and predicts success rates using RNNs.
problem Promoting healthy lifestyles through personalized exercise recommendations.
method Two interconnected recurrent neural networks (RNNs) using historical workout data.
result Interconnected-RNN model predicts exercise success rates with improved accuracy.
The paper introduces metrics and a regularization method to improve fairness in recommendation rankings.
problem Fairness risks in recommender systems that match users to products or information.
method Pairwise comparisons from randomized experiments to quantify and address fairness concerns in rankings.
result Improvement in pairwise fairness of a production recommender system through regularization.
AI-driven framework optimizes MCMC-based preconditioners for faster linear system solving.
problem Slow convergence of Krylov subspace solvers for ill-conditioned matrices.
method Graph neural surrogate and Bayesian optimization for AI-tuned MCMC parameters.
result 50% reduction in iterations to convergence on unseen system.
Industrial-scale podcast recommender system optimizes long-term listening journeys.
problem Optimizing long-term listening experiences in podcast recommendation systems.
method Reinforcement learning approach to optimize user listening journeys over months.
result Significantly improved long-term performance in A/B tests compared to short-term metrics.
Mobile app for neonatal EEG interpretation helps non-experts diagnose brain health.
problem Limited EEG interpretation skills among neonatal healthcare professionals.
method Low-cost, low-power EEG acquisition system with AI-assisted sonification.
result Improves diagnostic capabilities of non-expert clinicians.
Paper presents CeNN quantization for efficient CPS applications.
problem Efficient processing for CPS applications, especially in telemedicine and ADAS.
method Incremental quantization of CeNNs with various strategies.
result Achieved up to 7.8x speedup with no performance loss.
The paper aims to define a benchmark for deep learning recommendation models.
problem Insufficient benchmarking for deep learning recommendation models.
method Synthesizes modeling strategies, defines desirable characteristics, and summarizes advice from the MLPerf Recommendation Advisory Board.
result Defines an industry-relevant benchmark for deep learning recommendation models.
Study improves voice disorder detection system robust to channel effects.
problem Voice signals are sensitive to recording devices.
method Bidirectional LSTM network with domain adversarial training (DAT).
result Increased PR-AUC from 0.8448 to 0.9455 (and 0.9522 with labels).
The paper tackles selective labels in decision making, proposing a data augmentation approach to mitigate bias and discrimination.
problem Selective labels cause bias in decision making, making standard bias correction methods ineffective.
method Proposes a data augmentation approach to leverage expert consistency or empirically validate models under selective labels.
result Data augmentation can mitigate the bias caused by selective labels and prevent unreliable models.
A new method detects hidden driving forces in systems with multiple observables.
problem Hidden driving forces in systems with multiple observables cannot be detected by scalar statistics.
method Cross-spectral witness for hidden nonequilibrium.
result Two simultaneously observed channels retain an off-diagonal cross-spectral sector inaccessible to scalar reductions.
New online learning algorithms improve cyberattack detection in industrial control systems.
problem Detecting cyberattacks in industrial control systems with limited resources.
method Online learning algorithms to process continuous data streams and address class imbalance.
result Improved detection rate of cyberattacks in industrial control systems.
Paper proposes synthetic data generator to study and mitigate bias in machine learning.
problem Bias in machine learning data can lead to unfair outcomes.
method Developed a synthetic data generator to introduce and analyze various types of bias.
result Demonstrated how synthetic data can be used to study and mitigate bias in machine learning models.
Framework classifies machine learning failures into intentional and unintentional.
problem Understanding and preventing failures in machine learning systems.
method Developed a taxonomy of machine learning failure modes.
result Stakeholders found the framework useful for discussing machine learning failures.
This paper tackles online strategic decision making with asymmetry and knowledge transportability.
problem Strategic decision making with information asymmetry and knowledge transportability challenges.
method Developed a sample-efficient algorithm for online learning under these conditions.
result Proved sample complexity of O(1/ε2) for learning an ε-optimal policy. DQN optimizes traffic light control policies in ITSs.
problem Challenges in scalable real-time actuation mechanisms for smart traffic management.
method Exploration of Deep Q-Networks (DQN) for traffic light control policies.
result DQN algorithms produce intelligent behavior, such as greenwave patterns.
A new distribution fixes a common error in VAEs, improving image quality.
problem Using a Bernoulli likelihood for pixel data in VAEs.
method Introducing the continuous Bernoulli distribution.
result The continuous Bernoulli improves image quality across various metrics and datasets.
NCPF model improves traffic data imputation with neural and tensor methods.
problem Pervasive missing data in traffic analysis due to sensor failures and gaps.
method Neural Canonical Polyadic Factorization (NCPF) integrating CP decomposition and deep learning.
result NCPF outperforms state-of-the-art baselines in urban traffic datasets.
A wrist-worn device system for user authentication using writing behavior analysis.
problem User authentication through writing behavior for wearable devices.
method Dynamic Time Wrapping and Savitzky-Golay filter for fine-grained writing metrics.
result The proposed system achieves high accuracy in user identification with low false-positive and false-negative rates.
The paper finds a pervasive and severe bias in accounting semi-identity models.
problem Bias in investment-cash flow sensitivity models.
method Augmented specification with a bias-capturing variable tested across multiple databases.
result The Accounting Semi-Identity (ASI) distortion is universal and severe, affecting 100% of databases and explaining more than 83% of total explained variance.
Directed networks are pervasive both in nature and engineered systems, often underlying the complex behavior observed in biological systems, microblogs and social interactions over the web, as well as global financial markets. Since their structures are often unobservable, in order to facilitate network analytics, one …
Paper introduces kernel methods for detecting anomalous changes in remote sensing imagery.
problem Detecting anomalous changes in remote sensing imagery.
method Nonlinear extension of Gaussian and elliptically contoured distribution algorithms using reproducing kernel Hilbert space.
result Improved detection accuracy and reduced false-alarm rates compared to linear formulations.
Develops methods to estimate high rank tensors from noisy data.
problem Estimating high rank tensors from noisy observations.
method Generative latent variable tensor model, polynomial-time spectral algorithm.
result Achieves computationally optimal rate for signal tensor estimation.
Estimates peeking effects in p-values to correct bias.
problem Data peeking biases reported p-values downward.
method Develops mechanisms to estimate running extrema of test statistics.
result Corrects bias in p-values due to peeking.
Matrices of (approximate) low rank are pervasive in data science, appearing in recommender systems, movie preferences, topic models, medical records, and genomics. While there is a vast literature on how to exploit low rank structure in these datasets, there is less attention on explaining why the low rank structure ap…
Method discovers user habits from mobile data.
problem Understanding human mobility patterns and habits.
method Density-based clustering for spatio-temporal data and Gaussian Mixture Model (GMM).
result Many unique habits were identified from the datasets.
AquaSight detects water impurity using deep learning.
problem Water pollution and lack of affordable water quality assessment.
method Convolutional Neural Networks for automated water impurity detection.
result Deep learning model achieved 96% accuracy in detecting water contamination.
Research develops a DSS for stock selection and asset allocation using fundamental data.
problem Complex financial markets and limited use of fundamental data analysis.
method Data gathering, cleaning, and modeling of fundamental data; integration with macroeconomic conditions.
result Enhanced predictive model for mid- to long-term stock returns.
Study detects and explains positional bias in financial LLMs.
problem Positional bias in financial decision-making using LLMs.
method Unified framework and benchmark for detecting and quantifying bias in Qwen2.5 models.
result Positional bias is pervasive, scale-sensitive, and resurfaces under nuanced prompt designs.
We break down transformer embeddings into interpretable components revealing hidden geometric structures.
problem Understanding the hidden geometry and interpretability of transformer models.
method Decomposed transformer embeddings into position, context, and residual components.
result Pervasive mathematical structure in transformer embeddings, including position and context vectors.
Paper presents an energy-efficient RL method for sensor networks.
problem Energy consumption in sensor networks for health monitoring.
method Adaptive Reinforcement Learning framework using SARSA algorithm.
result Achieves performance enhancement and energy savings over time.
Proposes Pareto efficient fairness for supervised learning models.
problem Ensuring fairness in machine learning models without sacrificing accuracy.
method Formulates a bilevel optimization problem to find Pareto efficient classifiers.
result Guaranteed solution on Pareto frontier for convex and non-convex objectives.
This thesis tackles bias in AI decision-making in banking.
problem Bias in AI-driven banking decisions.
method Understanding, mitigating, and accounting for bias in AI systems.
result Establishment of Responsible AI practices for fair decision-making.
COHORTNEY groups web users based on activity patterns.
problem Lack of academic discussion on cohort analysis for user behavior.
method Unsupervised non-parametric machine learning approach.
result COHORTNEY outperforms traditional methods in cohort analysis.
For applications in computing, Bezier curves are pervasive and are defined by a piecewise linear curve L which is embedded in R^3 and yields a smooth polynomial curve C embedded in R^3. It is of interest to understand when L and C have the same embeddings. One class of counterexamples is shown for L being unknotted, wh…
Study shows skewed data labels significantly impact decentralized ML accuracy.
problem Skewed data labels across devices/locations cause significant accuracy loss in decentralized ML.
method Detailed experimental study on skewed data labels, presenting SkewScout system-level approach.
result Skewed data labels are a fundamental challenge for decentralized learning, affecting many applications and models.
Study reveals pervasive label errors in test sets, affecting machine learning benchmarks.
problem Label errors in test sets destabilize machine learning benchmarks.
method Identified label errors in 10 common datasets using confident learning algorithms and human validation.
result Lower capacity models may be more useful in real-world datasets with high proportions of erroneously labeled data.
While the use of volatilities is pervasive throughout finance, our ability to determine the instantaneous volatility of stocks is nascent. Here, we present a method for measuring the temporal behavior of stocks, and show that stock prices for 24 DJIA stocks follow a stochastic process that describes an efficiently pric…
New method estimates treatment effects in complex interference settings.
problem Challenges in estimating treatment effects due to unknown interference.
method Higher-order causal message passing for non-linear feature learning.
result Effective estimation of treatment effect dynamics in complex interference.
New model values CDS contracts considering multiple credit risks and collateralization.
problem Valuation of CDS contracts affected by multiple credit risks and collateralization.
method Developed a new model to value CDS contracts, considering default dependency and collateralization.
result Default dependency significantly impacts asset pricing and full collateralization does not eliminate counterparty risk.
WiPIN uses Wi-Fi signals to identify people without requiring them to walk.
problem Identification requires walking and is unreliable with many users.
method Extracts body information from Wi-Fi signals without user movement.
result Achieves 92% accuracy with 30 users, robust to various settings.
A hybrid neural network optimizes AI deployment on edge and cloud for energy efficiency.
problem Energy and resource constraints in edge devices for deep learning models.
method Conditionally deep hybrid neural network with quantized layers at edge and full-precision layers at cloud.
result Early classification at the edge reduces energy consumption by 5.5x on CIFAR-10 dataset.
Unified framework TSS certifies robustness against semantic transformations.
problem Certifying robustness of ML models against semantic transformations.
method Unified framework TSS categorizes transformations into resolvable and differentially resolvable, proposing randomized smoothing and stratified sampling strategies.
result Significantly outperforms state of the art on over ten types of semantic transformations.
Making inferences from data streams is a pervasive problem in many modern data analysis applications. But it requires to address the problem of continuous model updating and adapt to changes or drifts in the underlying data generating distribution. In this paper, we approach these problems from a Bayesian perspective c…