New algorithms improve hiring decisions by optimizing tiered interview processes.
problem Optimizing tiered interview allocation for better hiring decisions.
method Casted tiered hiring as a combinatorial pure exploration (CPE) problem in the stochastic multi-armed bandit setting, presenting new algorithms in PAC and fixed-budget settings.
result Our algorithms select a near-optimal cohort with provable guarantees, making better hiring decisions or using less budget than the status quo.
Former physicists share insights on derivatives in interviews.
problem Understanding physics in finance interview questions.
method Interviews with former physicists in finance.
result Compilation of physics-related interview answers.
Tiered graph autoencoders improve molecular graph representation.
problem Representing and utilizing groups in molecular graphs.
method Adapting tiered graph autoencoders for PyTorch Geometric.
result Molecular graphs have tiered latent representations.
Study uses interviews to automatically detect BD and BPD with good accuracy.
problem Challenges in distinguishing BD and BPD from clinical interviews.
method Developed a multi-modal dataset and used a linear classifier with selected features from interviews.
result Different sets of features characterize BD and BPD, providing insights into their differences.
New method simplifies causal inference with tiered background knowledge.
problem Large equivalence classes of DAGs limit causal information.
method Integrates tiered background knowledge to create 'tiered MPDAGs' with simplified structure.
result Tiered MPDAGs are chain graphs with chordal components, simplifying causal effect estimation.
Model detects depression from transcribed interviews using affective language.
problem Detecting depression from transcribed clinical interviews.
method Hierarchical Attention Network with affective conditioning.
result Model achieves state-of-the-art F1 scores in depression detection.
Paper tackles robust knowledge transfer in parallel RL tasks.
problem Transfer knowledge from low-tier to high-tier tasks in parallel RL without shared dynamics or reward functions.
method Identifies Optimal Value Dominance condition and proposes online learning algorithms for both tasks.
result Achieves constant regret on partial states and near-optimal regret when tasks are dissimilar.
AI-assisted interviews allow respondents to describe experiences naturally, but mapping those accounts into structured survey variables is fallible.
problem Mapping AI-assisted interview responses into structured survey variables is fallible.
method Adaptive Matrix Validation (AMV) is proposed, which involves mapping responses into tabular data and using a small set of structured questions for statistical adjustment.
result The estimator calibrates mapped values using validation answers from other respondents and corrects remaining error with validation answers observed for the target respondent.
Interview study reveals considerations for designing semi-automated bias detection tools.
problem Detecting and mitigating machine learning biases.
method Interview study with 11 machine learning practitioners.
result Four considerations identified for tool design.
Multimodal analysis assesses job interview performance and provides feedback.
problem Assessing candidate performance in interviews for professional roles.
method Multimodal analytical framework using video, audio, and text data.
result The proposed methodology achieved promising results in predicting behavioral cues.
Paper presents new algorithms for causal discovery with latent variables and overlapping datasets.
problem Causal discovery with latent variables and overlapping datasets.
method Introduces tiered FCI and tIOD algorithms for constraint-based causal discovery.
result The tIOD algorithm is more efficient and informative than the IOD algorithm.
TiFL divides clients into tiers to improve federated learning performance.
problem Heterogeneity in resource and data quality impacts FL performance.
method TiFL divides clients into tiers based on training performance and selects clients from the same tier in each training round.
result TiFL achieves faster training performance with comparable or better test accuracy.
Industry lacks tools to secure ML systems, study finds.
problem Insufficient security tools for ML systems in industry.
method Interviews with 28 organizations to identify gaps.
result Need revised Security Development Lifecycle for ML.
Study blockchain's impact on primary financial market challenges.
problem Challenges of blockchain in securities issuance and trading.
method Hybrid method combining interviews and surveys.
result Complex due diligence, mismatch, and difficult monitoring are significant challenges.
Method augments CTNs for ICD coding with neural network imputation.
problem Time-consuming manual annotation of CTNs for ICD coding.
method Semi-self-supervised neural network imputation of clinical features.
result Data augmentation improves ICD coding performance significantly.
A new reinforcement learning framework separates users into risk-tolerant and risk-averse groups for better performance.
problem Improving performance for risk-averse users in reinforcement learning.
method Introducing a tiered reinforcement learning approach with two policies: πextO and πextE. result Achieving constant regret for risk-averse users, independent of the number of episodes.
Model analyzes RFQ markets using stochastic control to optimize dealer performance and inventory.
problem Optimizing market making in aggregator-routed RFQ markets with varying dealer performance scores.
method Two-tier stochastic control model that separates RFQ-level price competition from macro routing.
result Optimal controls can be expressed through derivatives of reduced Hamiltonians, leading to interpretable mappings from optimal win probabilities to optimal offsets.
ASCEND discovers causal relationships in multi-omics data by leveraging known hierarchical structure.
problem Causal inference in high-dimensional multi-omics data, especially when ignoring the hierarchical structure.
method Two-tiered divide-and-conquer strategy with ancestral conditioning sets.
result Achieves polynomial-time complexity and accurately recovers ancestral relationships.
TIER uses extended strain data to improve gravitational wave detection sensitivity.
problem Improving gravitational wave detection sensitivity using extended strain data.
method TIER framework using machine learning to capture extended strain data features.
result Up to 20% improvement in sensitive volume time in LIGO-Virgo-Kagra O3 data.
This paper optimizes UAV and FeICIC locations in a three-tier LTE-Advanced network.
problem Optimizing UAV and FeICIC locations in a three-tier LTE-Advanced network.
method Integrates UAVs as both UE and BS in LTE-Advanced HetNet, uses CRE, ICIC, 3D beamforming, and genetic algorithms for optimization.
result Heuristic algorithms outperform brute-force techniques in achieving better 5pSE and coverage probability.
This work facilitates ensuring fairness of machine learning in the real world by decoupling fairness considerations in compound decisions. In particular, this work studies how fairness propagates through a compound decision-making processes, which we call a pipeline. Prior work in algorithmic fairness only focuses on f…
Based on 46 in-depth interviews with scientists, engineers, and CEOs, this document presents a list of concrete machine research problems, progress on which would directly benefit tech ventures in East Africa.
Two-tier approach optimizes RL hyper-parameters for better agent learning.
problem Optimizing hyper-parameters in reinforcement learning to improve agent performance.
method Two-step optimization: first categorical hyper-parameters, then solution-level hyper-parameters.
result Promising results in simulated control tasks, suggesting user-independent reinforcement learning applications.
Young investors, especially students, dominate Indonesian stock exchanges.
problem Investment behavior of young and rookie investors in the stock market.
method Qualitative approach with descriptive analysis and interviews.
result Perception of behavioral control influences investment decisions.
New algorithm improves causal discovery in biomedical data.
problem Stability and accuracy issues in causal discovery algorithms.
method Exploits temporal structure and tiered background knowledge.
result Increases accuracy in finite samples for causal structure estimation.
A firm learns to efficiently screen candidates for interviews.
problem Minimizing interviews while ensuring optimal matching of candidates to departments.
method Online assignment problem with two variants: independent draws and access to a training set.
result Exponential reduction in the number of retained items with access to a training set.
A model for learning customer preferences in a dynamic product launch setting.
problem Learning customer preferences in a setting with new product launches.
method Proposes a sequential multinomial logit (SMNL) model and a learning algorithm with a regret bound.
result Demonstrates the tier structure can mitigate risks associated with learning new products.
The paper proposes Tier Balancing for dynamic fairness in decision-making.
problem Achieving long-term fairness in decision-making processes.
method Causal modeling with DAGs to investigate dynamic fairness.
result Tier Balancing is a more natural approach to achieve long-term fairness, capturing latent causal factors.
The paper audits trading filters, finding a high save-to-miss ratio.
problem Improving the efficiency and accuracy of trading filters in decentralized exchanges.
method A precision audit of filter rules against real trading data, classifying rejection events.
result Conservative save-to-miss ratio of 3.7 : 1, with wider interpretation of 14.8 : 1.
Achieving machine intelligence requires a smooth integration of perception and reasoning, yet models developed to date tend to specialize in one or the other; sophisticated manipulation of symbols acquired from rich perceptual spaces has so far proved elusive. Consider a visual arithmetic task, where the goal is to car…
Paper presents algorithm for optimal job selection with dynamic scoring.
problem Optimal job assignment in a sequential selection process with dynamic scores.
method Developed using dynamic programming, with extensions for partial and no-information cases.
result Algorithm allows for optimal job assignment with limited information.
ATMSeer improves AutoML by making it more transparent and controllable.
problem Users distrust automatic AutoML results and increase search budgets.
method Interactive visualization tool to refine search space and analyze results.
result ATMSeer enhances efficiency and user trust in AutoML.
This study examines non-retail trading on Polymarket, revealing unique behavior patterns and structural limitations.
problem Lack of address-level quote-lifecycle data in Polymarket prediction markets.
method Empirical analysis of 13 million order-filled events using DBSCAN clustering on a six-feature fill-side vector.
result Non-retail behavior is uni-modal, contradicting previous archetypal hypotheses.
Funnelling improves cross-lingual text classification accuracy.
problem Classifying documents in multiple languages more accurately than individual language classifiers.
method A two-tier classification system using posterior probabilities from language-dependent classifiers.
result Funnelling significantly outperforms state-of-the-art baselines in multilingual text classification.
The latest global financial tsunami and its follow-up global economic recession has uncovered the crucial impact of housing markets on financial and economic systems. The Chinese stock market experienced a markedly fall during the global financial tsunami and China's economy has also slowed down by about 2\%-3\% when m…
Procedure verifies if machine learning models assign fixed predictions that preclude access.
problem Models assign fixed predictions that preclude access to credit and employment.
method Model-agnostic recourse verification with reachable sets.
result Models can inadvertently preclude access by assigning fixed predictions.
Study analyzes factors affecting capital adequacy in Bangladesh's banks.
problem Factors influencing capital adequacy in commercial banks in Bangladesh.
method Fixed Effect, Random Effect, and Pooled Ordinary Least Square (POLS) methods.
result Several independent variables significantly affect capital adequacy, with specific relationships between leverage, liquidity risk, and other factors.
ActiLabel learns activity patterns across diverse sensor devices.
problem Limited adoption of activity recognition models across different domains due to diverse sensor devices.
method Combination of graph model and optimal tiered mapping for learning activity labels.
result Superior performance compared to state-of-the-art methods on public datasets.
We explore the effect of introducing prior information into the intermediate level of neural networks for a learning task on which all the state-of-the-art machine learning algorithms tested failed to learn. We motivate our work from the hypothesis that humans learn such intermediate concepts from other individuals via…
TEASER improves early time series classification accuracy and speed.
problem Early and accurate classification of time series data.
method TEASER models eTSC as a two-tier classification problem, using a first-tier classifier to assess class probabilities and a second-tier to decide reliability.
result TEASER is two to three times faster at predictions than competitors while maintaining or improving accuracy.
Study bond market making with hit-ratio target using optimal control and HJB equations.
problem Optimizing bond market making with hit-ratio target in OTC markets.
method Stochastic optimal control approach, dualizing hit-ratio target, HJB equation, Riccati equation, linearization.
result Explicit quote decompositions into riskless spread, inventory-risk correction, and hit-ratio correction.
This thesis evaluates text-based vs audio-based classification of mental health interviews.
problem Classifying psychiatric illness using text-based methods.
method Design and evaluate a text classification network on mental health interviews, using belabBERT.
result Text-based classification is a strong alternative to audio-based methods.
NFTs raise concerns like scams, racism, and sexism; centralization vs decentralization debate.
problem Concerns and value judgments of stakeholders in NFT market.
method Mixed quantitative and qualitative methods: social media analysis and interviews.
result Identified financial scams, counterfeit NFTs, hacking, and unethical NFTs as major issues.
These are the lecture notes for an advanced Ph.D. level course I taught in Spring'02 at the C.N. Yang Institute for Theoretical Physics at Stony Brook. The course primarily focused on an introduction to stochastic calculus and derivative pricing with various stochastic computations recast in the language of path integr…
Paper develops a method for valid inference using language model predictions from verbal autopsy narratives.
problem Valid inference from verbal autopsy narratives for public health decision-making.
method Develops multiPPI++ method for valid inference using NLP techniques for COD prediction.
result Demonstrates the effectiveness of multiPPI++ in handling transportability issues and recovering ground truth estimates.
New protocol identifies impossible edge orientations in causal graphs.
problem Causal-discovery algorithms cannot distinguish edge directions without assumptions.
method Discrete impossibility certificates and oracle queries.
result Upper bound of 1+K expert interactions for DAG recovery. TransFall uses transfer learning to improve activity recognition from mobile sensors.
problem Performance degradation due to platform and user movement differences.
method Two-tier data transformation, label estimation, and model generation layers.
result TransFall enhances activity recognition accuracy for new scenarios.
Method predicts NAFLD risk with high accuracy and distribution-free coverage guarantees.
problem Insufficient population-level screening tools for NAFLD.
method Gradient-boosted decision trees with conformal prediction.
result Method achieves AUROC of 0.912 internally and 0.891 externally, superior to other models.