You can find my working papers and projects below. My research focuses on how we can leverage AI to best benefit those who use it - with evidence from unique, granular datasets.
Editing an LLM draft can feel better compared to writing from scratch — but sensors and physiological metrics on the body show that fatigue is accumulating even faster.
The use of large language models (LLMs) has grown rapidly over the past few years, especially in health care, where documentation burden remains a persistent problem for practicing clinicians. However, despite the deployment of many ambient scribes for documentation, it is still unclear how LLMs impact those who work with them. Across two preregistered studies, we examine how this shift—from writing documentation to editing an LLM-generated draft—impacts workers. We study this “first draft problem”: eliminating the blank page does not remove the work of documentation but relocates it into revising and editing. In Study 1, an online experiment with 16 clinicians producing documentation following a teletherapy patient encounter, we establish the puzzle: counter to our hypotheses, editing an LLM-generated draft of the documentation as opposed to writing it from scratch does not result in time savings or quality improvements on average. Instead, while some clinicians experience no time savings but large improvements in quality after making substantial revisions, others experience meaningful time savings and no change in quality after making limited revisions. In other words, we observe significant heterogeneity in how clinicians engage with the same draft, which leads to variation in productivity and quality. In Study 2, a lab experiment with 100 university students summarizing video interviews, we unpack the mechanisms underlying this heterogeneity using granular keystroke logs, validated self-reports of workload, and micro-level physiological and eye-tracking measurements. In this setting, we find the expected productivity improvements: editing reduces task completion time by roughly 70%, raises quality, and lowers self-reported workload. Yet the physiological evidence tells a more nuanced story. Physiological markers indicate that fatigue accumulates faster when editing an LLM-generated draft than when writing from scratch. Thus, we find that (a) time savings can be real, but are setting dependent; (b) quality improvements are possible, but the relationship between how heavily a draft is revised and the quality of the result differs sharply across our two settings; and (c) while editing an LLM draft lowers reported cognitive load, the body still tires faster over time.
Using wearable sensors from 400+ Air Force training flights, we develop a framework that learns the pilot's state to intervene before danger occurs. We focus on uncovering the latent state to determine crucial moments that can make or break a training flight.
Operational decision-makers often cannot perform well without substantial experience, and learning is slow when feedback is delayed and attention is a scarce resource allocated over many competing priorities. We propose to augment the decision-making process with intelligent contextual cues, identified from data, that train the decision-maker to achieve rapid optimality. We study this AI-based augmentation in the context of fatigue and safety risk management for military aviators, where latent physiological vulnerability can build during a sortie and rarely surfaces as a safety hazard—but when a hazard does occur, the risks are extreme. Partnering with a company that outfits aviators at 22 U.S. Air Force bases with wearable sensors recording synchronized multistream data, we construct a label-free risk proxy that flags “high-stress” episodes in which physiological reactions (e.g., heart rate increases or blood oxygen decreases) are unusually strong relative to contemporaneous maneuver-induced stress; we discover a small number of interpretable situations from the sensor streams using self-supervised learning; and we design a budget-constrained intervention policy through a decision-feedback training loop that upweights the episodes most influential for the policy. Two minutes prior to a hypoxia alarm, high-stress episodes appear 104% more often than in matched control windows with similar maneuver-induced stress, and given a hard budget of 5 interventions during a 2-hour flight, prescribed interventions decrease projected time spent in high-stress episodes by 14.20%.
Learning for long term retention has always been a difficult task: just how much practice can a platform require without driving users away? With our new memory model that balances learning and attrition, review burden falls by half — and students are projected to stick around for much longer.
Digital learning platforms face an important trade-off: raising student workload can improve learning but might reduce engagement. We study this trade-off on a large Chinese language-learning platform with thousands of active students and roughly 6 million reviews per month, where each learned item generates four related skills (reading, writing, definition, and tone), so the platform must jointly manage forgetting and cross-task spillovers. We first estimate student-specific memory and transfer parameters in an offline memory model, then design a two-stage adaptive review policy that allocates daily review capacity via fast meta-adaptation and selects items using a spillover-aware bandit algorithm. Incorporating contextual transfer improves predictive accuracy for 85% of students; personalization lengthens review intervals by 15.8% without lowering predicted recall; and policy simulations reduce reviews to 48% of baseline while reducing projected attrition by 11%.