Hi! I'm an assistant professor of Operations Research and Information
Engineering (ORIE) at Cornell Tech and Cornell University, and an ORIE, Computer Science, and Information Science field member.
Our work: "Full Stack, Public Interest AI"
I develop public interest AI systems and analyze the "human-AI-society" relationship. Methodologically, my work spans the research-to-deployment pipeline: formulating mathematical models; developing statistical, optimization, and modern language modeling methods; conducting empirical analyses; building and deploying systems; evaluating those systems through RCTs and qualitative interviews; and policy analysis.
As detailed below, my research spans three broad, overlapping areas:
See a recent
"manifesto" on the challenges caused by
heterogeneous participation in participatory systems, which also surveys my work broadly. See a recent
talk video.
Our work has received several awards, including the NSF CAREER, William T. Grant Foundation Scholars Award, INFORMS George Dantzig Dissertation award, ACM SIGecom
Dissertation Award (Honorable Mention), Forbes 30 under 30 for Science, the NSF graduate research fellowship, and paper awards from EC, CSCW, INFORMS, and others. It has been supported by the Sloan Foundation, NSF, NASA, the Cornell Tech Urban Tech Hub, Google, Meta, and
Amazon.
Full bio.
Human-Algorithmic Market Design
We design, analyze, and deploy recommender systems, with thousands of active users in high-stakes public interest
settings. We complement deployments with theoretical modeling of phenomena such as strategic behavior, fairness, and homogeneity.
High school applications (with NYC Public Schools). We are studying disparities in how students apply to high schools in New York City, as a result of a complex process (Nature Cities, Accepted). For the 2025 cycle, we worked with NYC to help students in the application process (ACM EC, Best Paper with a Student Lead Author).
Bluesky algorithmic feeds. We are building feeds on Bluesky for the academic community and beyond! We have launched Paper Skygest; it has about 8,000 daily feed views by about 1,100 daily active users, with over 3 million overall views. We'd love for you to use it! We have a paper describing the feed, and are currently experimenting with novel feed algorithm designs.
Platform to help place discharged hospital patients into long-term-care facilities. In Hawaiʻi, a PhD advisee built and deployed a platform to help place discharged hospital patients into one of more than 1,000 long-term-care facilities, many of which are run by single individuals out of their homes. The platform texts homes to ask for updated capacity and preference information, and then provides this information to about 10 hospital social workers; it has helped place hundreds of patients (CSCW, Best Paper Award). We also experimentally studied how preferences and incentives shape placements (CSCW, Impact Recognition), and then deployed and evaluated active information acquisition with language model support.
Theory. We theoretically model algorithmic monoculture and the wisdom of crowds in matching markets (NeurIPS; ACM EC) and the downstream strategic, diversity, and fairness implications of ranking algorithms in recommender ecosystems (NeurIPS; The Web Conference).
Foundations and Systemic Effects of Modern AI
We develop language modeling methods to use text as data and study the systemic effects of LLMs.
Methods and applications of language models. During my PhD, I developed methods to use word embeddings to study historical societal stereotypes (PNAS). LLM-based word embeddings have since dramatically improved, but still face an interpretability challenge: individual vector dimensions are not inherently interpretable, complicating their use. More recently, we used a modern LLM interpretability method — sparse autoencoders —
to tackle this issue (ICML; ICML). We have since begun applying these techniques to social-science questions and studying their theoretical foundations through connections to compressed sensing (COLT).
LLM homogeneity and its market implications.
Algorithmic homogeneity is a concern in the LLM age: even as AI systems improve individual decisions, they may homogenize outputs across users,
increasing systemic risk and reducing intellectual diversity. We showed that LLMs make correlated errors, with implications for LLM-as-judge and hiring markets (ICML). On the other hand, homogeneity may also be leveraged to detect LLM text in a semi-supervised manner, using test-time adaptation for robustness to distribution shift and strategic behavior.
Operations and Algorithms for Government
We develop statistical and optimization tools for government operations and policy design.
Resident crowdsourcing, with the NYC Department of Parks and Recreation (Talk video). Do some neighborhoods report more than others for the same underlying conditions on the ground, thus receiving better government services? Answering this question requires new methods because we do not directly observe those conditions. First, we used duplicate 311 reports to estimate reporting rates without external ground-truth data, finding that higher-income and more educated neighborhoods report the same types of problems faster (Nature Computational Science). We also leveraged spatial correlation (AAAI), combined government ratings and crowdsourced reports (AAAI), and used vision-language models (Nature Communications) to identify incidents and quantify underreporting. We have further optimized downstream decision-making, including capacity plans and service-level agreements (ACM EC) and built and transferred a data dashboard and nine-year tree-planting scheduler to NYC DPR.
Other applications. With the New York Public Library, we analyzed how heterogeneous use of digital holds leads the most desired books to flow from low-income to high-income branches (AAAI), and then developed an operational intervention to mitigate this inequity by optimizing from which branch to pull each requested book. We have also used operations research tools for policy design: showing how multi-member districts with ranked-choice voting can curtail partisan gerrymandering (Operations Research; ACM EC), and studying equity in congestion pricing (ACM EC).