Sitemap

A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.

Pages

Posts

Future Blog Post

less than 1 minute read

Published:

This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.

Blog Post number 4

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 3

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 2

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 1

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

portfolio

publications

JAWS: Auditing Predictive Uncertainty Under Covariate Shift

Published in Advances in Neural Information Processing Systems (NeurIPS), 2022

JAWS is a collection of wrapper methods for distribution-free predictive inference when the common data exchangeability (e.g., i.i.d.) assumption is violated due to shifts in the input data distribution (standard covariate shifts). JAWS is based on our core method JAW–the JAckknife+ Weighted for standard covariate shift–and also includes computationally efficient Approximations of JAW (JAWA) using higher-order influence functions.

Recommended citation: Prinster, D., Liu, A., & Saria, S. (2022). JAWS: Auditing Predictive Uncertainty Under Covariate Shift. In Advances in Neural Information Processing Systems. https://papers.nips.cc/paper_files/paper/2022/hash/e944bacecce6b06374ac39b260348db0-Abstract-Conference.html

JAWS-X: Addressing Efficiency Bottlenecks of Conformal Prediction Under Standard and Feedback Covariate Shift

Published in The Fortieth International Conference on Machine Learning (ICML), 2023

Accepted for an Oral Presentation at ICML 2023 (top ~2% of submissions) and building on our previous “JAWS” framework (NeurIPS 2022), this paper presents JAWS-X, a collection of methods for efficiently estimating predictive confidence intervals for black-box predictors under standard and feedback covariate shift. Our JAWS-X methods achieve distribution-free, finite-sample guarantees while flexibly balancing statistical and computational efficiency.

Recommended citation: Prinster, D., Saria, S., & Liu, A. (2023). JAWS-X: Addressing Efficiency Bottlenecks of Conformal Prediction Under Standard and Feedback Covariate Shift. In International Conference on Machine Learning. PMLR. https://proceedings.mlr.press/v202/prinster23a.html

Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)

Published in The International Conference on Machine Learning (ICML), 2024

Paper at ICML 2024. Demonstrates how conformal prediction can theoretically extend to any data distribution (i.e., not only exchangeable or quasi-exchangeable ones), with practical experiments focused on common settings of AI/ML agents including multiround synthetic protein design and active learning.

Recommended citation: Prinster, D., Stanton, S., Liu, A., & Saria, S. (2024). Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them). In International Conference on Machine Learning. PMLR. https://arxiv.org/abs/2405.06627

Care to Explain? AI Explanation Types Differentially Impact Chest Radiograph Diagnostic Performance and Physician Trust in AI

Published in Radiology, 2024

In this multisite prospective study of simulated artificial intelligence (AI)–assisted chest radiograph diagnosis involving 220 physicians, AI explanation type (local vs global) differentially impacted physician diagnostic performance and trust in AI advice, even if physicians were not aware of these effects.

Recommended citation: Prinster, D., Mahmood, A., Saria, S., Jeudy, J., Lin, C. T., Yi, P. H., & Huang, C.-M. (2024). Care to Explain? AI Explanation Types Differentially Impact Chest Radiograph Diagnostic Performance and Physician Trust in AI. Radiology, 313(2), e233261. https://pubs.rsna.org/doi/abs/10.1148/radiol.233261

WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales

Published in In International Conference on Machine Learning (ICML), 2025

Paper at ICML 2025. This paper develops a framework for the continual safety monitoring of AI deployments called “WATCH” (for Weighted Adaptive Testing for Changepoint Hypotheses). This framework is centered on a weighted generalization of conformal test martingales (WCTMs). WATCH addresses three main challenges in post-deployment AI monitoring: (1) Adaptation: WATCH enables monitoring under test-time adaptation to mild (covariate) shifts, to minimize unnecessary alarms. (2) Fast Detection: Empirically, WATCH rapidly detects more extreme or harmful shifts. (3) Root-Cause Analysis: WATCH aids in diagnosing the root-cause of performance degradation (as a covariate shift in $X$, or a concept shift in $Y \mid X$).

Recommended citation: Prinster, D., Han, X., Liu, A., & Saria, S. (2025). WATCH: Adaptive monitoring for AI deployments via weighted-conformal martingales. International Conference on Machine Learning (ICML). arXiv preprint arXiv:2505.04608. https://arxiv.org/abs/2505.04608

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing

Published in arXiv preprint arXiv:2512.03109, 2025

Converts agent verifier (e.g., LLM judge) scores into decision rules with statistical guarantees by framing trajectory evaluation as a sequential hypothesis testing problem, which enables continuous monitoring of agent action trajectories. Improves false alarm control and detection power while allowing early termination of unsuccessful trajectories to conserve computation.

Recommended citation: Sadhuka, S., Prinster, D., Fannjiang, C., Scalia, G., Berger, B., Regev, A., & Wang, H. (2025). E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing. arXiv preprint arXiv:2512.03109. https://arxiv.org/abs/2512.03109

Conformal Policy Control

Published in International Conference on Machine Learning (ICML), 2026

A framework to directly instruct an arbitrary AI agent on how far to advance its capabilities while still guaranteeing respect for user-specified safety constraints (e.g., safety violations <= 5%). The key requirements are access to any “safe policy” (e.g., a previous agent version verified to be safe) and any “optimized policy” (e.g., a new agent version trained for capabilities but not yet verified for safety). Conformal policy control (CPC) then searches “between” these policies to find an intermediate one that is optimized but guaranteed to respect the safety constraint. Experiments on applications ranging from medical question answering to biomolecular engineering show it is not only possible to safely advance AI abilities from the first moment of deployment, but that such safety can even counterintuitively improve performance.

Recommended citation: Prinster, D., Fannjiang, C., Park, J. W., Cho, K., Liu, A., Saria, S., & Stanton, S. (2026). Conformal Policy Control. International Conference on Machine Learning (ICML). arXiv preprint arXiv:2603.02196. https://arxiv.org/abs/2603.02196

talks

teaching

Teaching experience 1

Undergraduate course, University 1, Department, 2014

This is a description of a teaching experience. You can use markdown like any other post.

Teaching experience 2

Workshop, University 1, Department, 2015

This is a description of a teaching experience. You can use markdown like any other post.