Welcome!

My research develops computational methods that transform health data into trustworthy and clinically meaningful evidence for clinical decision support. Currently, I work as a research engineer at the Royal Melbourne Hospital, dedicated to improving the quality of hospital care for patients with dementia. In addition, I am an independent peer reviewer for multiple journals and conferences, with more than 40 completed reviews.

📝 Publications

IJCAI 2026
sym

X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models

  • Foundation Models for Electronic Health Records (FEMRs) are pretrained on large-scale structured patient data, enabling them to convert longitudinal patient trajectories into generalizable representations for diverse clinical prediction tasks. Despite their effectiveness, FEMRs remain black-box models, raising concerns about bias, interpretability, and clinical trust. To address this, we propose the first token-level explainability approach for FEMRs. We train a Transformer-based surrogate model on input-output pairs from the FEMR across two prediction tasks, approximating its behavior while preserving temporal dynamics. We identify the most influential tokens, providing insights into how FEMRs leverage different aspects of patient history for predictions. To evaluate clinical relevance, we introduce a novel clinical alignment metric that quantifies the correspondence between the surrogate model’s key tokens and clinically validated features. Our results demonstrate that the surrogate closely approximates FEMR predictions and that token-level explanations align well with clinical knowledge, offering a practical framework for interpretable and trustworthy clinical AI.

Reference: Huang, J.†, Yin, P.†, Xu, Z., Capurro, D., Conway, M., & Dang, T. (2026). X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models. arXiv preprint arXiv:2607.06163.
† These authors contributed equally. Accepted by In Proceedings of International Joint Conferences on Artificial Intelligence(IJCAI 2026)(AI and Health Special Track, Accept Rate:18%)

JBI 2025
sym

Measuring and Visualizing Healthcare Process Variability

  • Understanding factors that contribute to clinical variability in patient care is critical, as unwarranted variability can lead to increased adverse events and prolonged hospital stays. Determining when this variability becomes excessive can be a step in optimizing patient outcomes and healthcare efficiency. In this study, we presents a standardized way to measure and visualize variability in clinical processes and measure its impact on patient-relevant outcomes.

Reference: Yin, P., Cervantes, A. A., & Capurro, D. (2025). Measuring and visualizing healthcare process variability. Journal of biomedical informatics, 104918. Advance online publication. https://doi.org/10.1016/j.jbi.2025.104918

BIBM 2023
sym

Parsing Eligibility Criteria for Cohort Query by a Multi-Input Multi-Output Sequence Labeling Model

  • To enable electronic screening of eligible patients for clinical trials, free-text clinical trial eligibility criteria should be translated to a computable format. Natural Language Processing (NLP) techniques have the potential to automate this process. In this study, we explored a supervised multi-input multi-output (MIMO) sequence labelling model to parse eligibility criteria into combinations of fact and condition tuples. Our experiments on a small manually annotated training dataset showed that that the performance of the MIMO framework with a BERT-based encoder using all the input sequences achieved an overall lenient-level AUROC of 0.61.
  • This study was funded in part by the National Institutes of Health (NIH) through awards: R21 AG068717 and R21 CA253394.

Reference: Tian, S., Yin, P., Zhang, H., Erdengasileng, A., Bian, J., & He, Z. (BIBM 2023). Parsing Clinical Trial Eligibility Criteria for Cohort Query by a Multi-Input Multi-Output Sequence Labeling Model.2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Istanbul, Turkiye, pp. 4426-4430

AMIA 2021
sym

Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets

  • During the coronavirus disease pandemic (COVID-19), social media platforms such as Twitter have become a venue for individuals, health professionals, and government agencies to share COVID-19 information. Twitter has been a popular source of data for researchers, especially for public health studies. However, the use of Twitter data for research also has drawbacks and barriers. Biases appear everywhere from data collection methods to modeling approaches, and those biases have not been systematically assessed. In this study, we examined six different data collection methods and three different Machine Learning (ML) models-commonly used in social media analysis-to assess data collection bias and measure ML models’ sensitivity to data collection bias. We showed that (1) publicly available Twitter data collection endpoints with appropriate strategies can collect data that is reasonably representative of the Twitter universe; and (2) careful examinations of ML models’ sensitivity to data collection bias are critical.
  • This work was supported in part by NSF Award #1734134

Reference: Zhao, Y., Yin, P., Li, Y., He, X., Du, J., Tao, C., Guo, Y., Prosperi, M., Veltri, P., Wu, Y., & Bian, J. (AMIA 2021). Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets. American Medical Informatics Association Annual Symposium - Computing Research and Education Association of Australasia (CORE) Ranking A

SEPDA 2021
sym

Towards Formal Computable Representation of Clinical Trial Eligibility Criteria for Alzheimer’s Disease

  • Ambiguity and misunderstanding of free-text clinical trial eligibility can affect the accuracy of translating trial investigators’ mental model of the study population to the correct cohort identification queries. In this pilot study, to eliminate the ambiguity when parsing eligibility criteria, we built ontology-based representations to standardize clinical trial eligibility criteria. We analyzed 10 Alzheimer’s disease (AD) trials’ eligibility criteria and categorized them into general query patterns using an annotation schema borrowed from the literature on constructing knowledge graphs. Then, for each pattern, we built the corresponding ontological representations, linked them to real-word electronic health record (EHR) data, and constructed cohort identification queries using the Neo4j graph database. Our evaluation results of these cohort queries verified the accuracy of our ontology representation; and interestingly, we found that graph-queries achieved better runtime performance for complex study traits.
  • This study is funded in part by the National Insulites of Health (NIH) through awards: R21 AG068717 and R21 CA253394.

Reference: Yin, P., Zhang, H., He, X., Diller, M., Li, Q., Tian, S., Erdengasileng, A., He, Z., Tao, C., & Bian, J. (SEPDA 2021). Towards Formal Computable Representation of Clinical Trial Eligibility Criteria for Alzheimer’s Disease. The 6th International Workshop on Semantics-Powered Health Data Analytics.

🎉 News

  • 2026.09:   Invited to serve as a reviewer for the NeurIPS 2026 Workshop on Agents in the Wild.
  • 2026.06:   Awarded IJCAI Travel Grant (USD 800).
  • 2026.05:   Invited to serve as a reviewer for the ICML 2026 Workshop on Agents in the Wild.
  • 2026.05:   Invited to serve as a reviewer for the CVPR 2026 Workshop Multi-Modal Reasoning for Agentic Intelligence(MMRAgI).
  • 2026.04:   Invited to serve as a reviewer for the Frontiers in Cardiovascular Medicine.
  • 2026.02:   Invited to serve as a reviewer for the ICLR 2026 Workshop AIWILD and ES-Reasoning.
  • 2026.02:   Invited to serve as a reviewer for the Journal of Pharmaceutical and Healthcare Marketing.
  • 2025.11:   Invited to serve as a reviewer for the BMC Medical Genomics,Discover Artificial Intelligence,The journal of Supercomputing.
  • 2025.10:   Invited to serve as a reviewer for the Journal of Biomedical Informatics.
  • 2023.04:   Invited to serve as a reviewer for the American Medical Informatics Association Annual Symposium(AMIA) 2023.

đź“– Educations

  • 2024.03 - 2026.12, Doctor of Philosophy in Engineering and IT , University of Melbourne, Australia.
  • 2019.08 - 2021.05, Master of Science in Electrical and Computer Engineering , University of Florida(GPA:3.8/4.0), United States.