Acknowledgements#
Papers where I am credited for infrastructure contributions.
Virchow: A Million-Slide Digital Pathology Foundation Model#
E. Vorontsov, et al. · arXiv preprint, 2023
A 632M-parameter self-supervised vision transformer for computational pathology, trained on a million-slide whole-slide image archive from Memorial Sloan Kettering Cancer Center. Later published in Nature Medicine (2024).
Contribution: Built and operated the GPU compute infrastructure and high-throughput storage environment used to train and validate the model.
Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology#
E. Zimmermann, S. Liu, et al. · arXiv preprint, 2024
Three vision transformer foundation models trained on 3.1 million histopathology whole-slide images: Virchow2 (632M), Virchow2G (1.9B), and Virchow2G Mini (22M distilled). State-of-the-art on 12 tile-level tasks.
Contribution: Designed and operated the GPU compute infrastructure and high-throughput storage environment used to train and validate all three models.
PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology#
S. Liu, et al. · arXiv preprint, 2024
A multi-modal generative foundation model operating at the slide level for computational pathology, jointly modeling histology and clinical text.
Contribution: Built and maintained the HPC infrastructure supporting large-scale whole-slide image preprocessing, model training, and validation.
PRISM2: Unlocking Multi-Modal General Pathology AI with Clinical Dialogue#
E. Vorontsov, G. Shaikovski, A. Casson, et al. · arXiv preprint, 2025
A multi-modal slide-level pathology foundation model trained on 700,000 diagnostic specimen-report pairs and 14M question-answer pairs, the largest vision-and-language histopathology dataset to date. Aligns histomorphologic features with diagnostic language via clinical-dialogue supervision.
Contribution: Built and maintained the HPC infrastructure supporting large-scale whole-slide image preprocessing, model training, and validation.
Screen Them All: High-Throughput Pan-Cancer Genetic and Phenotypic Biomarker Screening from H&E Whole Slide Images#
Y. K. Wang, L. Tydlitatova, J. D. Kunz, et al. · arXiv preprint, 2024
OmniScreen, a high-throughput system predicting a broad range of clinically relevant molecular biomarkers across cancers from H&E whole-slide images, built on Virchow2 embeddings from 60,529 patients with paired MSK-IMPACT panels.
Contribution: Built and maintained the HPC infrastructure supporting large-scale whole-slide image preprocessing, embedding generation, model training, and validation.
High Speed Scientific Data Transfers Using Software Defined Networking#
H. Newman, et al. · INDIS ‘15 (SC15), Austin, Texas
Second Workshop on Innovating the Network for Data-Intensive Science, co-located with SC15: The International Conference for High Performance Computing, Networking, Storage and Analysis.
Contribution: Built and operated the SDN testbed and high-speed data-transfer infrastructure used in the demonstrations.
Co-authored#
Papers I co-authored from the Caltech CMS / Large Hadron Collider era.
SDN-NGenIA: A Software Defined Next Generation Integrated Architecture for HEP and Data Intensive Science#
J. Balcas, et al. · Journal of Physics: Conference Series, Vol. 898 (CHEP 2016)
A software-defined next-generation integrated architecture supporting high-energy physics and data-intensive science. Presented at the 22nd International Conference on Computing in High Energy and Nuclear Physics (CHEP 2016), San Francisco.
HTTP as a Data Access Protocol: Trials with XrootD in CMS’s AAA Project#
J. Balcas, et al. · Journal of Physics: Conference Series, Vol. 898 (CHEP 2016)
Evaluation of HTTP as a data access protocol for the CMS Any-data Anytime Anywhere (AAA) project, comparing performance and operational characteristics against XrootD’s native protocol.