👋 Lead Scientist — 💌 shahr2@mskcc.org — 🇺🇸 New York
Ronak H Shah
Bio
I am responsible for leading a team of Computational Biologists and Bioinformatics Software Engineers who develop, maintain, and operate bioinformatics pipelines and databases in the Center for Molecular Oncology. We also perform collaborative research with other labs and clinicians both within MSKCC and in the broader research community. On a daily basis, we analyze blood samples from patients with tissue-based cancers, using their circulating tumor DNA extracted from the blood, thereby avoiding the need for tumor biopsies. More specifically, I lead the team in designing, developing, and implementing software tools to process and analyze high-throughput next-generation sequencing data for liquid biopsy applications (MSK-ACCESS & Clonal Hematopoiesis Panel). Previously at MSK, I helped develop the workflow for analyzing MSK-IMACT data for both clinical and research implementation, which is still in use. You can read more about my background here. Also, here is the link to my google scholar profile
👋 Computational Biologist — 💌 charalk@mskcc.org — 🇺🇸 New York
I am a Computational Biologist as part of the Center of Molecular Oncology Informatics cfDNA group. I am responsible for the development of new and existing tools, maintaining the cfDNA pipeline, and processing samples for both clinical and research work. Previously, I have worked at the NHS trust Cambridge University Hospitals as a Bioinformatician with primary responsibility for the development of the Hemato-Oncology clinical assay. You can read more about my educational background, work experience, and research interests and you can access my Linkedin .
👋 Bioinformatics Engineer II — 💌 buehlere@mskcc.org — 🇺🇸 New York
The CMO Cell-Free DNA Informatics (CCI) group’s mission is to develop and apply computational methods to organize, analyze and understand genomic data generated from cfDNA assays such as MSK-ACCESS. The group is responsible for all computational infrastructure needed to deploy, run, and deliver results for research cfDNA assays in a production setting for CMO.
CCI's Mission
Our Values
Be Compassionate
We treat everyone we encounter with compassion, seeing the humanity behind their problems and experiences.
Be Mindful
We do not take advantage of our users' attention and adopt mindful working practices so that we can create safe spaces both in our working environment and in our products themselves.
Research First
We challenge our own and others' assumptions through qualitative and quantitative research. Not sure about an idea? Test it.
Applications
Applications/Tools the CCI is responsible for at MSKCC in CMO
This has workflows for BAM generation based on , , small variant calling, micro-satellite instabilty calling, copy number variant calling & structural variant calling for Version 1 of MSK-ACCESS assay
ACCESS_SV
This has the core workflow used for structural variant calling in MSK-ACCESS assay
ADMIE
This is the algorithm used for calling MSI status for sample associated with MSK-ACCESS assay
Nucleo
This is the BAM generation workflow for any assay that deals with Unique Molecular Indexs (UMIs) based on
ACCESS Quality Cotrol (For version 1 of the Assay)
This is the version 2 of the ACCESS QC generation and you can read more about it here
ACCESS data analysis
This repos helps with downstream data analysis of MSK-ACCESS data, you can read more about it here:
Biometrics
Python package to calculate various sample contamination metrics.
sequence_qc
Package for doing various ad-hoc quality control steps from MSK-ACCESS generated FASTQ or BAM files
Krewlyzer
Krewlyzer is a high-performance toolkit for extracting biological features from cell-free DNA (cfDNA) sequencing data. Designed for cancer genomics, liquid biopsy research, and clinical bioinformatics
Kreview
kreview is a production-grade, notebook-first (nbdev) evaluation engine designed for high-throughput cancer liquid biopsy fragmentomics feature analysis. Developed at Memorial Sloan Kettering (MSKCC), it processes cohorts containing tens of thousands of samples using an embedded DuckDB query engine with chunked I/O and automatic retry logic.
gbcms
A high-performance orientation-aware genotype counting system for genomic variants
STRiDE
Microsatellite Instability prediction for MSK-ACCESS cfDNA sequencing.MSK-ACCESS MSI calling tool
Developed by scientists in the CMO Technology Innovation Lab and Department of Pathology, this high-sensitivity assay is offered by the CMO to MSK researchers for profiling circulating tumor DNA derived from blood plasma. The inclusion of matched buffy coat DNA enables the identification and elimination of germline variants and mutations associated with clonal hematopoiesis, a significant confounder of most commercial assays. The assay is available for clinical use in the Molecular Diagnostics Service and for research projects in the Integrated Genomics Operations (IGO). CCI supports the data processing and analysis of research projects utilizing MSK-ACCESS in IGO and leads the ongoing development of the MSK-ACCESS pipeline for all applications. The current version of the pipeline is available here: mskcc/ACCESS-Pipeline: cfDNA Sequencing Pipeline with UMI (github.com), and more details about the assay and analysis are described in this paper below as well as here:
Developed by scientists in the CMO Technology Innovation Lab in collaboration with CCI and Clonal Hematopoiesis (CH) program, Diagnostic Molecular Pathology, Precision Interception, and Prevention Initiative & CCI, this assay utilizes the same barcoding and ultra-deep sequencing technology as MSK-ACCESS to detect CH mutations in white blood cells at high sensitivity. CMO-CH is offered by the CMO to MSK researchers for profiling white blood cell DNA to detect mutations in the most commonly altered CH-associated genes. The assay is run in IGO, and CCI supports the data processing and analysis. You can learn more about it
For both projects, additional analysis packages and development versions of the workflows can be found here:
This wiki will help you to get insights into CMO Cell-Free DNA Informatics Team (CCI) at MSKCC
Welcome aboard!
Welcome to CCI wiki! Here you'll find everything you need to know about CCI.
You can read more about the Center for Molecular Oncology (CMO) here:
General
What to do first?
CMO onboarding
The first thing to do once you have joined is to visit https://mskcc.github.io/on-boarding/ and finish of task necessary for compliance & initiate the process of getting access to various systems.
Cluster Guide
To learn more about the cluster and its resources visit the MIRO board
CCI specific onboarding
Guides on Confluence
You need to be on the internal network to access msk-confluence, you can request access to it on The Spot
Workflows V1
Workflows associated with version 1 of the Assay
BAM Generation & Quality Control
Overview of the BAM Generation and Quality Control workflow
Refer to Bioinformatics Pipeline to Detect CNA's sectionin this paper for details:
Github Location ->
Tool Used:
Github Location ->
Tool Used:
Voyager has all our configurations in the jinja template, it includes all the paths for various files and tools associated with the workflows, all location are on JUNO:
It is a hybrid capture panel designed for Analysis of Circulating cfDNA to Evaluate Somatic Status using the Unique Molecular Index (UMIs) for high sensitivity. MSK-ACCESS is 13% as large, captures 47% of all mutations detected by MSK-IMPACT.
Memorial Sloan Kettering Cancer Center is currently validating MSK-ACCESS v2 (2021). Building on the foundation of our original 2017 assay, Version 2 streamlines processing into a single probe pool while significantly expanding our genomic target territory.
Version Comparison: Advancing from v1 to v2
The updated panel increases our gene coverage and bait territory while simplifying the workflow.
To ensure we capture the most clinically relevant genomic data, several key upgrades have been integrated into v2:
Refined Regions: Updated hotspots and high TMB (Tumor Mutational Burden) regions.
New Tumor Suppressor Genes (TSGs): 5 additional TSGs have been included: CHEK2, ERCC2, PALB2, BAP1, and CDK12.
The MSK-ACCESS v2 panel consists of a carefully curated set of targets designed to provide comprehensive genomic insights. Below is the breakdown of targets, probes, and cumulative territory by category:
Note on Category A Exclusions: The OncoKB targets explicitly exclude Heme and L1 prostate markers (BARD1, BRIP1, CHEK1, RAD51B, RAD51C, RAD51D).
Refer to Bioinformatics Pipeline to Detect CNA's sectionin this paper for details:
Github Location ->
Tool Used:
Github Location ->
Tool Used:
Voyager has all our configurations in the jinja template, it includes all the paths for various files and tools associated with the workflows, all location are on JUNO:
CMO-CH V1
This wiki explains the CMO-CH V1 assay
CMO-CH is offered by the Center for Molecular Oncology (CMO) to MSK researchers for profiling white blood cell DNA to detect mutations in the most commonly altered (CH) associated genes.
596 targets capturing 58% of CH and 90.4% of CH-PD mutations identified in the latest CH dataset from 40K patients
Total size = 1,143 probes (0.14 Mb)
Full gene coverage for TP53, TET2, ASXL1, DNMT3A, PPM1D, CHEK2, ASXL1, ATM, SF3B1, SRSF2, U2AF1, and U2AF2•Additional targets with hotspot positions from IMPACT heme assay
SNP tiling around TP53, CBL, MPL, JAK2, EZH2, TET2, RUNX1, and ATM (+/-10kb) to identify allelic-imbalances
40 fingerprint SNPs that are shared with all other NGS assays (IMPACT, ACCESS, WES etc.) to detect sample mismatches
It is a hybrid capture panel designed for Analysis of Circulating cfDNA to Evaluate Somatic Status using the Unique Molecular Index (UMIs) for high sensitivity. MSK-ACCESS is 13% as large, captures 47% of all mutations detected by MSK-IMPACT.
-> information associated with a sample. This is a current work-in-progress written in Java/Neo4J that is supposed to be the one source of truth for metadata associated with samples. Currently, these responsibilities are managed by Beagle.
LIMS -> Laboratory information management system -- when sequencing is complete, metadata and file information on the sequences is first input into this system.
-> triggers and monitors workflows. There's also a part of it that tracks files and metadata associated with those files which may be splintered out in the future. A request typically comes from LIMS which begins the workflow process.
-> a nicer interface for Beagle (sample workflow tracking) and soon the MDB (make updates to metadata)
-> An HTTP API for Toil which Beagle is reliant on for interacting with workflows.
-> A workflow engine that interfaces well with LSF.
Voyager -> The suite of applications built by the Voyager team, including Ridgeback, Beagle, and Hermes.
-> Distributed high-performance computing that's managed in-house. Most data processing and files are housed here.
-> IBM's Platform Computing (Load Sharing Facility) tool for scheduling workflows on HPC. It has a specific to managing workloads.
-> a YAML-like language for defining workflows.
-> Reactive workflow framework and a programming that eases the writing of data-intensive computational pipelines.
-> Sequencing of the Deoxyribonucleic acid ()
-> Sequencing of the Ribonucleic acid ()
-> Cell-Free DNA assay for patients with solid tumors
CMO-CH -> Assay for profiling clonal hematopoiesis mutations from blood
-> Assay for profiling patients tissue DNA for solid tumors
-> Assay for profiling patients blood DNA for Heme malignancies
-> Allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing
-> Whole Exome Sequencing
-> Whole Genome Sequencing
-> Whole Transcriptome Sequencing
-> one of the companies that provides sequencing machine
-> Type of sequencer from Illumina
-> Type of sequencer from Illumina
-> Type of sequencer from Illumina
-> Type of sequencer from Illumina
-> a company that sells instruments for long-read sequencing based on
-> a company that sells instruments for long-read sequencing based on technology
Requesting Time Off
To request time off, just fill in things at , and also please email your manager of the same.
This will help you to find correct people to connect with from the subgroups.
In all the groups there are multiple amazing individuals involved, but listing just a few
Confluence knowledgebase across collabrators
CMO Project Managers
CCI works closely with CMO project managers to track the progress of cfDNA projects submitted to CMO by MSK researchers. CMO project managers support CCI by facilitating and coordinating project initiation, sample collection, and metadata collection.
CCI works with to make sure that the data generated for the cfDNA assays is of the highest standards.
CCI works closely with CMO Software Engineering (CSE) team within CI to support the development of Voyager & Hermes. The CSE/CAS team supports the development, integration, processing of CCI’s workflows implemented in Voyager.
CCI collaborates with scientists in various CMO groups to improve the existing workflows and to analyze the data in a consistent manner.
CCI works closely with the ClinBx group in the Molecular Diagnostics Service to deploy core workflows and maintain consistency among analyses performed on research and clinical cfDNA samples. We work together to improve the workflows in sync and to learn from one another about how to identify artifacts and interpret the data. CCI is also working with the ClinBx Software group to port an instance of the mPATH system used by the ClinBx team to sign-out cases for research.
CCI works with HPC to request resources and support w.r.t JUNO cluster and virtual machines that enable various aspects of our goals.