Cleveland Clinic Research Logo
Cleveland Clinic Research Logo
  • About
  • Science
    • Laboratories
    • Office of Research Development
    • Clinical Research
      Participating in Research
    • Departments
      Biomedical Engineering Cancer Sciences Computational Life Sciences Florida Research & Innovation Center Genomic Sciences & Systems Biology
      Heart, Blood & Kidney Research Inflammation & Immunity Microbial Sciences in Health Neurosciences Ophthalmic Research Quantitative Health Sciences
    • Centers & Programs
      Advanced Musculoskeletal Imaging Angiogenesis Center Cardiovascular Diagnostics & Prevention Cell Therapy & Immunoengineering Program Consortium for Pain Genitourinary Malignancies Research Genome Center Global Center for Immunotherapy & Precision Immuno-Oncology Microbiome & Human Health
      Musculoskeletal ResearchNeuro-Oncology ProgramNorthern Ohio Alcohol Center Pathogen & Human Health Research Populations Health Research Quantitative Metabolic Research Therapeutics Discovery Translational and Correlative Sciences Service Tumor Pathophysiology & Evolution Program Women's Cancers Program
  • Core Services
    • Ohio
      3D Printing Bioimage AnalysisBioRobotics & Mechanical Testing Cell Culture Cleveland Clinic BioRepository Computational Oncology Platform Discovery Lab Electron Microscopy Electronics Engineering
      Flow CytometryGenomic Medicine Institute Biorepository Genomics Glassware Histology Hybridoma Immunohistochemistry Immunomonitoring Lab Instrument Refurbishing & Repair Laboratory Diagnostic
      Lerner Research Institute BioRepository Light MicroscopyMechanical Prototyping Microbial Culturing & Engineering Microbial Sequencing & Analytics Media Preparation Molecular Biotechnology Nitinol Polymer Proteomics & Metabolomics SomaScan & Biomarker
    • Florida
      Flow Cytometry
      Imaging
  • Education & Training
    • Graduate Programs Molecular Medicine PhD Program Postdoctoral Program
      Global Research Education Research Intensive Summer Experience (RISE) Undergraduate & High School Programs
  • News
  • Careers
    • Faculty Positions Research Associate & Project Staff Postdoctoral Positions Technical & Administrative Engagement
  • Donate
  • Contact
  • About
  • Science
    • Scientific Programs
    • Laboratories
    • Office of Research Development
    • Clinical Research
      Participating in Research
    • Departments
      Biomedical Engineering Cancer Sciences Computational Life Sciences Florida Research & Innovation Center Genomic Sciences & Systems Biology
      Heart, Blood & Kidney Research Inflammation & Immunity Microbial Sciences in Health Neurosciences Ophthalmic Research Quantitative Health Sciences
    • Centers & Programs
      Advanced Musculoskeletal Imaging Angiogenesis Center Cardiovascular Diagnostics & Prevention Cell Therapy & Immunoengineering Program Consortium for Pain Genitourinary Malignancies Research Genome Center Global Center for Immunotherapy & Precision Immuno-Oncology Microbiome & Human Health
      Musculoskeletal Research Neuro-Oncology ProgramNorthern Ohio Alcohol Center Pathogen & Human Health Research Populations Health Research Quantitative Metabolic Research Therapeutics Discovery Translational and Correlative Sciences Service Tumor Pathophysiology & Evolution Program Women's Cancers Program
  • Core Services
    • All Cores
    • Ohio
      3D Printing Bioimage Analysis BioRobotics & Mechanical Testing Cell Culture Cleveland Clinic BioRepository Computational Oncology Platform Discovery Lab Electron Microscopy Electronics Engineering
      Flow CytometryGenomic Medicine Institute BiorepositoryGenomics Glassware Histology Hybridoma Immunohistochemistry Immunomonitoring Lab Instrument Refurbishing & Repair Laboratory Diagnostic
      Lerner Research Institute BioRepository Light MicroscopyMechanical Prototyping Microbial Culturing & Engineering Microbial Sequencing & Analytics Media Preparation Molecular Biotechnology Nitinol Polymer Proteomics SomaScan & Biomarker
    • Florida
      Flow Cytometry
      Imaging
  • Education & Training
    • Research Education & Training Center
    • Graduate Programs Molecular Medicine PhD Program Postdoctoral Program
      Global Research Education Research Intensive Summer Experience (RISE) Undergraduate & High School Programs
  • News
  • Careers
    • Faculty Positions Research Associate & Project Staff Postdoctoral Positions Technical & AdministrativeEngagement
  • Donate
  • Contact
  • Search

Cancer Sciences

Computational Oncology Platform

❮Translational & Correlative Sciences Service Computational Oncology Platform
  • Computational Oncology Platform
  • About Us
  • Our Team
  • Analyses & Technology
  • Services
  • Precision Oncology Program
    Overview Magic Data Warehouse cBioPortal Pythia Our Team Our Partners Highlights Contact

About Us

About Us

The Cancer Sciences Computational Oncology Platform group uses various core technologies for processing and analyzing data based on high-throughput next-generation sequencing from the raw data to ready-to-use results. Through partnerships with scientists, clinicians and industry, Cancer Sciences accelerates the development and approval of novel immunotherapy treatments and applications. Our partners benefit from early access to scientific advances, expansive clinical-genomic datasets and expert advice and analysis.

Our scientific & technology areas include:

  • Mutation Calling
  • Neoantigen Analysis
  • Copy-Number Analysis
  • Tumor Clonality Analysis
  • Gene Expression Analysis
  • Immune Infiltration and Immune Activity Analysis
  • Differentially Expressed Gene and Pathway Analysis
  • Single Cell RNA Sequencing

Learn more about our scientific & technology areas.

Services

The Computational Oncology Platform offers bioinformatics & computational biology support as a core service to investigators. Learn more about our services. To view prices and request services, visit our iLab page.

Contact Us

Interested in learning more? Contact us today at [email protected].

Our Team

Our Team

Vladimir Makarov Headshot

Vladimir Makarov, MD, MS

Staff Scientist

Read Bio

Dr. Vlad Makarov is the Scientific Director of the Computational Immunology Platform (CIP), a subdivision of CITI’s Precision Oncology Program. Since joining the Cleveland Clinic in 2020, Dr. Makarov and his team have launched bioinformatic computational fee-for-project shared resource services for analysis of high throughput DNA/RNA/single cell RNA-sequencing data output. Dr. Makarov earned his MD in St. Petersburg, Russia. His background in population genetics, cancer genomics, immuno-oncology, and software development has poised him to lead a team of computational genomic data scientists. He has been involved in various genomic consortium projects and has led the development and optimization of variant-calling pipelines for whole genome (exome and transcriptome) analyses. He has also co-authored three published genome analysis tools: AnnTools, RDXplorer, and NAseek. 

Tyler Alban Headshot

Tyler Alban, PhD

Assistant Staff
Director, Translational and Correlative Sciences Service
[email protected]
Lab Profile

Prerana  Bangalore Parthasarathy Headshot

Prerana Bangalore Parthasarathy, MS

Data Scientist
[email protected]

Ivan Juric Headshot

Ivan Juric, PhD

Data Scientist
[email protected]

Benjamin Kovacic Headshot

Benjamin Kovacic

Research Data Scientist
[email protected]

John Metzcar Headshot

John Metzcar, PhD

Project Staff
[email protected]

Phong Nguyen Headshot

Phong Nguyen

Software Developer
[email protected]

Ying Ni Headshot

Ying Ni, PhD

Assistant Staff
[email protected]
Lab Profile

Shikha Parsai Headshot

Shikha Parsai, MS

Research Data Scientist
[email protected]

Anurag Saxena Headshot

Anurag Saxena, MS

Software Developer
[email protected]

Salendra Singh Headshot

Salendra Singh, MS

Software Developer
[email protected]

Read Bio

Salendra Singh is a software developer and data scientist on CITI’s computational immunology team. He obtained his Master of Science in Biochemistry and Molecular Biology from Georgetown University and has been working in the field of computational biology for over 11 years. He most recently served as a Senior Bioinformatics Scientist for the Case Comprehensive Cancer Center. At CITI, Salendra plays a leading role in the development and management of departmental web resources while also continuing research efforts in immunotherapy, genomics, oncology, and biomedical sciences. His research focuses on developing systems biology tools and contributing to integrative genomics, digital pathology, radio genomics, and single cell/nuclei and spatial genomics. Find Salendra on LinkedIn, Google Scholar, and Twitter.

Analyses & Technology

Analyses & Technology

The Cancer Sciences Computational Oncology Platform utilizes large-scale technologies and high-throughput profiling for collaborative projects. These analyses are also offered as core services to support research community. Through partnerships with scientists, clinicians, and industry collaborators, Cancer Sciences computational faculty focus state-of-the-art computational biology and NGS-based discovery. Our partners benefit from early access to scientific advances, expansive clinical-genomic datasets, and expert advice and analysis.

The web-based immunogenomics analysis platform developed by the Department of Cancer Sciences, IOExplorer, allows collaborators to easily visualize their results, explore analytic variations, and compare their cohorts with published datasets.

The Cancer Sciences computational division, in collaboration with Cleveland Clinic’s Pathology & Laboratory Medicine Institute (PLMI) and the Taussig Cancer Institute (TCI), operates the Precision Oncology Program. As part of this program, we support CCF cBioPortal, which provides easy access to genomic data from CCF patients in a scalable, mineable, easy to use portal environment linked with treatment recommendations.

Below are some of our primary pipelines and services.


Having been a key early member of the TCGA analysis team, our group has extensive experience with the analysis of cancer genomes. Our mutation calling pipeline processes whole-exome, whole genome, and targeted gene sequencing panels to identify somatic and germline mutations. Illustrative example analyses are shown below.

Mutation calling pipelines are fully automated following the best practices recommended by the National Institutes of Health.

Shown below is an example of genomic output for a recent clinical trial.

Representative integrated mutational data from melanoma. Data types are labeled above. Riaz et al, Tumor and Microenvironment Evolution during Immunotherapy with Nivolumab, Cell, 2017 Nov 2;171(4):934-949

Neoantigens are mutated peptides that form the basis of how immune cells recognize cancer cells. Our research group was the first to show that tumor mutations and neoantigen burden drive immunotherapy treatment benefit in patients, a finding that is a cornerstone of the understanding of immunotherapy’s mechanism of action. We can use computational methods to predict neoantigens using various methods and incorporate the predictions into integrated genomic analyses. By using concurrent HLA genotyping, we can also characterize HLA divergence (HED) and immunopeptidome content and use this information to elucidate cancer cell epitope presentation.

Neoantigen formation. Different types of mutations can form neoantigens in cancer cells that are recognized by the immune system.

HLA evolutionary divergence predicts for immunotherapy efficacy in melanoma patients treated with immune checkpoint blockade.

Allele-specific copy-number analysis is performed using DNA sequencing data. The fraction of the copy-number-altered genome is defined as the fraction of the genome with either non-diploid copy-number or evidence of loss of heterozygosity. Sample purity and ploidy are also estimated for use in downstream analysis. This is helpful to characterize copy number alterations in normal and disease states.

Example of copy number analysis identifying genomic changes in cancer. Ganly et al. Cancer Cell 2018.

Cancer cell fraction (CCF) represents an estimate of the fraction of cancer cells carrying a given mutation. Analysis of CCF can help identify cell subclones that independently develop over the lifetime of a tumor and estimate their relative fitness and susceptibility to immune targeting. For each mutation, we calculate the CCF based on variant allele frequency, copy number, and sample purity estimated in previous steps. Furthermore, we classify single nucleotide variants (SNVs) into clonal and sub-clonal variants depending on the confidence interval (CI) of the CCF estimation. SNVs for which the lower bound of the CI exceeded 95% are considered clonal mutations, others sub-clonal.

Bulk RNA-seq experiments measure the average expression of each gene across an entire transcriptome. Reads obtained by the experiment are aligned to the latest build of the human or model organism genome. Raw gene-level count values are normalized by sample specific size factor and FPKM (Fragments Per Kilobase Million) values are reported. The normalized values are used to find significantly different expression between specified groups of samples. Canonical pathway analysis of differentially expressed genes is then performed by pathway analysis software. Tools like GSEA and network analyses are used to define the biological pathways of altered transcriptional programs. Additionally, more advanced analyses can be performed from bulk RNA experiments, such as gene fusion expression and TCR clonality. We have recently added the Immunarch package to further analyze T-cell receptor (TCR) and B-cell receptor (BCR) repertoires.

Example of unsupervised clustering analysis of gene expression data from thyroid cancers.

In mixed cell populations, bulk RNA-seq experiments lack the resolution required to identify the cell types responsible for altering gene expression between groups. We deploy high resolution single cell RNA-seq (scRNA-seq) to gain a better understanding of how individual cell populations are altered during therapeutic response. In scRNA-seq experiments, reads are tracked to individual cells and used to construct expression profiles, enabling accurate detection of different cell types. ScRNA-seq using 10X genomics 3’ and 5’ protocols with VDJ and Feature library preparations are used in the laboratory with both tumor and PBMC samples. Downstream processing includes the standard 10X genomics Cell Ranger 6.0.0 software pipeline, followed by additional quality control checks, data normalization, and batch correction using the R package Seurat v4.0. These methods have been validated on multiple immunotherapy studies including on renal cell carcinoma and breast cancer. With verified cell signatures and reference datasets provided from Seurat, we can confidently identify immune cell types to identify changes under immunotherapy, including therapeutic resistance.

Mutational signatures are patterns of nucleotide alterations in tumor genomes that are characteristic of various mutational processes, including carcinogenic insult, aging, and DNA repair defects. Our mutational signatures computational pipeline utilizes mathematical methods to estimate, from NGS samples, the contribution of various known mutational signatures. The pipeline includes mutation calling, tri-nucleotide context matrix generation and normalization, negative matrix factorization, non-negative least squares regression, prediction of mutational signatures, and transcriptional strand-based mutational signature analysis (reference).



Immunogenomics

To support research and clinical operations, the department collaborates with multiple partners within and beyond Cleveland Clinic to develop computational systems that facilitate the efficient utilization of our extensive data collections. The basic components of our core immunogenomics data systems are: multiple sources of clinical and genomic data, a central data warehouse where the data are integrated (MaGiC), and applications to access and use the data, such as cBioPortal and IOExplorer.

The Department of Cancer Sciences works closely with collaborators from multiple Cleveland Clinic divisions, including the Pathology and Laboratory Medicine Institute (PLMI), the Taussig Cancer Institute (TCI), the Enterprise Analytics (EA) division, and departments within the Cleveland Clinic Research. Through this joint effort, we integrate disparate data sources into a unified precision oncology data warehouse called the Molecular and Genomics Integrated at Cleveland Clinic (MaGiC) database. Data sources include databases and files overseen by PLMI, TCI, and EA, which are hosted on various data systems such as Teradata, MS SQL Server, Oracle, and flat files. Key types of data include sample and diagnostic, such as cancer types, staging, and procedure dates; molecular, such as raw sequencing results and mutation calls from various panels and providers such as Tempus and Caris; and curated treatment and outcome data. After integration, anonymized data extracts are provisioned for exploratory analysis to client applications such as cBioPortal and IOExplorer, or for advanced analysis directly to IRB authorized clinicians or researchers. MaGiC is intended and designed to facilitate future expansion to any disease type.

cBioPortal is an open source, interactive graphical cancer genomics web app developed in association with Memorial Sloan Kettering Cancer Center and used by major cancer centers. Cancer Sciences and our partners are providing cBioPortal to the Cleveland Clinic community in support of both clinical and research uses. The cBioPortal includes, in addition to public datasets such as TCGA, genomics data from major gene panels used at CCF, such as Tempus and Caris, and will expand to include all CCF panels and other genomics data. In addition to genomics data, cBioPortal contains associated clinical data such as diagnoses, treatments, and outcomes. Access is restricted according to IRB and clinical authorization. For additional details about cBioPortal functionality, please see https://www.cbioportal.org/tutorials.

MAGIC and cBIOPortal are key elements that support precision medicine in research and clinical activities. They are integrated with the CCF cancer center genomics tumor board. The system provides genomics reporting for ordering clinicians and enables scalable and minable genomics research for clinical trials.

The pace of cancer immunotherapy research is accelerating, increasing the volume of data available for developing novel therapies, discovering and refining biomarkers for more precise targeting of existing therapies, and making other advances. But utilizing existing data to formulate new hypotheses, address clinical questions, etc. continues to require extensive bioinformatics expertise and time, which is an impediment to the efficiency, scope, and pace of research.

While powerful, cBioPortal (above) is not designed to specifically investigate immuno-oncology datasets and thus does not provide certain functionality key to this field of research, such as analysis of HLA types and diversity. To fill this gap and make exploratory IO analyses more accessible, rapid, and powerful, we have developed an IO-specific interactive graphical web application, called IOExplorer.

Key features of IOExplorer include:

  • An interactive graphical user interface with convenience features, such as point-and-drag data selection and data type tagging.
  • Meta-analysis ready datasets and features.
  • Pre-/on-treatment model for samples.
  • Multiple analysis modules: distribution, correlation, mutation, expression, volcano (beta), HLA, survival.
  • Analyses can be saved, resumed, and shared.
  • Interactive tutorials and contextual help.
  • Collaborative development with researchers and clinicians.


Datasets & Standardization

The first obstacles to exploratory and meta-analysis are simple in concept but complicated to implement: acquiring and standardizing available datasets. Whereas acquisition is a bureaucratic exercise, meaningful analyses that pool or compare data between or across studies require that the data be standardized. Effective standardization demands extensive technical expertise and computing resources. IOExplorer includes key published IO datasets (see below) that we have acquired and standardized by reprocessing the raw data through our established pipelines (where permitted by data use agreements ) (see above). In addition, IOExplorer includes features to facilitate analyses of multiple datasets, including capabilities to:

  • Pool and filter data according to any user-specified criteria. Filtering criteria are specified via analysis modules, either by predefined groupings such as quartiles, or visually by click-and-drag selection.
  • Correct expression data in real-time across studies or batches.
  • Use built-in standardized clinical attributes (e.g., response to therapy).
  • Analyze cohorts created from pooled data or compare between cohorts.


The following datasets are currently included in IOExplorer (as of December 2021) and we frequently update as new studies are published, processed, and standardized. In general, data use agreements restrict use to analyses within IOExplorer, not redistribution of original data.

  • Snyder et al., NEJM 2014
  • Riaz et al., Cell 2015
  • Rizvi et al., Science 2015
  • Van Allen et al., Science 2015
  • Gao et al., Cell 2016
  • Hugo et al., Cell 2016
  • Ott et al., Nature 2017
  • Roh et al., Sci. Transl. Med. 2017
  • Auslander et al., Nat. Med. 2018
  • Cristescu et al., Science 2018
  • McDermott et al., Nat. Med. 2018
  • Miao et al., Nat. Gen. 2018
  • Miao et al., Science 2018
  • Kim et al., Nat. Med. 2018
  • Cloughesy et al., Nat. Med. 2019
  • Gide et al., Cancer Cell 2019
  • Liu et al., Nat. Med. 2019
  • Zhao et al., Nat. Med. 2019
  • Valero et al., Nat. Gen. 2021

Analysis Modules

Distribution

Exploratory data analysis is an essential first step of rigorous statistical analysis. It involves an examination of basic features of a dataset to understand the characteristics of its data points (e.g., samples) and its relationship to other datasets, helping to define the applicability of various statistical models and tests, etc. IOExplorer provides two analysis modules dedicated to examining the characteristics of individual variables (the Distribution module) or pairs of variables (the Correlation module). As with all modules, analyses may be confined to a single dataset or applied to cohorts consisting of combinations of 1-to-N datasets and 0-to-N filters. In addition to these to purpose-built modules, all IOExplorer modules provide functionality for exploratory for specific attributes relevant to immuno-oncology.

Mutation

Mutation analysis is a cornerstone of immuno-oncology. Loss-of-function mutations in tumor suppressor genes and gain-of-function mutations in proto-oncogenes are associated with specific cancer histologies and may be predictive of specific treatment outcomes and thus of high clinical relevance. Furthermore, the overall mutation load of a tumor (TMB) is generally predictive of the effectiveness of immunotherapy: high TMB is associated with improved immunotherapy outcomes in many patients, possibly via generating effective neoantigen targets for immune system activity. IOExplorer permits the user to select one or more predefined gene sets relevant to immunotherapy (e.g., CD8+ T-cells), or to enter any gene(s) of interest. Single and short nucleotide variants are displayed on oncoprint displays with customizable axis, ordering and scaling, permitting the user to visualize associations between gene sets and any clinical or genomic variable(s) of interest. Copy-number variation analysis is under development.

Expression

Analysis of altered tumor expression patterns is another cornerstone of immunogenomics and goes hand in glove with mutation analysis. Characteristic histologies and predicted treatment outcomes may be associated with under-expression of tumor suppressor genes or over-expression of oncogenes, which may be associated with detected mutations or mutations to regulatory regions not included in gene sequencing panels or by more complex genomic alterations.

RNA-seq results are notoriously sensitive to batch effects, making it particularly challenging to implement general comparisons between or pooling among batches or studies. IOExplorer implements two general approaches to expression data normalization. First, all studies include user-selectable options for both raw and customary per-study normalization (VST, TPM, CPM). Furthermore, IOExplorer provides capabilities for real-time cross-study batch correction, which is currently limited to user selected pairs of studies (Volcano module), but with more generalized real-time normalization under development.

As with mutations, IOExplorer provides the user to interactively explore patterns of altered expression in tumor samples using predefined or custom gene sets, displayed as heatmaps, and to associate altered expression patterns with user-configurable clinical and molecular variables.

HLA Divergence

If mutation and expression go hand-in-glove, then HLA is the hand. HLA-I is key in determining which (if any) antigens are displayed on the surface of tumor cells for immune recognition. Until recently, the capacity of cells to display neoantigens was typically analyzed by simply examining whether HLA-A, HLA-B, and HLA-C were homozygous or heterozygous. Such an approach ignores the great variability between different allele pairs, some of which are nearly identical and others of which are highly diverged. Members of our team developed an improved approach to characterizing HLA diversity based on the physiochemical differences between heterozygous alleles, called HLA evolutionary divergence (HED) (reference). IOExplorer provides multiple interactive displays for both HLA-I heterozygosity and HED of cohorts, making it an extremely powerful tool for HLA analyses. We have created and maintain the HED R package which has been updated to the latest Release 3.46 of IEDB (Oct 2021)

Survival

Finally, IOExplorer endpoint analysis is currently provided as Kaplan-Meier overall survival plots with configurable Wilcoxon weighting, optional confidence interval display, and statistics.

Services

Services

Bioinformatics & Computational Biology Services

The Cancer Sciences Computational Oncology Platform offers bioinformatics support as a core service. We use industry standard tools for data processing and analysis. Having been original core members of the Cancer Genome Atlas (TCGA) team, our pipelines are extensively benchmarked and based on best practices. Users are encouraged to include the description to their papers to support the methods sections within their manuscripts. All deliverables are in industry standard formats such as BAM, VCF or MAF. Quality control (QC) metrics are also provided, allowing investigators to gauge the quality of sequencing in a unified manner. We encourage researches to schedule a meeting with us prior to data processing/analysis.

The Cancer Sciences computational team currently offers services for the following data types:

  • STAR alignment to reference genome, Pre/Post alignment QC collection, raw and normalized read counts over genes.
  • Differential gene expression as specified by users. Users provide metadata files according to provided templates.
  • Gene set enrichment analysis (GSEA) and pathway analysis and xCell deconvolution for 64 immune cell types.

  • Alignment to reference and Pre/Post alignment QC collection.
  • DNA profiling (genetic fingerprinting). Optional, add to post-alignment QC.
  • Post-alignment WES/WGS data processing.
  • Somatic or germline variant calling and annotation (SNV/INDELS).
  • Allele Specific Copy Number calling.
  • Microsatellite instability (MSI) status.

  • Standard Cell Ranger pipeline, Loupe objects per sample and per project, Pre/Post alignment QC metrics collection.
  • Post-alignment Single Cell RNASeq analysis. Includes standard Seurat pipeline, generating Integrated Seurat Object and automated cell typing.

  • T-cell/Antibody Immune Repertoires Analysis. Data normalization, CDR3 identification, clonal dynamics, and repertoire features such as diversity and Shannon metrics. We will need to discuss metadata, comparison, and other parameters prior to analysis. For TCR Sequencing, please see the Discovery Lab.
  • Uploading genomics data to dbGAP or other data repositories for publication.
  • Additional custom analysis may be possible for additional fees, please contact us for availability. Services beyond the basic analyses provided above will be available based on feasibility and availability.

Researchers interested in continuing working with us through more extensive collaborative efforts after basic data processing/analyses has been completed may contact us today.

Future Services

Work in Progress coming soon
We are currently in the process of adding the new pipelines.

ChIP-Seq
ATACseq

Precision Oncology Program

Precision Oncology Program

Overview

Precision Oncology seeks to match the best therapies with each individual patient. At the Cleveland Clinic, our Precision Oncology Program helps cancer patients access advanced testing to analyze their tumor on a molecular level.

What is Precision Oncology at Cleveland Clinic?

Precision Oncology seeks to match the best therapies with each individual patient. At the Cleveland Clinic, our Precision Oncology Program helps cancer patients access advanced testing to analyze their tumor on a molecular level. The program focuses on identifying the unique features of each patient's tumor to provide the most effective, targeted therapies based on personalized evaluations and treatment plans. The past few decades have seen the development of hundreds of targeted agents and Immunotherapeutics.

The mission of our program is to empower clinicians and investigators across the Cleveland Clinic Health System to integrate precision therapeutics into oncology practice by providing decision support, genomics infrastructure, and expertise. Our goal is to provide personalized treatment options to give every patient the best possible care.

About Our Program

The Precision Oncology Program is a multi-institute initiative between Cleveland Clinic Research, the Cancer Institute, and the Diagnostic Institute. Oncologists in the Cancer Institute as well as other physicians offer precision oncology options to our patients. The Diagnostic Institute uses state-of-the-art sequencing to generate data. The Cleveland Clinic Research program runs state-of-the-art platforms to power a comprehensive set of end-to-end software solutions that streamline decision support and research. 

Research Computing, CC Research, Information Technology Department, Diagnostic Institute, Cancer Institute, Data and Analytics 

Cancer patients at the Cleveland Clinic receive standard-of-care genomics testing and trial-based assays with the reporting of clinically actionable genomic alterations and associated therapeutic targets.

Our Efforts

Centralized Data for Better Care
We bring together genetic and medical information in one place so that doctors can see a complete picture of each patient's situation. This helps us create personalized treatment plans for every patient.

Matching Patients with Clinical Trials
Our program helps match patients with clinical trials that are right for their specific genetic profile. This means patients have access to new and emerging treatments that may not be available otherwise.

Easy Access for Researchers
We have created an easy-to-use system for our researchers to access both our own data and public data. This helps our researchers make important discoveries that can improve patient care. We are the first to routinely use whole exome and RNAseq for tumor sequencing.

Tools to Help Doctors Make the Best Decisions
We provide doctors with tools to understand the genetic details of each patient's cancer, helping them make the best decisions about treatment. This means that every patient's treatment plan is based on the latest and most relevant information.



Magic Data Warehouse

Empowering Research and Clinical Applications with Integrated Genomic Data

The MaGIC (Medical and Genomics Data Integrated at Cleveland Clinic) is a comprehensive data warehouse within the Cleveland Clinic enterprise system, serving as the foundational platform for numerous research and clinical applications. Through associated front-end interfaces, such as cBioPortal and Pythia, data is made accessible to researchers and clinicians in a secure, de-identified, or controlled manner.

The MaGIC system is made up of automated processes that update every day. These processes collect data from clinical results and genetic profiles and then organize it into a standard format. This data is stored in a special database, which follows strict rules to keep the information confidential.

MaGIC holds both clinical results and processed sequencing data, creating a valuable resource for research and patient care. Through easy-to-use tools like cBioPortal and Pythia, MaGIC provides secure access to data for doctors and researchers.

All data is kept safe behind the Cleveland Clinic firewall, with daily backups and strict access controls. Currently, MaGIC contains approximately 27,000 molecular profiles from multiple sequencing platforms, with an ongoing addition of 250 patients per month.

If you have inquiries regarding the MaGIC Data Warehouse, please contact [email protected] .

cBioPortal

Cleveland Clinic cBioPortal Instance

cBioPortal is the main tool we use to present and visualize genomics data. It is an open-source platform that allows users to explore, interact with, analyze, and download large-scale cancer genomics datasets.

The Cleveland Clinic version of cBioPortal includes de-identified data from Cleveland Clinic cancer patients, as well as public datasets. Access to Cleveland Clinic-specific data is restricted to authorized users who apply via [email protected].

For trialists, we provide a service that lets them create a private study, accessible only to specified users. This way, they can use cBioPortal's features to analyze your data and compare it with other studies.

For additional details about cBioPortal functionality, please see https://www.cbioportal.org/tutorials.

Questions? Email [email protected]

Pythia

Pythia is a clinical decision support tool that uses real-time genomic testing results from the MaGIC Data Warehouse, along with clinical trial information and patient data, to match patients with suitable clinical trials. By combining multiple data sources into one easy-to-use platform, Pythia helps oncologists make informed and timely decisions about trial recommendations, ultimately improving cancer treatment and patient care.

If you have inquiries regarding the Pythia platform, please contact [email protected].

Our Team


Our Partners
Wen Wee Ma, MD

Wen Wee Ma, MD

Director, Novel Therapeutics Clinic
Cancer Institute Lead

Sarah Johnson

Sarah Johnson

Program Manager of Precision Oncology
Cancer Institute Lead

Scott Robertson, MD, PhD

Scott Robertson, MD, PhD

Medical Director of Image Analytics
Diagnostic Institute Lead

Elizabeth Azzato, MD

Elizabeth Azzato, MD

Section Head of Molecular Pathology and Cytogenomics
Diagnostic Institute Lead

Jacob Miller, MD

Jacob Miller, MD

Associate Director of Precision Oncology
Cancer Institute Lead



Highlights

The Precision Oncology Program carries out the development and implementation of comprehensive end-to-end data management and software solutions designed to streamline decision support for patient care and scientific research. Our achievements have not only been recognized by investigators seeking support for their research studies and clinical providers utilizing the tools to improve their work productivity, but also throughout the CCF enterprise for their innovation in digital solutions. Our achievements are represented in a variety of ways:

  • Catalyst Award (2023-2024): "MaGIC clinical trial matching platform for patients with cancer"
  • Mandel Accelerator Award (2024-2025): "Transformative AI-Driven Clinical Trial and Care Path Automation"
  • Oral presentations:
    • CCF Analytics & AI Summit (AAIS):
      • "MaGIC: Facilitating Precision Oncology through Multi-Institutional Collaboration and Data Dissemination"
      • "Pythia – A cancer clinical trial matching and cohort building platform"
    • IBM Discovery Accelerator Workshop:
      • "Pythia – A cancer clinical trial matching and cohort building platform"
    • Cancer Center Grand Rounds: Clinical Translational Partnership Lecture Series:
      • "Leveraging Clinical Data to Power Precision Oncology: Building Smart Systems to Support Our Caregivers and Researchers"
    • Center for Clinical Research Roundtable
      • "Bridging Data and Clinical Decision Support: Cancer Genomics to Precision Oncology"
  • Featured research studies:
    • Glioblastoma clinical trials
    • Clonal hematopoiesis of indeterminate potential (CHIP) study
    • Ovarian cancer study
    • GU NGS and pathomics


Contact

If you want to use our data, please contact our program manager Sarah Johnson at [email protected].

Please click here to download the Genomic Data Request Form.

If you have questions or want to set up a private study for cBioPortal, please contact [email protected].

If you have inquiries regarding the MaGIC Data Warehouse or Pythia platform, please contact [email protected] or [email protected], respectively.

About Cleveland Clinic Research

About Us Careers Contact Us Donate People Directory

Science

Clinical & Translational Research Core Services Departments, Centers & Programs Laboratories Research News

Education & Training

Graduate Programs Global Research Education Molecular Medicine PhD Program Postdoctoral Program RISE Program Undergraduate & High School Programs

Site Information & Policies

Privacy Policy Search Site Site Map Social Media Policy

9500 Euclid Avenue, Cleveland, Ohio 44195 | © 2026 Cleveland Clinic Research