Machine Learning Biomarker Discovery and Research Validation Service

A statistically significant feature list is not yet a biomarker panel. Candidates must remain stable across samples and analysis choices, add value beyond a simple baseline, make biological sense, and be measurable in a follow-up study.

CD Genomics coordinates wet-lab omics experiments, rigorous bioinformatics, machine-learning-assisted prioritization, and research validation. Start with biospecimens, existing datasets, or both, and build evidence from discovery through an interpretable candidate panel and a practical next-study plan.

  • Sample-to-Insight, Data-to-Insight, or Hybrid Study
  • Genomics, transcriptomics, epigenomics, single-cell, and multi-omics options
  • Nested assessment, feature stability, confounder review, and calibration
  • Independent-cohort and targeted experimental validation pathways

Integrated biomarker discovery solution connecting biospecimens, omics experiments, machine learning, and research validation

Decision-Oriented Deliverables

  • Quality-controlled omics data and analysis-ready feature matrices
  • Baseline and machine-learning model comparison
  • Feature-stability and confounder assessment
  • Interpretable candidate panel with evidence tiers
  • Independent assessment or targeted validation results
  • Limitations and recommended next experiments
Table of Contents

    Study-design map linking research endpoint, cohort, omics technology, machine learning, and validation evidence

    Design backward from the research decision: define the evidence a candidate must meet before selecting the assay or model.

    Why Biomarker Studies Need More Than Feature Ranking

    High-dimensional omics data can produce attractive signatures that do not survive a new cohort. An integrated study is needed to separate repeatable biological signal from batch effects, cohort imbalance, overfitting, and chance correlations.

    Study design determines what a model can legitimately answer. Wet-lab quality determines whether the intended signal reaches the data. Bioinformatics turns raw measurements into comparable features. Machine learning can then help prioritize combinations and quantify their robustness—but it cannot repair an unsuitable cohort or replace independent evidence.

    We therefore begin with the intended decision: early discovery, treatment-response research, molecular stratification, longitudinal risk exploration, or validation of an existing hypothesis. That decision guides group definitions, molecular layers, sample allocation, locked evaluation strategy, panel size, and follow-up assay.

    Questions the study should answer

    • Is the signal stronger than a simple baseline?
    • Is it stable across resampling and preprocessing?
    • Could technical or clinical variables explain it?
    • Can the shortlist be measured in new samples?
    • What evidence is still missing?

    Concrete Research Scenarios

    The following examples show how experiments, analysis, and validation can be connected. Cohort sizes and methods are illustrative and are refined after feasibility review.

    1

    Blood-Based Biomarker Discovery for Earlier Disease Research

    Client question: Can a compact blood-based molecular panel distinguish the target condition from both healthy and clinically relevant benign controls?

    Illustrative design: Use serum or plasma from a discovery cohort containing target cases, benign-disease controls, and healthy controls. Generate discovery proteomic, metabolomic, cfDNA, or transcriptomic profiles; retain a locked subset or independent cohort for assessment. Relevant covariates may include age, sex, collection site, fasting status, hemolysis, storage time, and processing batch.

    Machine-learning role: Compare a conventional baseline with regularized and ensemble approaches; perform feature selection inside resampling; report stability, ROC and precision-recall performance, calibration where relevant, and the incremental value of a compact panel. Shortlisted proteins or molecular features can move to targeted measurement in retained samples.

    What the client receives: Evidence-ranked candidates, a compact panel proposal, model and stability report, confounder analysis, targeted-validation plan or results, and a go/refine/stop recommendation.

    Research basis: Ney et al. used 539 serum samples, diagnosis-specific base learners, feature elimination, and held-out assessment to develop a pancreatic cancer protein panel. Xing et al. demonstrated a staged discovery-verification-validation proteomics workflow across 1,002 participants. These studies support the workflow logic, not guaranteed performance in another cohort.

    2

    Pre-Treatment Biomarkers of Research Response

    Client question: Which baseline molecular features distinguish responders from non-responders, and does multi-omics information improve on clinical variables alone?

    Illustrative design: Collect pre-treatment biopsies and a clearly defined response endpoint. Depending on mechanism and sample availability, combine variant profiling, expression, immune-state, epigenetic, or pathology-derived features. Reserve evaluation data by patient, site, or study cohort so that related samples cannot cross partitions.

    Machine-learning role: Compare clinical-only, single-omics, and integrated models under the same evaluation plan; review class imbalance, treatment arm, tissue composition, and site effects; identify stable features that add measurable value beyond the baseline.

    Validation route: Test the locked signature in an external dataset or follow-up cohort, then use targeted sequencing or expression measurement to verify a manageable candidate set.

    Research basis: Sammut et al. integrated clinical, digital pathology, genomic, and transcriptomic features from pre-treatment breast tumor biopsies and evaluated the response model in an external cohort, illustrating why the entire tumor ecosystem and an independent dataset matter.

    3

    Single-Cell Discovery of Cell-State Biomarkers

    Client question: Is the apparent bulk-tissue signal driven by a specific immune, stromal, or disease-associated cell population?

    Illustrative design: Profile carefully balanced tissue or peripheral-blood samples at single-cell resolution. Define the donor—not the cell—as the independent sample unit. Discover cell populations, state programs, and abundance shifts associated with the endpoint, while controlling donor and processing effects.

    Machine-learning role: Prioritize reproducible cell-state features, compare donor-level signatures, and avoid inflated performance caused by randomly splitting cells from the same donor. Candidate populations can be verified with flow-based, targeted expression, or orthogonal tissue assays in additional samples.

    Research basis: A published melanoma study used single-cell RNA sequencing for discovery and then flow-based analysis in additional samples to investigate S100A9-positive monocytes as a response-associated biomarker. The design demonstrates the value of moving from high-dimensional discovery to a practical orthogonal assay.

    4

    Multi-Omics Stratification and Risk Research

    Client question: Do genomic, transcriptomic, epigenomic, protein, or metabolite layers define reproducible subgroups or improve a risk model beyond established variables?

    Illustrative design: Harmonize sample identity and metadata across layers, define which samples have complete or partial data, and establish whether integration is early, intermediate, or late. Evaluate each molecular layer separately before testing the combined model.

    Machine-learning role: Identify stable latent factors or cross-layer features, compare clinical-only and single-layer baselines with the integrated model, perform sensitivity analyses for missing layers and batch effects, and translate the result into an interpretable panel or subgroup definition.

    Validation route: Reproduce the subgroup or risk association in a separate cohort and verify key features with a smaller targeted assay. The objective is not to maximize the number of omics layers but to show which layer changes the research decision.

    Research basis: Hoadley et al. integrated aneuploidy, DNA methylation, mRNA, microRNA, and protein measurements across approximately 10,000 tumors and showed how individual and integrated molecular layers reveal both shared and tissue-of-origin patterns. The study supports evaluating each layer before interpreting an integrated subtype.

    Technology and Service Options

    The technology is selected around the biological question, sample type, expected signal, and downstream validation route—not around algorithm novelty.

    Service TechnologyWhat It ContributesCommon Biomarker UseValidation Consideration
    RNA SequencingGene, transcript, splice, and pathway-level expression featuresResponse signatures, molecular subgroups, disease-state programsTargeted expression measurement or independent transcriptomic cohort
    Single-Cell RNA SequencingCell populations, cell states, and donor-level cellular compositionRare-cell and immune-state biomarker discoveryFlow-based or targeted expression verification in additional donors
    ATAC-SeqChromatin accessibility and regulatory-state featuresRegulatory biomarkers and mechanism-linked signaturesTargeted regulatory-region or expression follow-up
    Whole-Exome SequencingCoding variants, mutational patterns, and copy-number-related featuresVariant-informed stratification and response studiesOrthogonal variant confirmation and independent cohort assessment
    Multi-Omics ServicesIntegrated evidence across molecular layersCross-layer subtyping, response, and risk researchShow incremental value over clinical and single-layer baselines
    Bioinformatics ServicesData review, processing, harmonization, modeling, and reproducible reportingData-to-Insight and hybrid projectsClaims depend on metadata, quality, and available validation data

    Proteomics, metabolomics, targeted measurement, spatial profiling, and other project-specific assays may also be incorporated after feasibility review.

    Project Entry Modes

    Entry ModeStarting MaterialIntegrated Scope
    Sample-to-InsightBiospecimens plus research question and metadataStudy-design support, omics experiments, QC, bioinformatics, machine-learning-assisted discovery, and validation planning
    Data-to-InsightRaw or processed omics data and metadataData audit, harmonization, confounder review, feature engineering, model comparison, interpretation, and validation strategy
    Hybrid StudyExisting data plus samples for a missing layer or follow-up cohortTargeted new experiments integrated with client data, followed by discovery and research validation

    Integrated Biomarker Discovery Workflow

    One coordinated workflow connects the intended research decision to cohort design, wet-lab execution, modeling, and validation.

    Machine learning biomarker discovery workflow from research question and cohort design through omics experiments, model comparison, independent assessment, and research validation

    Step 1 — Define the decision: Specify the comparison, endpoint, intended sample unit, candidate-panel constraints, and evidence needed for advancement.

    Step 2 — Design the cohort: Review inclusion criteria, balance, covariates, batches, paired or longitudinal structure, and options for held-out or independent assessment.

    Step 3 — Generate fit-for-purpose omics data: Select and execute the wet-lab assay or assay combination, with quality thresholds linked to downstream modeling needs.

    Step 4 — Build analysis-ready features: Process raw data, annotate features, review missingness and technical variation, and lock outcome-independent preprocessing.

    Step 5 — Compare baselines and models: Evaluate conventional statistics and justified machine-learning approaches under the same nested or held-out plan.

    Step 6 — Test robustness and interpret biology: Quantify feature stability, review confounders, assess calibration where relevant, and connect candidates to pathways or cell context.

    Step 7 — Validate the research finding: Evaluate the locked signature in independent data and/or verify candidates with a targeted orthogonal assay.

    Step 8 — Deliver the decision package: Report the candidate panel, supporting evidence, limitations, reproducible outputs, and recommended next experiments.

    Machine Learning with Research Safeguards

    Machine learning is an assistive research layer, not an automatic answer. Depending on the endpoint and sample structure, methods may include regularized linear models, tree-based models, kernel approaches, ensemble learning, or justified integration methods. Every complex model should be compared with a simpler, interpretable baseline.

    • Feature processing inside the resampling loop
    • Nested cross-validation for tuning and internal estimation
    • Grouped or site-aware splitting when samples are related
    • Training-only handling of class imbalance
    • Feature-selection frequency and stability reporting
    • Confounder and sensitivity analyses
    • Calibration assessment when probabilities matter
    • Locked independent-cohort evaluation when feasible

    A model is not advanced on one metric alone.

    Discrimination, error profile, calibration, stability, biological interpretation, and practical assay feasibility are reviewed together.

    Study Inputs and Sample Considerations

    There is no universal minimum cohort size. Feasibility depends on endpoint prevalence, heterogeneity, effect size, feature dimensionality, cohort structure, missingness, and the validation claim.

    • Sample type, preservation, input quantity, and anticipated quality
    • Endpoint definition, group balance, time points, and paired measurements
    • Age, sex, treatment, site, batch, and other relevant covariates
    • Raw-data availability and metadata completeness
    • Discovery, tuning, held-out, and independent-cohort options
    • Candidate-panel size and downstream assay constraints
    • Data-transfer, privacy, and reproducibility requirements

    Deliverables

    • Experimental QC and omics data reports
    • Analysis-ready matrices and metadata review
    • Locked splitting and validation plan
    • Baseline and machine-learning comparisons
    • Cross-validation and held-out summaries
    • Discrimination, error, and calibration outputs
    • Feature stability and selection-frequency report
    • Evidence-tiered candidate biomarker panel
    • Biological annotation and pathway context
    • Targeted-validation results or recommendations
    • Reproducible outputs and methods-ready report
    • Limitations and recommended next experiments

    References

    1. Walsh I, Fishman D, Garcia-Gasulla D, et al. DOME: recommendations for supervised machine learning validation in biology. Nature Methods. 2021.
    2. Diaz-Uriarte R, Gómez de Lope E, Giugno R, et al. Ten quick tips for biomarker discovery and validation analyses using machine learning. PLOS Computational Biology. 2022.
    3. Ney A, Nené NR, Sedlak E, et al. Identification of a serum proteomic biomarker panel using diagnosis specific ensemble learning and symptoms for early pancreatic cancer detection. PLOS Computational Biology. 2024.
    4. Xing X, et al. Proteomics-driven noninvasive screening of circulating serum protein panels for the early diagnosis of hepatocellular carcinoma. Nature Communications. 2023.
    5. Sammut SJ, Crispin-Ortuzar M, Chin SF, et al. Multi-omic machine learning predictor of breast cancer therapy response. Nature. 2022.
    6. Rad Pour S, Pico de Coaña Y, Martinez Demorentin X, et al. Predicting anti-PD-1 responders in malignant melanoma from the frequency of S100A9-positive monocytes in the blood. Journal for ImmunoTherapy of Cancer. 2021.
    7. Hoadley KA, Yau C, Hinoue T, et al. Cell-of-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer. Cell. 2018.

    For Research Use Only. Not for use in diagnostic or clinical procedures.

    Example Decision-Oriented Output

    A typical report combines performance, stability, calibration, and biological interpretation rather than presenting one accuracy number in isolation.

    Illustrative biomarker report combining nested model performance, calibration, feature stability, and candidate panel interpretation

    Illustrative composite output: nested assessment distributions, held-out ROC and precision-recall curves, calibration, feature-selection frequency, and an evidence-tiered compact panel. Final metrics and plots depend on the study endpoint and design.

    Machine Learning Biomarker Discovery FAQs

    1. Can the project start from biospecimens?

    Yes. Sample-to-Insight projects may include study-design support, omics experiments, quality control, bioinformatics, machine-learning-assisted discovery, and validation planning. The assay scope is selected after sample and endpoint review.

    2. Can you analyze data generated elsewhere?

    Yes. We first review raw or processed files, metadata, quality information, cohort structure, and compatibility with the intended claim. Missing metadata or major batch differences may narrow the feasible analysis.

    3. Which omics layer should we choose?

    The choice depends on mechanism, sample accessibility, expected abundance, relationship to the endpoint, budget, and validation route. We prioritize the smallest defensible design rather than adding layers without a decision purpose.

    4. Do you always use deep learning?

    No. Regularized or tree-based models are often more appropriate for limited omics cohorts and easier to interpret. Method choice follows data size, endpoint, structure, and validation needs.

    5. How do you reduce overfitting?

    The plan may use nested cross-validation, grouped splits, locked held-out data, preprocessing inside resampling, stability analysis, and independent assessment. The design is agreed before model tuning.

    6. Can selected biomarkers be validated experimentally?

    Potentially. Follow-up may use targeted sequencing, targeted expression, flow-based measurement, or another orthogonal assay in retained or independent research samples, subject to feasibility.

    7. Can performance be guaranteed?

    No. Biomarker discovery is uncertain. The service reduces avoidable bias, quantifies robustness, and helps determine whether a candidate should advance, be refined, or be discontinued.

    Published Case Study

    Independent Research Highlight

    Serum Proteomic Biomarker Panel Using Diagnosis-Specific Ensemble Learning

    This independent publication is presented as a study-design example and is not a CD Genomics customer project.

    Research question

    Could a compact serum protein signature distinguish pancreatic cancer from healthy individuals and clinically relevant benign conditions among people with concerning symptoms?

    Study design

    The researchers analyzed 539 serum samples using an oncology protein panel plus additional markers. Sixteen specialized base learners were combined in a stacked ensemble. Feature elimination and cross-validation were used during development, while a held-out set was reserved for assessment.

    Reported result

    In the held-out validation set, the ensemble achieved an area under the ROC curve of 0.95 (95% confidence interval 0.91–0.99) and sensitivity of 0.86 (95% confidence interval 0.68–1.00) at 90% specificity.

    Published study performance comparison for base learners and a stacked serum proteomic biomarker ensemble

    Transferable lesson

    The same performance should not be expected in another cohort. The useful lesson is the connected workflow: clinically relevant controls, broad protein measurement, feature reduction inside model development, complementary learners, held-out evaluation, and a compact panel suitable for further verification.

    Reference

    1. Ney A, Nené NR, Sedlak E, et al. Identification of a serum proteomic biomarker panel using diagnosis specific ensemble learning and symptoms for early pancreatic cancer detection. PLOS Computational Biology. 2024.

    Selected Publications

    These customer-related publications illustrate omics datasets and research questions relevant to biomarker studies. Inclusion does not imply that each paper used the complete solution described here.

    1. Iparraguirre L, Alberro A, Iñiguez SG, et al. Blood RNA-Seq profiling reveals a set of circular RNAs differentially expressed in frail individuals. Immunity & Ageing. 2023.
    2. Van Goubergen J, Peřina M, Handle F, et al. Targeting the CLK2/SRSF9 splicing axis in prostate cancer leads to decreased ARV7 expression. Molecular Oncology. 2025.
    3. Joruiz SM, Von Muhlinen N, Horikawa I, et al. Distinct functions of wild-type and R273H mutant Δ133p53α differentially regulate glioblastoma aggressiveness and therapy-induced senescence. Cell Death & Disease. 2024.
    À des fins de recherche uniquement, non destiné à un diagnostic clinique, un traitement ou des évaluations de santé individuelles.
    Services connexes
    Demande de devis
    ! À des fins de recherche uniquement, non destiné à un diagnostic clinique, un traitement ou des évaluations de santé individuelles.