1. Home
  2. Services
  3. Whole-genome & exome sequencing
Genomics

Whole-genome and exome sequencing, delivered as a short list of variants that matter.

We read a whole genome or all its protein-coding genes, find where it differs from the reference and classify the variants that could matter for your question.

Blood collection tubes with grey, purple and red caps
How it works

Line the reads up, then look for differences.

Whole-genome sequencing reads the entire genome: genes, the stretches between them and the regulatory regions. Exome sequencing first pulls out the exons, the protein-coding parts that make up about 1.5% of the human genome, and reads only those. The exome costs less per sample; the genome also sees non-coding and regulatory variants.

The short reads are placed on a reference genome. Where many reads show a different letter at the same position, the sample carries a variant. Coverage, the number of reads stacked on each position, decides how confidently a variant is called: about 30 reads for a human genome, more for an exome.

One human exome yields well over a hundred thousand variants, and almost all of them are harmless. We keep the variants that independent calling methods agree on, set aside those that are common in the population, and classify the rest on the five-tier ACMG scale from pathogenic to benign.

Diagram: reads spread along a whole genome compared with reads stacked on exons; below, reads aligned to a reference sequence with one position where three of five reads show A instead of G.
Reads cover the whole genome or pile up on the exons. Lined up against the reference, positions where reads differ reveal the variants. Tap the diagram to enlarge it.

Key terms

Exome
All protein-coding parts of the genome, about 1.5% of human DNA.
Coverage
How many reads, on average, cover each position. 30× means 30 reads.
SNV and indel
A single-letter change, and a small insertion or deletion of letters.
Heterozygous
Only one of the two chromosome copies carries the variant, so about half the reads show it.
Population frequency
How common a variant is in large reference populations. Common variants rarely cause rare disease.
ACMG classification
The five-tier scale for grading variants: pathogenic, likely pathogenic, uncertain, likely benign, benign.
Questions this answers

Bring the question. We design the experiment around it.

  • Which variant in this genome could explain the phenotype?
  • Is this variant rare, and is it likely to damage the protein?
  • How should each candidate variant be classified?
  • Do I need the whole genome, or is the exome enough?

Where it is used

  • Rare and inherited disease researchSearching a genome or exome for the variant behind a phenotype.
  • Candidate gene screeningChecking the genes linked to a condition for rare, damaging variants.
  • Viral genomesConsensus genomes, variants and family trees of viruses.
  • Bacterial isolatesWhole genomes of isolates for characterisation.
  • Plant and animal genomesResequencing of lines, breeds and strains against their reference.
  • Species without a referenceWhen no reference exists, the genome is assembled first.
Your project

Four steps from DNA to a classified variant list.

  1. Step 1

    Choose genome or exome

    We match the method and coverage to the question: exome for coding variants, genome when non-coding regions matter too.

  2. Step 2

    Check the DNA

    DNA amount and integrity are checked before library preparation, so problems show up early.

  3. Step 3

    Sequence

    Libraries are sequenced as 150-letter read pairs to the agreed coverage.

  4. Step 4

    Call, filter and classify

    Variants from three callers are compared, filtered by population frequency and predicted impact, classified and summarised in an interactive report.

Results

What the analysis shows you.

  • Horizontal bars on a log scale showing variant numbers at four filtering steps, from 141,000 calls to 62 shortlisted variants in one exome.

    From raw calls to a shortlist

    Each filter removes variants that are common or unlikely to matter. Of about 140,000 calls in one exome, a few dozen are left to classify.

  • Histogram of the share of reads carrying each variant, with one peak near 50% for heterozygous and one near 100% for homozygous variants.

    One copy or both?

    The share of reads carrying each variant. Variants on one chromosome copy sit near 50%, variants on both copies near 100%.

  • Horizontal bar chart of 62 shortlisted variants by ACMG class: 2 pathogenic, 3 likely pathogenic, 11 uncertain, 22 likely benign and 24 benign.

    Classified variants

    The shortlist graded on the five-tier ACMG scale, with the reasoning behind each classification in the report.

Included as standard every project

  • Alignment and clean-upReads are placed on the reference genome, with duplicates marked and base qualities recalibrated.
  • Consensus variant callingThree independent callers; variants found by at least two are kept.
  • AnnotationEach variant is linked to its gene, its effect on the protein, and population and disease databases.
  • FilteringRare variants that are predicted to be damaging or known as pathogenic are shortlisted.
  • Classification and reportACMG classification with reasoning and literature, in an interactive report with a list of critical findings.

Added for your question custom

  • Viral genomesConsensus genomes, variants and phylogenetic trees from viral sequencing data.
  • Somatic tumour reportsTumour variants reported following ClinGen, CGC and VICC guidelines.
  • HLA typing and copy numberHLA alleles, and gains or losses of DNA segments.
  • Genome size checkFor less-studied organisms, genome size is looked up in public databases before the depth is set.
Typical project

What goes in, and how we run it.

Sequencing
Illumina NovaSeq X Plus or NextSeq 2000, 150 bp read pairs
Typical coverage
About 30× for a human genome · 50–100× for germline exomes · 100× or more for tumour exomes
Variants
Single-letter changes (SNVs) and small insertions and deletions
Samples
Blood, tissue and extracted DNA · bacterial isolates, plants and animals
Delivered
Consensus VCF, annotated and classified variant tables, interactive report
Main tools
nf-core/sarek, BWA-MEM, HaplotypeCaller, FreeBayes, OpenCRAVAT
Two-panel figure: a, bar chart on a log scale of variants found by each combination of three variant callers, with sets found by two or more callers highlighted; b, horizontal bars counting coding and splice variants by predicted effect.
Fig. 1 | Calling and annotating variants. a, Variants found by each combination of the three callers; those found by two or more are kept. b, Coding and splice variants by predicted effect on the protein.
What you receive

Figures, interpretation and Methods, ready for the manuscript.

  • ReportInterpretation written against your hypothesis
  • FiguresPublication-ready figures for the manuscript
  • MethodsMaterials & Methods text for the paper
  • Tables and dataAll results as tables, plus the processed data files
Choosing a method

Genome or exome?

Send us the question and we recommend one, including when the cheaper option is enough.

AspectWhole genomeWhole exome
CoversEvery part of the genome, including non-coding regionsProtein-coding exons only
Data per sampleAbout 30× over 3 billion letters50–100× over the exome, far less data in total
FindsCoding and non-coding variantsCoding variants
Best whenRegulatory or non-coding variants matter, or an exome found nothingMany samples, with coding variants as the main suspects
Questions

Before you send samples

Related services

Whole genome or exome?

The exome covers the protein-coding genes, where most known disease variants sit, at a lower cost per sample. The genome adds introns, regulatory regions and the space between genes.

How much coverage do I need?

About 30× for a human germline genome and 50–100× for a germline exome. Tumour material, mosaic variants and bacterial isolates are sequenced deeper.

How is coverage calculated?

Read length times the number of reads, divided by genome size. For a human genome, 300 million read pairs of 150 letters each give about 30×.

How do you decide which variants matter?

We keep variants that at least two independent callers agree on, set aside those common in population databases, and shortlist variants predicted to be highly damaging or known as pathogenic. The shortlist is then classified on the ACMG scale.

Can you sequence bacteria, plants or animals?

Yes. The analysis is set up for your organism and its reference genome. If no good reference exists, the genome is assembled first.

How big is my genome, and how much sequencing does that mean?

For less-studied organisms we check the genome size in public databases before setting the depth. In one project, two genomes turned out about three times larger than expected, which would have left them under-sequenced.

I already have genome or exome data. Can you analyse it?

Yes. Send the raw FASTQ files, and for exomes tell us which capture kit was used.

The 100% Complete Project Guarantee

If something goes wrong on this project, you pay €0 to fix it.

  • Sample problems €0
  • Library prep redo €0
  • Resequencing €0
  • Reviewer re-analysis €0