1. Home
  2. Services
  3. De novo genome assembly
Genomics · long reads

De novo genome assembly: a genome sequence for an organism that has none.

We sequence long DNA molecules on PacBio or Oxford Nanopore, join them into a genome, check how complete and clean it is, and confirm which species it is.

Agarose gel with DNA bands glowing under blue light
How it works

Long reads keep a genome in one piece.

De novo means from scratch: there is no reference genome to compare against, so the genome is rebuilt from the reads alone. Overlapping reads are joined into continuous stretches called contigs. The fewer and longer the contigs, the better the assembly.

Every genome contains repeats, stretches of DNA that occur in several places. A short read that falls inside a repeat could belong to any copy, so short-read assemblies break at repeats. A long read runs across the repeat into the unique DNA on both sides and shows how the pieces connect.

PacBio HiFi reads are long and highly accurate; Oxford Nanopore reads can be even longer. Short Illumina reads can be added to polish remaining errors. Before sequencing, the expected genome size is checked in public databases so the amount of sequencing fits the genome.

Diagram: a genome with two copies of a repeat; short reads that cannot be placed and long reads spanning the repeats; overlapping reads joined into contigs; summary values for contig N50, gene completeness and contamination screen.
Short reads inside a repeat could belong to either copy. Long reads span the repeats and overlap into contigs, and the result is checked for size, completeness and contamination. Tap the diagram to enlarge it.

Key terms

Contig
A continuous stretch of assembled sequence without gaps.
N50
Half of the assembly sits in contigs this long or longer. Higher is better.
BUSCO
A completeness check that looks for genes almost every species in a lineage carries.
HiFi read
A PacBio long read made by reading one DNA molecule many times, more than 99.9% accurate.
HMW DNA
High-molecular-weight DNA: long, intact molecules, needed for long reads.
ANI
Average nucleotide identity between two genomes. Microbes that share 95% or more are usually the same species.
Questions this answers

Bring the question. We design the experiment around it.

  • Nobody has sequenced this species. Can I get a genome to map my other data against?
  • How big and how heterozygous is this genome?
  • Is my bacterial chromosome closed, and which plasmids does the strain carry?
  • Is my isolate really the species I think it is, and is the culture pure?

Where it is used

  • Reference genomes for new speciesA first genome for organisms without one, or a better one than a fragmented draft.
  • Complete bacterial and yeast genomesClosed chromosomes and plasmids from long reads.
  • Isolate identity and puritySpecies confirmation and detection of mixed or contaminated cultures.
  • Protists and other eukaryotesGenomes with gene models for less-studied single-celled organisms.
  • Methylation from the same run5mC and 6mA marks read directly from native DNA.
  • A base for later studiesThe assembly becomes the reference for resequencing, RNA-seq and population work.
Your project

Four steps from DNA to a checked genome.

  1. Step 1

    Size the project

    We check the expected genome size in public databases and plan the DNA extraction and sequencing depth with you.

  2. Step 2

    Check the DNA

    Long reads need long, intact DNA, so amount and fragment length are checked before library preparation.

  3. Step 3

    Sequence long reads

    PacBio HiFi or Oxford Nanopore, with Illumina reads added when polishing helps.

  4. Step 4

    Assemble and check

    Reads are assembled, then checked for contiguity, gene completeness and contamination, and delivered with a report and Methods.

Results

What the analysis shows you.

  • K-mer spectrum with a peak for DNA words found on one copy at about 15× and a larger peak for words on both copies at about 30×; estimated genome size 412 Mb, heterozygosity 0.9%.

    Sizing the genome

    Counting short DNA words in the reads reveals the genome size and how different the two parental copies are, before the assembly starts.

  • Bar chart of genome identity between a new assembly and six related reference genomes; the closest reaches 87.1%, below the 95% species line.

    Is it the species you think?

    Genomes of one species share at least 95% identity. This isolate’s closest match falls below that line, which points to a species not yet described.

  • Scatter plot of contig length against read coverage: one closed circular chromosome, two circular plasmids with higher coverage, and three short contigs from another organism.

    Every contig accounted for

    The chromosome and plasmids close into circles. Plasmids have higher coverage because cells carry several copies, and stray contigs from another organism are removed.

Included as standard every project

  • AssemblyLong reads joined into contigs, with an assembly graph that shows how they connect.
  • Contiguity statisticsNumber of contigs, total length and N50.
  • Gene completenessBUSCO check against the lineage your organism belongs to.
  • Report and MethodsA readable report, a Materials & Methods document and checksums for every file.

Added for your question custom

  • Genome surveyGenome size and heterozygosity estimated from the reads.
  • Species identity and contamination16S and whole-genome identity to reference genomes, and a check of every contig.
  • Closed circles for microbesCircular chromosomes and plasmids marked, with read coverage per contig.
  • Polishing with short readsIllumina reads correct remaining errors in a long-read draft.
  • Gene modelsGene structures predicted for eukaryotic genomes.
  • Methylation5mC and 6mA calls from PacBio HiFi data, delivered as genome tracks.
Typical project

What goes in, and how we run it.

Sequencing
PacBio Revio HiFi · Oxford Nanopore PromethION · Illumina for polishing
Typical depth
About 60× HiFi coverage of the expected genome size
Before sequencing
Genome size checked against public databases
Input
High-molecular-weight DNA in microgram amounts · extraction planned with you
Organisms
Bacteria, yeasts and fungi, protists · plants and animals
Main tools
hifiasm, Flye, SPAdes, QUAST, BUSCO, GenomeScope2
Two-panel figure: a, Nx curves comparing HiFi, hybrid and Nanopore assemblies with contig N50 values; b, BUSCO completeness bars for the three assemblies.
Fig. 1 | Assembly contiguity and completeness. a, How much of the genome sits in contigs of a given length, for three sequencing strategies. b, Share of expected genes found complete (BUSCO).
What you receive

Figures, interpretation and Methods, ready for the manuscript.

  • ReportInterpretation written against your hypothesis
  • FiguresPublication-ready figures for the manuscript
  • MethodsMaterials & Methods text for the paper
  • Tables and dataAll results as tables, plus the processed data files
Choosing a method

PacBio HiFi or Oxford Nanopore?

Send us the question and we recommend one, including when the cheaper option is enough.

AspectPacBio HiFiOxford Nanopore
Read accuracyAbove 99.9% per readLower per read, corrected by depth and polishing
Read lengthLong, typically 15–25 kbThe longest reads of any platform
DNA neededHigh-molecular-weight DNAHigh-molecular-weight DNA, more for the longest reads
MethylationRead from native DNARead from native DNA
Best whenAccurate, contiguous assemblies of most genomesVery long repeats, or combined with short reads in a hybrid assembly
Questions

Before you send samples

Related services

PacBio HiFi, Nanopore or hybrid?

HiFi gives very accurate long reads and suits most assemblies. Nanopore gives the longest reads. A hybrid adds Illumina reads to a Nanopore draft to correct the remaining errors. We recommend one in the design review.

How much DNA do I need, and what quality?

Microgram amounts of long, intact DNA. Handle it gently: wide-bore tips, no vortexing, no repeated freezing and thawing, and measure it fluorometrically. If extraction is difficult for your organism, we plan it with you or do it for you.

Do I need to know my genome size first?

It sets the amount of sequencing, so we check it in public databases before quoting. After sequencing, the reads give a second estimate.

How do you judge assembly quality?

By contiguity (number of contigs and N50), completeness (BUSCO genes found) and a contamination check of the contigs.

Will my bacterial genome come out as one closed circle?

Each contig is marked as circular or linear, so you see at once which chromosomes and plasmids are closed.

Is my isolate what I think it is?

We compare the assembly with reference genomes by 16S and by whole-genome identity, and check every contig and the unassembled reads for other organisms. Past projects have revealed a likely new species and a culture of three organisms.

Can you detect methylation?

Yes. 5mC and 6mA are read directly from PacBio HiFi data and delivered as genome tracks.

The 100% Complete Project Guarantee

If something goes wrong on this project, you pay €0 to fix it.

  • Sample problems €0
  • Library prep redo €0
  • Resequencing €0
  • Reviewer re-analysis €0