Independent bioinformatics contractor

Metagenomic data,
turned into answers.

We design and run analysis pipelines for shotgun and amplicon metagenomic sequencing — from raw reads to publication-ready figures — for research groups, public health labs, and biotech teams working across Africa, and beyond.

Foci — gut, soil, water & clinical microbiomes
Inputs — Illumina, Nanopore, PacBio HiFi data
Turnaround — pipelines built in Nextflow / Snakemake, versioned and reproducible
Delivery — reports, figures, and handed-off code

Why metagenomics needs a specialist

A shotgun run doesn't arrive as an answer. It arrives as several hundred million short reads, a constellation of databases, and a deadline.

We sit between the sequencer and the paper — cleaning, assembling, classifying, and quantifying microbial communities so the biology, not the bioinformatics, is what your team spends its time on. Every pipeline is version-controlled, documented, and handed to you at the end.

50M–2B
reads per run handled, QC to report
10+
taxonomic & functional profiling frameworks in active use
100%
reproducible — containerised, versioned pipelines

Our specialists

Domain expertise on tap.

Portrait of Dr Marc Van Goethem
Dr Marc Van Goethem
Bioinformatician & Metagenomics Researcher

Marc's research focuses on desert and dryland soil microbiomes, biocrust ecology, and long-read metagenomic assembly, with previous work at Lawrence Berkeley National Laboratory, the DOE Joint Genome Institute, the University of Pretoria, and KAUST. He advises on assembly strategy, MAG recovery, and functional annotation for environmental sample sets.

View full profile on Google Scholar →
Services

Analysis, sized to the question you're asking.

Engagements range from a single diversity analysis on an existing OTU table to a full pipeline built from raw FASTQ files upward.

01

Read QC & preprocessing

Adapter and quality trimming, host and contaminant removal, error correction — the unglamorous work that determines whether everything downstream can be trusted.

fastp · Trimmomatic · BBDuk
02

Taxonomic profiling

Species- and strain-level community composition from shotgun or 16S/18S/ITS amplicon data.

Kraken2 · MetaPhlAn4 · QIIME2 · DADA2
03

Assembly & MAG recovery

Metagenome assembly, binning and refinement into draft genomes, with completeness/contamination scoring and dereplication across samples.

MEGAHIT · metaSPAdes · CheckM2 · dRep
04

Functional annotation

Gene prediction and pathway mapping — what the community can do, not just who's present, including AMR and virulence gene screening on request.

Prokka · DIAMOND · eggNOG-mapper
05

Comparative & diversity analysis

Alpha/beta diversity, differential abundance, ordination and statistical testing across cohorts, timepoints or treatment groups.

R / Bioconductor · phyloseq · scikit-bio
06

Custom pipeline development

A workflow built to your data, your cluster and your reviewers — documented, containerised and handed over as something your team can run without me.

Nextflow · Snakemake · Docker/Singularity
How a project runs

Raw reads to report, in five stages.

This is the standard order of operations for a shotgun metagenomics project. Stages tailored to you.

1

Raw reads

FASTQ intake, sequencing QC, sample sheet and metadata validation.

2

Quality control

Trimming, filtering, host/contaminant depletion, duplicate removal.

3

Assembly & binning

Contig assembly, binning into draft genomes, quality scoring.

4

Classification & annotation

Taxonomic identity and functional/pathway annotation per bin or read set.

5

Statistics & report

Diversity metrics, group comparisons, figures and a written summary.

Toolchain

Established tools, not reinvented ones.

Pipelines are built from peer-reviewed, widely benchmarked software — chosen per project, never a one-size-fits-all default.

QIIME2 Kraken2 / Bracken MetaPhlAn4 MEGAHIT metaSPAdes CheckM2 DIAMOND eggNOG-mapper Prokka / Bakta GTDB-Tk Nextflow Snakemake R / Bioconductor phyloseq Python / scikit-bio Docker & Singularity
Representative work

Published research behind the practice.

A selection of Dr Marc Van Goethem's peer-reviewed work in metagenomics and microbial ecology — the same methods brought to client projects. Full list on Google Scholar.

Long-read metagenomics

Long-read metagenomics of soil communities reveals phylum-specific secondary metabolite dynamics

Combines short- and long-read sequencing of biocrust samples to recover thousands of biosynthetic gene clusters, showing what integrated long-read approaches reveal that short reads alone miss.

Communications Biology · 2021
Read the paper →
Assembly methodology

Assembling metagenomes, one community at a time

A review of metagenome assembly strategies and the tradeoffs between assemblers — foundational reading for choosing an assembly approach per sample type.

BMC Genomics · 2017
Read the paper →
African dryland soils

Cyanobacteria and Alphaproteobacteria may facilitate cooperative interactions in niche communities

Structure and co-occurrence patterns of hypolithic microbial communities in the Namib Desert, comparing total and metabolically active community fractions.

Frontiers in Microbiology · 2017
Read the paper →
Viral ecology

Characteristics of wetting-induced bacteriophage blooms in biological soil crust

Genome-resolved metagenomics linking phage populations to their bacterial hosts during a soil-wetting event, the first characterisation of bacteriophage dynamics in biocrust.

mBio · 2019
Read the paper →
Multi-omics

Learning representations of microbe–metabolite interactions

A neural network method (mmvec) for inferring conditional relationships between microbes and metabolites across paired omics datasets.

Nature Methods · 2019
Read the paper →
Applied ecology

Diamonds in the rough: dryland microorganisms as ecological engineers

A synthesis of how dryland microbial communities can be applied to restore degraded land and mitigate desertification.

Microbial Biotechnology · 2023
Read the paper →

Get in touch

Tell me about your data.

Sample type, sequencing platform, read count and what you need at the end — that's enough to start scoping. A reply usually follows within two working days.