Independent bioinformatics contractor
Metagenomic data,
turned into answers.
We design and run analysis pipelines for shotgun and amplicon metagenomic sequencing — from raw reads to publication-ready figures — for research groups, public health labs, and biotech teams working across Africa, and beyond.
Why metagenomics needs a specialist
A shotgun run doesn't arrive as an answer. It arrives as several hundred million short reads, a constellation of databases, and a deadline.
We sit between the sequencer and the paper — cleaning, assembling, classifying, and quantifying microbial communities so the biology, not the bioinformatics, is what your team spends its time on. Every pipeline is version-controlled, documented, and handed to you at the end.
Our specialists
Domain expertise on tap.
Marc's research focuses on desert and dryland soil microbiomes, biocrust ecology, and long-read metagenomic assembly, with previous work at Lawrence Berkeley National Laboratory, the DOE Joint Genome Institute, the University of Pretoria, and KAUST. He advises on assembly strategy, MAG recovery, and functional annotation for environmental sample sets.
View full profile on Google Scholar →Analysis, sized to the question you're asking.
Engagements range from a single diversity analysis on an existing OTU table to a full pipeline built from raw FASTQ files upward.
Read QC & preprocessing
Adapter and quality trimming, host and contaminant removal, error correction — the unglamorous work that determines whether everything downstream can be trusted.
Taxonomic profiling
Species- and strain-level community composition from shotgun or 16S/18S/ITS amplicon data.
Assembly & MAG recovery
Metagenome assembly, binning and refinement into draft genomes, with completeness/contamination scoring and dereplication across samples.
Functional annotation
Gene prediction and pathway mapping — what the community can do, not just who's present, including AMR and virulence gene screening on request.
Comparative & diversity analysis
Alpha/beta diversity, differential abundance, ordination and statistical testing across cohorts, timepoints or treatment groups.
Custom pipeline development
A workflow built to your data, your cluster and your reviewers — documented, containerised and handed over as something your team can run without me.
Raw reads to report, in five stages.
This is the standard order of operations for a shotgun metagenomics project. Stages tailored to you.
Raw reads
FASTQ intake, sequencing QC, sample sheet and metadata validation.
Quality control
Trimming, filtering, host/contaminant depletion, duplicate removal.
Assembly & binning
Contig assembly, binning into draft genomes, quality scoring.
Classification & annotation
Taxonomic identity and functional/pathway annotation per bin or read set.
Statistics & report
Diversity metrics, group comparisons, figures and a written summary.
Established tools, not reinvented ones.
Pipelines are built from peer-reviewed, widely benchmarked software — chosen per project, never a one-size-fits-all default.
Published research behind the practice.
A selection of Dr Marc Van Goethem's peer-reviewed work in metagenomics and microbial ecology — the same methods brought to client projects. Full list on Google Scholar.
Long-read metagenomics of soil communities reveals phylum-specific secondary metabolite dynamics
Combines short- and long-read sequencing of biocrust samples to recover thousands of biosynthetic gene clusters, showing what integrated long-read approaches reveal that short reads alone miss.
Assembling metagenomes, one community at a time
A review of metagenome assembly strategies and the tradeoffs between assemblers — foundational reading for choosing an assembly approach per sample type.
Cyanobacteria and Alphaproteobacteria may facilitate cooperative interactions in niche communities
Structure and co-occurrence patterns of hypolithic microbial communities in the Namib Desert, comparing total and metabolically active community fractions.
Characteristics of wetting-induced bacteriophage blooms in biological soil crust
Genome-resolved metagenomics linking phage populations to their bacterial hosts during a soil-wetting event, the first characterisation of bacteriophage dynamics in biocrust.
Learning representations of microbe–metabolite interactions
A neural network method (mmvec) for inferring conditional relationships between microbes and metabolites across paired omics datasets.
Diamonds in the rough: dryland microorganisms as ecological engineers
A synthesis of how dryland microbial communities can be applied to restore degraded land and mitigate desertification.
Get in touch
Tell me about your data.
Sample type, sequencing platform, read count and what you need at the end — that's enough to start scoping. A reply usually follows within two working days.
- Email hello@omix.africa
- Typical scope 1 sample set to multi-cohort studies