Courses

Introduction to Bioinformatics and Computational Genomics

Eleven weeks from the command line to microbial community analysis, using one bacterium the whole way through.

Christian-Albrechts-Universitat zu Kiel

One organism, all term

Rhizobia and the legume symbiosis

Every exercise uses the same system: nitrogen-fixing bacteria that form root nodules on legumes. A bacterium is small enough to assemble, annotate and compare in a browser, and the analysis steps are the same ones you would run on wheat. The nod, nif and fix genes you meet in week 1 are the ones you annotate in week 5, compare in week 7 and measure the expression of in week 10.

  1. Week1

    Bioinformatics, genomics, and our model organism

    What the field is, how genome sequencing got from a 3,569-base phage to a 17-gigabase wheat, and the bacterium we will use for everything that follows.

    SlidesCheck
    28 min
  2. Week2

    The command line

    The shell is the interface to every tool in this course. Read files, filter them, count things, and chain small tools into one answer.

    SlidesLabCheck
    98 min
  3. Week3

    Biological databases and sequence search

    Where sequence data lives, why the archives are full of errors, and how BLAST finds a needle in two trillion bases fast enough to be useful.

    SlidesExternalCheck
    90 min
  4. Week4

    Sequencing technologies and genome assembly

    How reads are made and what they cost you in quality, then the two graph algorithms that put them back together, and the metrics that tell you whether it worked.

    SlidesLabCheck
    101 min
  5. Week5

    Genome annotation

    An assembly is a string until you say where the genes are. Find open reading frames, model them with hidden Markov models, then interrogate the annotation file.

    SlidesLabCheck
    61 min
  6. Week6

    Variant calling and structural variants

    Once you have reads aligned to a reference, the question becomes which differences are real. This is the unit where the tools you run are the actual tools the field uses.

    SlidesLabCheck
    63 min
  7. Week7

    Gene content and pan-genomes

    The E. coli result from week 1, done properly: cluster genes into families across strains, then split the result into core, shell and cloud.

    SlidesLabCheck
    61 min
  8. Week8

    Phylogenetic trees

    How to read a tree without over-reading it, why you cannot simply try every tree, and what a bootstrap value is actually measuring.

    SlidesNotebookCheck
    66 min
  9. Week9

    R

    Everything after this point is statistics on tables. R is how the field does that, and it runs in your browser.

    SlidesNotebookCheck
    48 min
  10. Week10

    Transcriptomics and differential expression

    From a count table to a list of genes that changed, and the three sources of noise that stop you believing the first answer.

    SlidesLabNotebookCheck
    98 min
  11. Week12

    Microbiomes and microbial communities

    Most bacteria have never been cultured. Sequencing lets you study them anyway, and almost all of the analysis is arithmetic on one table.

    SlidesNotebookCheck
    88 min