Introduction to Bioinformatics and Computational Genomics

Week 5Genome annotation

An assembly is a string until you say where the genes are. Find open reading frames, model them with hidden Markov models, then interrogate the annotation file.

Questions this week answers

  • How do you find a gene in a string of four letters?
  • An open reading frame scan already finds genes. What does a hidden Markov model add?
  • Prokka finishes in a few minutes. What are you supposed to do for the rest of the afternoon?

By the end of this week you can

  • Find open reading frames and explain why six frames must be searched
  • Say what a hidden Markov model contributes that a plain ORF scan does not
  • Pull counts, coordinates and gene names out of a GFF with command-line tools
0 of 4 done
  1. Why a length cutoff cannot find genes

    Raise the GC content and watch the number of things that look like genes, but are not, go up by two orders of magnitude.

    At 50% GC a random frame runs 21.3 codons before hitting a stop, so reaching 100 by chance is unlikely: about 3,546 across the whole genome. Length is doing real work here, but it is doing it because the genome is AT-rich, not because length is a good criterion.

    ORFs this long by chance3,546against roughly 4,600 real genes
    Stop codon frequency1 in 21.3codons, at this GC
    Mean random ORF21.3 codonsbefore a stop appears
    Junk per real gene0.8x
    Plasmodium sits near 0.2, E. coli near 0.5, Streptomyces near 0.72.
    The cutoff a naive gene finder would use.
    Things to try0 of 3

    ORFs and the move to hidden Markov models are from the week 5 annotation lecture.

What should survive this week

  • Six reading frames, three on each strand. In bacteria that gets you a long way; introns are what make eukaryotes hard.
  • The states in a hidden Markov model (coding, non-coding, intergenic) are not observed. Viterbi finds the most likely path through them, which is the gene structure.
  • Bacterial genes that work together are transcribed together, which is why finding one nif gene usually means finding its neighbours.
  • The heavy tool produces a file. The learning is in interrogating that file, and the course's own practical says so outright.