Introduction to Bioinformatics and Computational Genomics

Week 3Biological databases and sequence search

Where sequence data lives, why the archives are full of errors, and how BLAST finds a needle in two trillion bases fast enough to be useful.

Questions this week answers

  • GenBank is enormous and full of errors. Why has nobody fixed it?
  • An exhaustive alignment against every sequence in a database is too slow to use. What does BLAST give up to be fast?
  • You have a DNA sequence of unknown function. Should you search with the DNA or translate it first?

By the end of this week you can

  • Read a FASTA header and a GenBank flat file, and tell RefSeq accessions apart by prefix
  • Explain why GenBank is redundant and RefSeq is not
  • Score an alignment by hand from match, mismatch and gap penalties
  • Choose the right BLAST program for a given query and database
  • Retrieve a specific genome and its coding sequences from NCBI
0 of 5 done
  1. An alignment is a hypothesis, and the penalties decide which one wins

    Two alignments of the same two sequences. Drag the penalties and watch which one the scoring function prefers.

    The ungapped alignment wins by 2. Opening two gaps to avoid one mismatch costs more than the mismatch does. Push the mismatch penalty below -3 and that flips.

    Ungapped45 matches, 1 mismatch
    Gapped25 matches, 2 gap positions
    Margin2Ungapped wins
    How much a substitution costs. The lecture uses -0.5.
    Charged once per gap, however long. The lecture uses -1.
    Charged per gap position. The lecture uses -0.5.
    Things to try0 of 3

    Both alignments and the default penalties are from the week 3 sequence-search lecture.

What should survive this week

  • GenBank is an archive and keeps everything, redundancy and errors included. RefSeq is the curated, non-redundant answer to that.
  • The accession prefix tells you what you are holding, and an X means predicted by a model rather than observed.
  • BLAST is a heuristic: words of length 3 for protein and 11 for DNA, then extension from two hits on a diagonal.
  • Two random protein sequences are about 5% identical; two random DNA sequences about 25%. Protein searches carry more signal and hit fewer accidents.
  • An optimal alignment is only optimal for the penalties you chose, and near a boundary a small change rewrites the biology you would report.