Week 1 of this course introduces the nod, nif and fix genes; week 5 annotates them; week 7 asks why two strains of one species share so few genes. Each of those labs uses a teaching fixture, because until recently nothing else would run in a browser. This page does the same three questions on the actual genome — Sinorhizobium meliloti 1021, the strain the course is about, with NCBI's own annotation — and everything below is computed here, from those files, while you watch.
All of it. The genome is NC_003047.1 with plasmids NC_003037.1 (pSymA) and NC_003078.1 (pSymB), 5,945 annotated proteins, served from this site so nothing has to be downloaded by hand. The comparison sequences come from UniProt, live, and every one is a reviewed entry. The alignment is muscle, the tree is fasttree, and both run in this tab.
Start by bringing the annotation and the proteome onto the page, and counting what is in them. Three replicons, and the coding sequences are not spread evenly across them.
Reading sinorhizobium.gff…
That is a real GFF3: one row per feature, tab separated, with the attributes crammed into the ninth column. The product= in that column is the annotator's claim about what the protein does, and it is the field the next cell reads — because a locus tag tells you nothing you can check, and a product name does.
chr1$1BRCA1$2missense$3HIGH$4Before any Python, the shell. awk -F'\t' '$3=="CDS"{print $1}' sinorhizobium.gff | sort | uniq -c counts coding sequences per replicon, and grep -ci nitrogenase sinorhizobium.gff is a one-line version of the question this page is about. Type them into the terminal below.
Type a command, or click the suggested one. ↑/↓ recall history, Tab completes file names. New to this? Hit Explain this command to see it in plain English.
Now the count properly, with the answer drawn. The cell looks for products named for nodulation, nitrogenase or nitrogen fixation, tallies them per replicon, and prints any that are not where you would expect.
A gene on a megaplasmid can be lost without killing the cell, and can move between cells. So two strains of one species really can differ by hundreds of genes, and the ones they differ by are disproportionately the ones that decide what the organism does — which host it nodulates, what it resists, what it can eat. The core genome is the housekeeping; the accessory genome is the biology anybody notices.
Note the exception the cell printed rather than hid, and go and read what it actually is. Two questions settle it: does that product do nitrogen fixation, or does it regulate something about nitrogen? And did it match because of what it is, or because the word "nitrogen" appears in its name? The general nitrogen-regulation machinery is housekeeping and belongs on a chromosome; the machinery of the symbiosis is what sits on the plasmid. A search over product names is a blunt instrument, and knowing where it is blunt is the skill.
The second half is week 8's question, on real sequences. The nitrogenase iron protein does the same job in every organism that fixes nitrogen, whether it lives in a root nodule, in soil, or in a hot spring. So its tree can be compared with what we know about the organisms — and where a protein's tree disagrees with its host's, something has moved between them.
Reading nifH.faa…
Align them and build the tree, in the terminal above. muscle -align nifH.faa -output nifH.aln is the alignment; fasttree nifH.aln > nifH.nwk is the tree. Both are the tools themselves, and both take a second or two on ten sequences. While you are there, needle -asequence nifH.faa -bsequence nifH.faa -outfile pair.needle is EMBOSS, and it will show you what a percentage identity is actually counting.
Type a command, or click the suggested one. ↑/↓ recall history, Tab completes file names. New to this? Hit Explain this command to see it in plain English.
No tree on this page yet. fasttree -nt aligned.fa > tree.nwk in a terminal cell writes one, and it will appear here.
nifH.nwk is not on this page, or is not text
# The tree this cell drew, drawn by ape instead.
install.packages("ape")
library(ape)
tree <- read.tree("nifH.nwk")
plot(tree, use.edge.length = TRUE)
add.scale.bar()
nodelabels(tree$node.label, frame = "none", adj = c(1.1, -0.3), cex = 0.7)
nifH.nwk drawn to scale; the export draws it with ape.
fasttree writes SH-like local support onto its branches by default, not bootstrap percentages. Both are numbers between 0 and 1 in roughly the same place, which is exactly why they get confused. A bootstrap value is the fraction of resampled alignments in which a branch reappeared; an SH-like support is a test of that branch against the two alternative arrangements of the same four groups. Neither is the probability that the branch is true, and the course's week-8 warning applies to both.
Read it against the list you fetched. The rhizobia are five of the ten and they nodulate legumes; Frankia nodulates alders and is not a proteobacterium at all; Klebsiella and Azotobacter fix nitrogen free-living; the methanococcus is there to root the thing. If the tree grouped by lifestyle rather than by organism, that would be a claim about horizontal transfer — so check which it did before you say either.
Week 8's lab types a distance matrix for five strains into the cell and clusters it. Here is that same exercise with the typing removed: the distances are computed from the alignment you just built, and the clustering is base R, so nothing has to be installed.
One last look, at the protein itself. Nitrogenase iron protein is a dimer with an iron-sulfur cluster held between the two subunits, and it spends ATP to push electrons uphill into the enzyme that actually breaks the nitrogen triple bond. The model below is Sinorhizobium's own, by the accession you fetched first.
url=$(curl -s 'https://alphafold.ebi.ac.uk/api/prediction/P00460' | python3 -c 'import json,sys; print(json.load(sys.stdin)[0]["pdbUrl"])')
curl -sL "$url" -o NifH.cif
# Open NifH.cif in PyMOL, ChimeraX, or https://molstar.org/viewer/
AlphaFold's model of P00460, kept as NifH.cif
A page becomes a picture when you insert it, so this notebook needs no PDF reader to show it. It stops following the deck at that moment, which is why each one records the file and page it came from.
What this page did not do is decide anything. The annotation's product names are software's opinion and some of them are wrong; ten sequences is a small tree and a different ten would move it; and a protein tree that disagrees with an organism tree is evidence for horizontal transfer only after you have ruled out the alignment, the model and the sampling. The tools are free now. Knowing whether the answer is true is still the job.
The course's own week-5 and week-7 labs use invented files and say so. This page replaces them for one organism. What it does not replace is the annotation step itself: prokka and roary are not compiled for the browser yet, so the annotation here is NCBI's rather than one you produced. Assembling this genome from reads is the other missing half — see the build plan's assembler item.