TEM beta-lactamase alleles and what they hydrolyze

NCBI’s Reference Gene Catalog names 233 TEM beta-lactamase alleles and records a phenotype for many of them: broad-spectrum, extended-spectrum, inhibitor-resistant broad-spectrum, and inhibitor-resistant extended-spectrum. The proteins are 286 residues and differ from TEM-1 by a handful of substitutions, so this page puts 46 of them in one alignment, gives each row the catalog’s phenotype and the residue it carries at seven positions the literature names, and reads the two groups of positions against each other. The control is a polymorphic position picked for being the least associated with phenotype of any variable position in the data.

Prerequisites

Every figure below links to the live view it captured.

Where the data comes from

Two files of the AMRFinderPlus database, and four this page writes from them.

1. Pick the alleles

The catalog’s product_name carries the phenotype, and its gene_family column is blaTEM for every TEM allele. The selection takes the twelve lowest-numbered alleles of each phenotype, which are the ones described first and cited most:

number = re.fullmatch(r'blaTEM-(\d+)', row['allele'])
phenotype = row['product_name'].split(' class A')[0]
if not number or phenotype not in PHENOTYPES:
    continue
candidates[phenotype].append((int(number.group(1)), row))
broad-spectrum: 16 alleles in the catalog, 12 kept
extended-spectrum: 83 alleles in the catalog, 12 kept
inhibitor-resistant broad-spectrum: 30 alleles in the catalog, 12 kept
inhibitor-resistant extended-spectrum: 10 alleles in the catalog, 10 kept
46 alleles selected, lengths 285 aa x 1, 286 aa x 45

The catalog holds 94 further TEM alleles whose product name is a plain class A beta-lactamase, with no phenotype recorded, and the selection leaves them out. The inhibitor-resistant extended-spectrum class has ten members in all, so all ten are here.

2. Align and infer a tree

clustalw -INFILE=tem.fasta -ALIGN -TYPE=PROTEIN -OUTORDER=INPUT \
  -OUTPUT=FASTA -OUTFILE=tem.afa
clustalw -INFILE=tem.afa -TREE -TYPE=PROTEIN -OUTPUTTREE=phylip
tr -d '[:space:]' < tem.ph > tem.nwk
46 rows, 286 alignment columns

One column per residue of TEM-1: the only allele of another length is TEM-178, at 285 residues, and its deletion is the alignment’s one gap.

Forty-six alleles beside the tree ClustalW inferred from them. Each column draws in one color because the rows agree on it; the white and off-color specks are the 31 positions that vary.

3. The row table

Every row gets the phenotype the catalog records and the residue it carries at the seven positions named below, keyed by row name, which is the allele name:

record = {'phenotype': phenotype[row], 'subclass': subclass[row]}
for position in fields:
    record[f'Ambler {position}'] = aligned[row][ambler_column[position]]
wrote ./tem-rowdata.json: 46 rows, 10 fields, 11 kB

Eleven kilobytes percent-encode to more than the 8,192-character request line a ?data= link fits in, so the snapshot names the file and the viewer fetches it at startup:

{
  "type": "MsaView",
  "msaFilehandle": { "uri": "https://example.org/tem.afa" },
  "treeFilehandle": { "uri": "https://example.org/tem.nwk" },
  "treeMetadataFilehandle": { "uri": "https://example.org/tem-rowdata.json" },
  "encodings": [{ "channel": "tipLabel", "field": "phenotype" }]
}
The tipLabel channel reading the phenotype field: olive for extended-spectrum, cyan for inhibitor-resistant broad-spectrum, purple for inhibitor-resistant extended-spectrum, salmon for broad-spectrum. The tree puts most of the olive rows in the top half and most of the cyan rows in the bottom half, and 6 of its 44 internal nodes hold alleles of one phenotype.

4. Ambler numbering against the sequence index

The literature numbers class A beta-lactamases by Ambler’s scheme, which starts at 3, skips 239 and 253, and ends at 290 over a 286-residue protein. Ambler 238 is therefore residue 236 of the sequence, and reading a residue at index 238 lands two residues away with no error. The script maps every position through the TEM-1 row and checks the result against the residues the reviews name for TEM-1 before anything downstream uses it:

found = {p: reference[ambler_column[p]] for p in sorted(TEM1_RESIDUES)}
wrong = {p: r for p, r in found.items() if r != TEM1_RESIDUES[p]}
if wrong:
    raise SystemExit(f'Ambler mapping is off: {wrong}')
TEM-1 residues at the named Ambler positions: 69M 70S 73K 104E 130S 164R 166E 234K 238G 240E 244R 276N
the mapping is Ambler = residue index + 2 up to 238, + 3 from 240 to 252, + 4 from 254
Ambler [104, 164, 238, 240, 69, 244, 276] are columns 102, 162, 236, 237, 67, 241, 272

M69, S70, K73, E104, S130, R164, E166, K234, G238, E240, R244 and N276 are the residues TEM-1 carries in every review of the family, including the catalytic S70 and K73 and the omega-loop E166, and all twelve come out where the mapping says.

5. The positions as strips

Four positions carry the extended-spectrum substitutions (104, 164, 238, 240), three carry the inhibitor-resistance ones (69, 244, 276), and the eighth is the control of step 7. Each is a strip row panel over its field, and the eight share one legend so the figure carries one residue key:

"rowPanels": [
  { "kind": "strip", "field": "phenotype", "width": 14, "header": "phenotype" },
  { "kind": "strip", "field": "Ambler 104", "width": 12, "header": "104", "legend": "residue" }
]
The phenotype strip and the eight position strips between the tree and the alignment, in the order 104, 164, 238, 240, then 69, 244, 276, then 265. The rows whose phenotype cell is olive or purple carry a minority residue in the first four columns; the rows whose cell is cyan carry one in the next three.
Ambler 104: E 36, K 10
Ambler 164: R 29, S 11, H 6
Ambler 238: G 40, S 5, R 1
Ambler 240: E 41, K 5
Ambler 69: M 31, L 8, V 4, I 3
Ambler 244: R 42, S 2, C 1, L 1
Ambler 276: N 38, D 8
Ambler 265: T 42, M 4

6. The two groups of positions

For each allele, whether it differs from TEM-1 at any of the four extended-spectrum positions, and at any of the three inhibitor-resistance ones:

ESBL positions 104/164/238/240: broad-spectrum 0/12, extended-spectrum 12/12, inhibitor-resistant broad-spectrum 0/12, inhibitor-resistant extended-spectrum 10/10
inhibitor positions 69/244/276: broad-spectrum 0/12, extended-spectrum 0/12, inhibitor-resistant broad-spectrum 12/12, inhibitor-resistant extended-spectrum 7/10

All 22 alleles the catalog calls extended-spectrum carry a substitution at 104, 164, 238 or 240, and the 12 broad-spectrum and 12 inhibitor-resistant broad-spectrum alleles carry none. The three inhibitor-resistance positions split the other way: every inhibitor-resistant broad-spectrum allele is substituted at one of them, and no extended-spectrum allele is.

The same positions are bands over the alignment columns, which is what highlights draws. A band with no row takes 1-based alignment columns, so Ambler 238 is start: 236:

"highlights": [{ "start": 236, "end": 236, "label": "Ambler 238" }]
Eight bands over the 286 columns, seven red and the control grey, each labeled with its Ambler position.
Ambler 238 and 240 at residue resolution with every row read against TEM-1, so a dot is the same residue and a letter a different one. The boxed columns read S and K in olive and purple rows and a dot in every salmon and cyan row. Ambler 244, eleven columns to the right, reads C, L and S in cyan rows alone.

7. The control

For every variable position, how its residues split across the four phenotypes, scored by Cramer’s V over the residue-by-phenotype table:

position  residues                  Cramer's V
     104  36E,10K                         0.60
     164  29R,11S,6H                      0.57
     276  38N,8D                          0.48
      69  31M,8L,4V,3I                    0.47
      39  38Q,8K                          0.42
     240  41E,5K                          0.39
     238  40G,5S,1R                       0.37
     ...
     265  42T,4M                          0.22

The control is the lowest-scoring position whose minor residues at least three alleles carry, so the column has something to read:

control position: Ambler 265, T in TEM-1, minor residues in 4 alleles, Cramer's V 0.22
Ambler 265 at the same resolution as the figure above. The four alleles reading M are TEM-110, TEM-9, TEM-4 and TEM-68, whose phenotypes are broad-spectrum, extended-spectrum, extended-spectrum and inhibitor-resistant extended-spectrum.

Ambler 39 scores 0.42, between the two groups. Q39K is the one substitution separating TEM-2 from TEM-1, and both are broad-spectrum; the eight alleles carrying K here are one broad-spectrum, five extended-spectrum and two inhibitor-resistant extended-spectrum.

8. Check it against the raw data

The check reads the letters out of the unaligned NCBI sequences by the Ambler arithmetic alone and compares them with the row table the viewer loads:

letters = [unaligned[row][residue_of_ambler[p]] for p in fields]
from_table = [table[row][f'Ambler {p}'] for p in fields]
if letters != from_table:
    raise SystemExit(f'{row}: the FASTA reads {letters}, the row table {from_table}')
allele     104  164  238  240   69  244  276  265  phenotype
TEM-1        E    R    G    E    M    R    N    T  broad-spectrum
TEM-2        E    R    G    E    M    R    N    T  broad-spectrum
TEM-3        K    R    S    E    M    R    N    T  extended-spectrum
TEM-10       E    S    G    K    M    R    N    T  extended-spectrum
TEM-12       E    S    G    E    M    R    N    T  extended-spectrum
TEM-30       E    R    G    E    M    S    N    T  inhibitor-resistant broad-spectrum
TEM-32       E    R    G    E    I    R    N    T  inhibitor-resistant broad-spectrum
TEM-125      E    S    G    E    L    R    D    T  inhibitor-resistant extended-spectrum

TEM-3 reads E104K and G238S, TEM-10 reads R164S and E240K, TEM-12 reads R164S, TEM-30 reads R244S, TEM-32 reads M69I, and TEM-125 reads R164S with M69L and N276D. Those are the substitutions the primary literature gives for each allele.

9. Open the whole thing

The three hosted files, the tipLabel encoding, the nine strips and the eight bands in one view.

Reproduce it end to end

curl -O https://raw.githubusercontent.com/GMOD/JBrowseMSA/main/docs/tutorials/scripts/build_tem_alleles.sh
bash build_tem_alleles.sh

The script downloads the catalog and the reference proteins, selects the alleles, aligns them, infers the tree, writes the row table and prints every number on this page.

See also

References