CATH 3.40.50.620 · FunFam 145 · EC 6.3.1.14

Reading a protein family,
residue by residue

This is the functional annotation of diphthine—ammonia ligase (DPH6), the enzyme that puts the final amide onto diphthamide — a modified histidine on translation elongation factor 2. Most of its 47 database members carry no usable functional label. The pipeline below turns that raw membership list into a specific, testable claim: which residues form the active site, and one pocket residue that swaps identity only in archaea.

4741
members, after cleanup
11
catalytic-core residues
26
un-annotated, rescued
1
archaea-specific site

A · the pipeline

How the family was read

Five steps. Each row: what goes in → which tool → what comes out (the coloured strip is the real output) → what it tells us. Follow the rail top to bottom.

conserved (high scorecons) predicted pocket (P2Rank) functional core (both) variable / gap

What the family turns into

47→41
members after cleanup
11
catalytic-core residues
26
un-annotated, rescued
1
archaea-specific site (Q50H)

From 47 database IDs with mostly-missing labels, the pipeline ends at a specific, testable claim: 11 residues that are both evolutionarily conserved and sit in the predicted ATP pocket — the enzyme's likely active site — plus residue 50, which flips from Gln to His only in archaea.

B · deep dive on step 3

The 5 removed members: where the extra piece is, and why function is untouched

The overview says step 3 "dropped 5 outliers." This page opens that step: which 5, what the family's ATP-binding fingerprint says about them, exactly where their extra mass sits when superposed, and whether it can reach the active site.

1 · How the fingerprint was found (derived, not imported)

scorecons flags 15 columns at a perfect 1.0. Five are contiguous near the N-terminus (alignment columns 9–13). Their consensus across all 41 aligned sequences is G-G-K-D-S — identical in every member. With the conserved L6 and S8 just upstream, that block is the Walker A / PP-loop: the ATP α-phosphate-binding signature of the PP-loop ATPase family DPH6 belongs to. It is simply the most conserved stretch in the family's own alignment.

Frequency sequence logo, P-loop window (residues 5–14). Letter height = fraction of the 41 members carrying that residue; colour = amino-acid class (hydrophobic, polar, positive, negative, G/special). We plot frequency, not information-content (bits), so the report stays free of Shannon entropy. The G-G-K-D-S block is a full-height single letter — invariant across all 41.

2 · The P-loop test on all 5

MemberOrganismStruct.ExtraP-loopVerdict
D0A5F8Trypanosoma brucei gambiense DAL972274 aa+131 aa C-termISGGKDS ✓same enzyme · deleted entry
Q387L5Trypanosoma brucei brucei TREU927272 aa+129 aa C-termISGGKDS ✓same enzyme · EC 6.3.1.14
Q54Y82Dictyostelium discoideum264 aa+120 aa C-termISGGKDS ✓same enzyme · EC 6.3.1.14
M7WSX9Entamoeba histolytica HM-3:IMSS227 aa+90 aa C-termLSGGKDS ✓same enzyme · deleted entry
F6H602Vitis vinifera105 aanone ✗no P-loop; parses differently in CATH vs TED — flagged below

4 carry the fingerprint (same enzyme); only F6H602 (grape) lacks it — no P-loop anywhere in its 775-aa sequence. The same test that rescues 4 flags the 1 that needs a closer look.

3 · Where the extra mass sits — domain architecture, aligned on the shared core

core ends · 139Q966L4representative · chop 3–139ATP-hydrolase coreD0A5F8chop 3-276 · +131 aaATP-hydrolase coreC-terminal extensionQ387L5chop 3-274 · +129 aaATP-hydrolase coreC-terminal extensionQ54Y82chop 3-151_158-272 · +120 aaATP-hydrolase coreC-terminal extensionM7WSX9chop 4-230 · +90 aaATP-hydrolase coreC-terminal extension050100150200250residue number (same scale for every row)
ATP-hydrolase core (the FF145 domain, residues 3–139) C-terminal extension (a second domain, not part of FF145) P-loop (ATP-binding fingerprint) the 11 active-site residues

How to read it: every row is drawn on the same residue ruler (bottom axis), so the blue core sits in the same place in every row — that is what "aligned on the shared core" means. The representative (top) is chopped to the core and stops at the dashed line, residue 139. The 4 outliers carry an identical core (same blue block, same P-loop, same white active-site ticks) and then keep going: the hatched tail to the right of 139 is the extra 90–131 aa — a C-terminal extension (a second domain) that TED chopped into them but not into the representative.

Because all 11 active-site residues (up to residue 122) sit inside the shared blue core, to the left of the dashed line, the extra tail never touches the active site and cannot change the chemistry. It only dragged the global LDDT down to 0.50, since FoldMason was superposing a 137-aa core against these 227–274-aa two-part structures. Removing them cleans the per-column scoring; it does not exclude a different function. F6H602 is the true exception — no P-loop at all, and it does not resolve cleanly. It is flagged below.

⚠ Flagged for manual follow-up — F6H602 (grape)

This is the one member that fails the P-loop test. Three quick signals flagged it; I then compared its sequence and structure against the family directly — the P-loop scan, TM-align, and catalytic-residue projection below.

1 · databases annotate it as the enzyme — automatically
UniProt names it Diphthine—ammonia ligase, EC 6.3.1.14 (Rhea 19753); InterPro/Pfam carry the Diphthami_syn signatures; Gene3D lists 3.40.50.620 + 3.90.1490.10. But the entry is TrEMBL unreviewed, protein existence Predicted (PE 4), annotation score 2/5, and the reaction is tagged Automatic Annotation — homology transfer, no experiment.
2 · the catalytic fingerprint is missing
No P-loop (SGGKDS) anywhere in the full 775-aa sequence — the one motif all 45 other members carry. Searched GGKDS, ISGGKDS, and looser patterns: none.
3 · CATH and TED carve the region differently
CATH calls F6H602/63–191 one 3.40.50.620 domain. TED splits the same span into TED01 (3.30.160) + TED02 (1.20.5) — two domains, neither labelled 3.40.50.620. Disagreement on how to carve and name this region.
UniProt F6H602: Diphthine-ammonia ligase, TrEMBL unreviewed, protein existence Predicted, annotation score 2/5, Automatic Annotation
source: uniprot.org/uniprotkb/F6H602 — captured evidence

Signal 3, in full — CATH and TED annotate the same protein differently

CATH FunFam cathdb.info
Lists F6H602/63–191 as a member of FunFam 145, superfamily 3.40.50.620 — the N-terminal ATP-hydrolase family.
view the FF145 alignment →
TED ted.cathdb.info
Parses the same AlphaFold model into 6 structural domains — none is 3.40.50.620. CATH's 63–191 overlaps TED01 + TED02 (labelled 3.30.160 and 1.20.5). The diphthamide-related TED03 (3.90.1490.10) is a separate C-terminal domain that barely touches 63–191 (12 aa).
TED domainchoppingCATH superfamilyoverlap w/ 63–191
TED0112–1023.30.16040 aa ←
TED02125–1721.20.548 aa ←
TED03180–2843.90.1490.1012 aa
TED04337–4523.30.1330.400 aa
TED05455–6013.30.1330.400 aa
TED06631–663 / 675–7743.30.1330.400 aa
view the TED domains (API) →

Not a numbering artefact. F6H602 has no PDB structure (AlphaFold model), so both ranges are UniProt positions (1-based) — no PDB-vs-1-based offset. The gap is real: CATH treats 63–191 as one 3.40.50.620 domain; TED carves it into TED01 (3.30.160) + TED02 (1.20.5). A domain-parsing and labelling disagreement, not a coordinate convention. (Data from cathdb.info and the TED API; links above to verify.)

Structural check — I compared the folds (TM-align + P-loop scan)

To test the disagreement directly, I cut CATH's 63–191 region out of the AlphaFold model and structurally aligned it (TM-align) against the FF145 representative (Q966L4, a genuine 3.40.50.620) and against the C-terminal 3.90.1490.10 domain.

Cartoon superposition — the 3.40.50.620 representative carries all 11 DPH6 active-site residues; F6H602 63-191 keeps only the C-terminal hairpin and core, the P-loop and substrate pocket are absent
P-loop / Walker-A (binds ATP) substrate pocket C-terminal hairpin catalytic core site in the family, absent in F6H602

Both domains in the same orientation (TM-align superposition, PyMOL cartoon). Left: the representative carries all 11 DPH6 active-site residues, clustered in the ATP-and-substrate cleft. Right: F6H602 keeps only the C-terminal hairpin (cyan) and core (orange) in the same positions — the P-loop and substrate pocket (the 7 red markers) have no residue in F6H602 at all.

structural comparisonTM-scoreRMSDreading
genuine FF145 member vs representative (positive control)0.951.4 Åsame fold
F6H602 63–191 vs 3.40.50.620 representative0.452.3 Åpartial — about half a real member
F6H602 63–191 vs C-terminal 3.90.1490.10 (TED03)0.244.2 Åunrelated — not this fold

So 63–191 is a partial, degraded 3.40.50.620-like fold (TM 0.45 vs 0.95 for a genuine member; the resemblance sits in the 125–172 sub-piece, 0.4 Å over 48 residues) that has lost the catalytic P-loop — and it is not the C-terminal 3.90.1490.10 fold. Whether it is a divergent, catalytically dead DPH6 or a homology-driven mis-assignment cannot be settled here. (TM-align via tmtools; P-loop scanned over the full 775-aa sequence; CATH region and TED domains cut from the AlphaFold model.)

Catalytic-residue projection — does F6H602 keep the active site?

Final test: I mapped each of DPH6's 11 predicted active-site residues from the representative onto F6H602 through the TM-align superposition (nearest Cα in the common frame). Well-superposed positions land at 0.3–0.4 Å, so the register is trustworthy where it exists.

DPH6 active-site residues (rep)roleF6H602 equivalentresult
L6, S8, D12, S13Walker-A / P-loop — binds ATPno structural counterpart
M48, Y49, Q50substrate pocket — Q50 does the ammonia chemistryno structural counterpart
G115, A116, Q122C-terminal hairpin — substrate bindingG148, A149, Q155identical (0.3–0.4 Å)
L95catalytic coreM128 (0.4 Å)aligned, substituted

The catalytic machinery is gone. The entire ATP-binding P-loop and the substrate/ammonia pocket — 7 of the 11 active-site residues — have no structural counterpart in F6H602; only a peripheral C-terminal hairpin survives (3 identical residues, G/A/Q). A protein missing the whole ATP site and substrate pocket cannot run the amidation reaction. (Active-site residues projected via the TM-align superposition, nearest-Cα within 4 Å; matches shown land at 0.3–0.4 Å.)

Verdict — most likely a homology-annotation artefact, not a functional DPH6. The enzyme label is homology-only and unverified (automatic, PE 4, score 2/5); F6H602 has lost the entire ATP-binding P-loop and substrate pocket — 7 of 11 active-site residues have no structural counterpart, and the P-loop is absent in sequence too — keeping only a peripheral C-terminal hairpin. The 63–191 region is a partial 3.40.50.620-like fold (TM 0.45) that CATH and TED parse and label incompatibly. On this evidence it cannot perform the amidation chemistry, so it is excluded from the residue-level analysis; a genuinely divergent but active DPH6 is very unlikely. Confirming pseudogene-vs-mis-annotation would need curated (not automatic) review.

C · the payoff

The active site, on the structure

The 11 green residues are conserved across every member and inside the P2Rank-predicted pocket. On the representative structure (Q966L4, C. elegans, pLDDT 95.6) they cluster into one face — the ATP-and-substrate cleft. Residue 33 (amber) sits just under the conservation threshold. Drag to rotate; the buttons recolour by evidence.

Interactive structure · 3Dmol.js

The 11+1 residues · scorecons · role

ResScoreRegionLikely role
61.000Walker A / P-loopbinds ATP α-phosphate
80.967Walker A / P-loopbinds ATP
121.000Walker A / P-loopGKDS motif — ATP
131.000Walker A / P-loopGKDS motif — ATP
480.959Substrate pocketpositions diphthine
490.843Substrate pocketpositions diphthine
500.789Substrate pocketammonia chemistry — Gln→His in archaea
950.897Catalytic coreadjacent to conserved E97 base
1151.000C-terminal hairpinsubstrate binding (YQ motif)
1161.000C-terminal hairpinsubstrate binding (YQ motif)
1221.000C-terminal hairpinsubstrate binding (GA motif)
330.726Pocket loopthreshold-borderline (0.726)

No experimental active site exists for EC 6.3.1.14 — M-CSA has no entry, and UniProt lists no binding sites for Q966L4. These 11 residues are therefore the first residue-level active-site hypothesis for the family, built purely from conservation ∩ pocket geometry.

D · can we trust the labels?

Turning 26 blank entries into function

Of the 41 aligned members, only 15 carry any annotation; 26 are dead entries or blank. The question is whether the blank ones are the same enzyme. The test: do they keep the 11 catalytic-core residues? Read the two panels left→right — status, then the conservation check that resolves it.

Step 1 · what UniProt says about the 41

Annotated (EC/name)
15
Dead UniProt entry
14
Live, no annotation
12

26 of 41 members (63%) arrive with no usable function. Annotation alone would leave most of the family dark.

Step 2 · how many of the 11 core residues each keeps

11/11 core residues identical
17
10/11 core residues identical
15
9/11 core residues identical
8
8/11 core residues identical
1

Every one of the 41 keeps at least 8 of the 11 core residues and matches the CATH domain. Because the catalytic core is intact, the blank entries are the same enzyme: 40 high-confidence, 1 conserved-variant. Conservation rescues what annotation left blank.

E · honest labelling

How sure is each label?

A predicted function is not a proven one. Every member (all 47) carries an evidence tier on a four-step scale — Swiss-Prot reviewed → TrEMBL auto-annotated → predicted from structure → cannot confirm. This FunFam has no Swiss-Prot entry, so it fills the lower three. Each card below shows what puts a member in that tier and a real example from the family.

TrEMBL auto-annotated17 members
UniProt gives it an EC number, but from automatic pipeline annotation — not experimentally verified.
example
Q387L5
Diphthine--ammonia ligase · EC 6.3.1.14
Trypanosoma brucei brucei (strain 927/4 GUTat10.1)
UniProt lists EC 6.3.1.14, but the source is the automatic pipeline — no experiment behind it.
Predicted from structure29 members
UniProt gives no functional annotation at all; the function is inferred here from structure + conservation.
example
Q7LXN2
(no protein name)
Saccharolobus solfataricus
No UniProt annotation; the function is assigned here from the intact catalytic core. Keeps 10/11 catalytic-core residues.
Cannot confirm as FF1451 member
Function can't be inferred — the catalytic P-loop motif is missing and CATH and TED parse the domain incompatibly.
example
F6H602
Diphthine--ammonia ligase
Vitis vinifera
P-loop absent across the whole 775-aa protein; only a peripheral hairpin survives — fully dissected in section B.

The one cannot-confirm case (F6H602, grape) is instructive: its FF145 label is a homology-only automatic annotation (TrEMBL, protein-existence "Predicted", score 2/5), the catalytic P-loop is absent from the whole 775-aa protein, and its 63–191 region only partly resembles the 3.40.50.620 fold (TM 0.45 vs 0.95 for a genuine member) with CATH and TED carving it into different domains. It is flagged for manual follow-up in section B, not silently counted as a member.

F · the one thing homology would miss

Residue 50: Gln in eukaryotes, His in archaea

Position 50 sits in the substrate pocket (scorecons 0.789) — conserved enough to matter, variable enough to split. The split falls exactly on kingdom, and exactly where the ammonia chemistry happens.

Gln
25 eukaryotes — all Gln (human, yeast, worm, slime mould…).
+ 3 archaea, all Saccharolobus solfataricus strains.
His
13 archaea — every one His (Methanocaldococcus, Methanosarcina, Thermococcus, Pyrobaculum).
No eukaryote carries His here.

Why it could matter — the ammonia chemistry

Diphthineon eEF2–CH₂–COO⁻Diphthine–AMPreactive intermediate–C(=O)–O–AMPDiphthamidethe product–C(=O)–NH₂+ ATP– PPiadenylation · ATP cut at α-P+ NH₃ (from NH₄⁺)– AMPamidationPosition 50 sits at this stepGln (eukaryotes) / His (archaea)may help NH₄⁺ → NH₃ at the moment of attack

The catch, at the second arrow: only free NH₃ can attack the intermediate, but at cellular pH almost all ammonia sits as NH₄⁺ (the NH₄⁺/NH₃ equilibrium is near pH 9.2). So the enzyme has to strip a proton — NH₄⁺ → NH₃ — right where position 50 sits. That is exactly where Gln and His behave differently:

Gln · eukaryotes

a Gln side chain is a neutral amide. It can hydrogen-bond to orient the ammonia and the intermediate, but it cannot gain or lose a proton, so it cannot itself pull the proton off NH4+. Eukaryotic DPH6 must therefore rely on something else (water, phosphate, or another residue) for that proton transfer — and it clearly works, since human and yeast DPH6 are fully active with Gln.

His · archaea

a His imidazole has a pKa near 6, so at cellular pH it can exist in both protonated and neutral forms and shuttle a proton. A His at this substrate-facing position could act directly as the base that converts NH4+ to the attacking NH3, or stabilise the negatively-charged transition state as it forms. That is a catalytic role glutamine cannot play.

Same reaction, two ways — the substrate and product never change. this is not a change of substrate or of product — every member makes diphthamide from diphthine. It is the same reaction reached two ways. Eukaryotes appear to run the amidation with a non-ionisable Gln at the key pocket position; most archaea replace it with an ionisable His that could take a direct hand in the proton chemistry of ammonia. Whether the archaeal enzyme is actually faster, or simply tolerant of a different residue, is untested (see literature) — but the split is at exactly the position where the ammonia chemistry happens, which is why it is worth flagging.

Literature check — is this a real split, or an annotation artefact?

The sceptic's worry: maybe the archaeal sequences were misfiled into this family — then "Gln vs His" would just be two different proteins, not a real signal. Two papers rule that out:

PMID 23013770 de Crécy-Lagard et al., Biol Direct 7:32 (2012)

DPH6 (the DUF71/COG2102 family, a PP-loop ATPase) orthologues span nearly all sequenced Archaea and Eukaryotes — a distribution the authors read as origin in the two lineages' common ancestor. Either way, the archaeal His-carrying sequences are the SAME enzyme as the eukaryotic ones, not a different one.

PMID 23971743 Su, Lin & Lin, Crit Rev Biochem Mol Biol 48:515-521 (2013)

Diphthamide itself is a modified histidine on elongation factor 2, conserved across eukaryotes AND archaea — confirming both domains of life run this pathway.

Literature confirms (1) archaea and eukaryotes share the same DPH6 enzyme (orthologues across both domains of life), and (2) its active site has never been experimentally studied. The Q50H split is therefore a genuinely novel, untested observation — exactly the kind of homology-blind detail that PMID 23013770 warns pure-homology annotation misses.

Both papers were selected from the family's 115-paper literature pool (EuropePMC keyword search, 93 papers · UniProt-cited references, 22 papers), the same auditable pool behind every member's evidence in Methods.

Other kingdom-split pocket positions

res 47 · F→Y · sc 0.764res 33 · L→M · sc 0.726 · in pocketres 45 · D→E · sc 0.709

Three more positions split along the same eukaryote/archaea line, all in or near the pocket. Position 50 is the sharpest and sits closest to the ammonia step.

Status — a hypothesis, not an established mechanism. Observed pattern from the alignment. No paper found describes a distinct DPH6/amidation mechanism in archaea, nor this specific position. No paper reports a crystal structure, catalytic-residue assignment, or any experimental study of the DPH6 active site. Residue 50 has never been experimentally characterised in any organism.

G · what these actually are

One enzyme, 46 times — and one impostor

The family barely diversifies: 46 of the 47 members are the same enzyme. So instead of 46 near-identical entries, here is the one shared description — then the single exception, F6H602, whose real function is something else.

The shared function — all 46 members

Diphthine—ammonia ligase (DPH6), EC 6.3.1.14. Every one of the 46 catalyses the final, ATP-dependent amidation step of diphthamide biosynthesis on translation elongation factor 2 (eEF2): diphthine + NH₄⁺ + ATP → diphthamide + AMP + diphosphate PMID 23169644, 23468660. The family is the DUF71/COG2102 group, a PP-loop ATPase that spans Archaea and Eukaryotes PMID 23013770.

They are one enzyme, not a mixed bag, on the structural evidence: every scored member keeps the catalytic core intact — 32 of 41 keep 10–11 of the 11 core residues, all 41 keep ≥ 8, and each matches the CATH 3.40.50.620 fold. (The remaining 5 are inactive UniProt entries with no sequence left to score.)

How the label is earned splits the 46 in two: 17 carry EC 6.3.1.14 as a TrEMBL annotation; the other 29 have no UniProt functional annotation and are assigned from structure + conservation alone — the half of the family pure annotation would leave dark. The only within-function variation worth flagging is residue 50, Gln in eukaryotes vs His in archaea (section F).

The 47th — F6H602's likely real function (a hypothesis)

F6H602 (grape, Vitis vinifera) is annotated "diphthine—ammonia ligase" by homology, but it is not a functional DPH6. Current UniProt/InterPro shows a 775-aa multi-domain fusion: a DPH6-fold module (Gene3D 3.40.50.620 + 3.90.1490.10; InterPro DPH6/MJ0570 and Diphthamide-synthase signatures) fused to two RutC-like / YjgF-UK114 domains (Pfam PF01042, the RidA family).

The DPH6 module is catalytically dead: the family's PP-loop motif SGGKDS — the ATP-α-phosphate fingerprint present in every real member — is absent from the entire 775-aa sequence. The fold survives; the active site does not. So the "ligase" name is a fold-homology overreach.

Best guess at what it really does: its only catalytically-supported domains are the two RidA/YjgF deaminases, which elsewhere hydrolyse reactive enamine/imine metabolites (e.g. 2-aminoacrylate) to protect metabolism. So F6H602 most likely acts as a RidA-family reactive-intermediate deaminase that happens to carry a degenerate DPH6-like domain — not a diphthamide synthase. Caveat: it is a protein-existence-4 "Predicted" grape gene model, so a chimeric mis-prediction (two genes joined by a faulty model) cannot be excluded. Structural detail for this call is in section B.

H · dig in

Methods & data

Tools. TED (ted.cathdb.info) · AlphaFold DB model_v4 · FoldMason easy-msa (Science 391:485-488, 2026) · scorecons valdar01 / PET91mod on host ganon · P2Rank 2.5.1 (Java 17). Representative Q966L4, TED chopping 3-139, pLDDT 95.6.
Numbers. 41 aligned members · 145 columns · DOPS 94.4 · 38 columns ≥ 0.7 · pocket ligand-binding probability 0.431 (P2Rank pocket 1). All conservation from real scorecons; no Shannon-entropy proxy anywhere in this report.
Enzyme. diphthine—ammonia ligase / DPH6, EC 6.3.1.14, Rhea 19753. diphthine + NH₄⁺ + ATP → diphthamide + AMP + diphosphate.