This is the functional annotation of diphthine—ammonia ligase (DPH6), the enzyme that puts the final amide onto diphthamide — a modified histidine on translation elongation factor 2. Most of its 47 database members carry no usable functional label. The pipeline below turns that raw membership list into a specific, testable claim: which residues form the active site, and one pocket residue that swaps identity only in archaea.
A · the pipeline
Five steps. Each row: what goes in → which tool → what comes out (the coloured strip is the real output) → what it tells us. Follow the rail top to bottom.
From 47 database IDs with mostly-missing labels, the pipeline ends at a specific, testable claim: 11 residues that are both evolutionarily conserved and sit in the predicted ATP pocket — the enzyme's likely active site — plus residue 50, which flips from Gln to His only in archaea.
B · deep dive on step 3
The overview says step 3 "dropped 5 outliers." This page opens that step: which 5, what the family's ATP-binding fingerprint says about them, exactly where their extra mass sits when superposed, and whether it can reach the active site.
1 · How the fingerprint was found (derived, not imported)
scorecons flags 15 columns at a perfect 1.0. Five are contiguous near the N-terminus (alignment columns 9–13). Their consensus across all 41 aligned sequences is G-G-K-D-S — identical in every member. With the conserved L6 and S8 just upstream, that block is the Walker A / PP-loop: the ATP α-phosphate-binding signature of the PP-loop ATPase family DPH6 belongs to. It is simply the most conserved stretch in the family's own alignment.
Frequency sequence logo, P-loop window (residues 5–14). Letter height = fraction of the 41 members carrying that residue; colour = amino-acid class (hydrophobic, polar, positive, negative, G/special). We plot frequency, not information-content (bits), so the report stays free of Shannon entropy. The G-G-K-D-S block is a full-height single letter — invariant across all 41.
2 · The P-loop test on all 5
| Member | Organism | Struct. | Extra | P-loop | Verdict |
|---|---|---|---|---|---|
| D0A5F8 | Trypanosoma brucei gambiense DAL972 | 274 aa | +131 aa C-term | ISGGKDS ✓ | same enzyme · deleted entry |
| Q387L5 | Trypanosoma brucei brucei TREU927 | 272 aa | +129 aa C-term | ISGGKDS ✓ | same enzyme · EC 6.3.1.14 |
| Q54Y82 | Dictyostelium discoideum | 264 aa | +120 aa C-term | ISGGKDS ✓ | same enzyme · EC 6.3.1.14 |
| M7WSX9 | Entamoeba histolytica HM-3:IMSS | 227 aa | +90 aa C-term | LSGGKDS ✓ | same enzyme · deleted entry |
| F6H602 | Vitis vinifera | 105 aa | — | none ✗ | no P-loop; parses differently in CATH vs TED — flagged below |
4 carry the fingerprint (same enzyme); only F6H602 (grape) lacks it — no P-loop anywhere in its 775-aa sequence. The same test that rescues 4 flags the 1 that needs a closer look.
3 · Where the extra mass sits — domain architecture, aligned on the shared core
How to read it: every row is drawn on the same residue ruler (bottom axis), so the blue core sits in the same place in every row — that is what "aligned on the shared core" means. The representative (top) is chopped to the core and stops at the dashed line, residue 139. The 4 outliers carry an identical core (same blue block, same P-loop, same white active-site ticks) and then keep going: the hatched tail to the right of 139 is the extra 90–131 aa — a C-terminal extension (a second domain) that TED chopped into them but not into the representative.
Because all 11 active-site residues (up to residue 122) sit inside the shared blue core, to the left of the dashed line, the extra tail never touches the active site and cannot change the chemistry. It only dragged the global LDDT down to 0.50, since FoldMason was superposing a 137-aa core against these 227–274-aa two-part structures. Removing them cleans the per-column scoring; it does not exclude a different function. F6H602 is the true exception — no P-loop at all, and it does not resolve cleanly. It is flagged below.
This is the one member that fails the P-loop test. Three quick signals flagged it; I then compared its sequence and structure against the family directly — the P-loop scan, TM-align, and catalytic-residue projection below.
Signal 3, in full — CATH and TED annotate the same protein differently
| TED domain | chopping | CATH superfamily | overlap w/ 63–191 |
|---|---|---|---|
| TED01 | 12–102 | 3.30.160 | 40 aa ← |
| TED02 | 125–172 | 1.20.5 | 48 aa ← |
| TED03 | 180–284 | 3.90.1490.10 | 12 aa |
| TED04 | 337–452 | 3.30.1330.40 | 0 aa |
| TED05 | 455–601 | 3.30.1330.40 | 0 aa |
| TED06 | 631–663 / 675–774 | 3.30.1330.40 | 0 aa |
Not a numbering artefact. F6H602 has no PDB structure (AlphaFold model), so both ranges are UniProt positions (1-based) — no PDB-vs-1-based offset. The gap is real: CATH treats 63–191 as one 3.40.50.620 domain; TED carves it into TED01 (3.30.160) + TED02 (1.20.5). A domain-parsing and labelling disagreement, not a coordinate convention. (Data from cathdb.info and the TED API; links above to verify.)
Structural check — I compared the folds (TM-align + P-loop scan)
To test the disagreement directly, I cut CATH's 63–191 region out of the AlphaFold model and structurally aligned it (TM-align) against the FF145 representative (Q966L4, a genuine 3.40.50.620) and against the C-terminal 3.90.1490.10 domain.
Both domains in the same orientation (TM-align superposition, PyMOL cartoon). Left: the representative carries all 11 DPH6 active-site residues, clustered in the ATP-and-substrate cleft. Right: F6H602 keeps only the C-terminal hairpin (cyan) and core (orange) in the same positions — the P-loop and substrate pocket (the 7 red markers) have no residue in F6H602 at all.
| structural comparison | TM-score | RMSD | reading |
|---|---|---|---|
| genuine FF145 member vs representative (positive control) | 0.95 | 1.4 Å | same fold |
| F6H602 63–191 vs 3.40.50.620 representative | 0.45 | 2.3 Å | partial — about half a real member |
| F6H602 63–191 vs C-terminal 3.90.1490.10 (TED03) | 0.24 | 4.2 Å | unrelated — not this fold |
So 63–191 is a partial, degraded 3.40.50.620-like fold (TM 0.45 vs 0.95 for a genuine member; the resemblance sits in the 125–172 sub-piece, 0.4 Å over 48 residues) that has lost the catalytic P-loop — and it is not the C-terminal 3.90.1490.10 fold. Whether it is a divergent, catalytically dead DPH6 or a homology-driven mis-assignment cannot be settled here. (TM-align via tmtools; P-loop scanned over the full 775-aa sequence; CATH region and TED domains cut from the AlphaFold model.)
Catalytic-residue projection — does F6H602 keep the active site?
Final test: I mapped each of DPH6's 11 predicted active-site residues from the representative onto F6H602 through the TM-align superposition (nearest Cα in the common frame). Well-superposed positions land at 0.3–0.4 Å, so the register is trustworthy where it exists.
| DPH6 active-site residues (rep) | role | F6H602 equivalent | result |
|---|---|---|---|
| L6, S8, D12, S13 | Walker-A / P-loop — binds ATP | — | no structural counterpart |
| M48, Y49, Q50 | substrate pocket — Q50 does the ammonia chemistry | — | no structural counterpart |
| G115, A116, Q122 | C-terminal hairpin — substrate binding | G148, A149, Q155 | identical (0.3–0.4 Å) |
| L95 | catalytic core | M128 (0.4 Å) | aligned, substituted |
The catalytic machinery is gone. The entire ATP-binding P-loop and the substrate/ammonia pocket — 7 of the 11 active-site residues — have no structural counterpart in F6H602; only a peripheral C-terminal hairpin survives (3 identical residues, G/A/Q). A protein missing the whole ATP site and substrate pocket cannot run the amidation reaction. (Active-site residues projected via the TM-align superposition, nearest-Cα within 4 Å; matches shown land at 0.3–0.4 Å.)
Verdict — most likely a homology-annotation artefact, not a functional DPH6. The enzyme label is homology-only and unverified (automatic, PE 4, score 2/5); F6H602 has lost the entire ATP-binding P-loop and substrate pocket — 7 of 11 active-site residues have no structural counterpart, and the P-loop is absent in sequence too — keeping only a peripheral C-terminal hairpin. The 63–191 region is a partial 3.40.50.620-like fold (TM 0.45) that CATH and TED parse and label incompatibly. On this evidence it cannot perform the amidation chemistry, so it is excluded from the residue-level analysis; a genuinely divergent but active DPH6 is very unlikely. Confirming pseudogene-vs-mis-annotation would need curated (not automatic) review.
C · the payoff
The 11 green residues are conserved across every member and inside the P2Rank-predicted pocket. On the representative structure (Q966L4, C. elegans, pLDDT 95.6) they cluster into one face — the ATP-and-substrate cleft. Residue 33 (amber) sits just under the conservation threshold. Drag to rotate; the buttons recolour by evidence.
Interactive structure · 3Dmol.js
The 11+1 residues · scorecons · role
| Res | Score | Region | Likely role |
|---|---|---|---|
| 6 | 1.000 | Walker A / P-loop | binds ATP α-phosphate |
| 8 | 0.967 | Walker A / P-loop | binds ATP |
| 12 | 1.000 | Walker A / P-loop | GKDS motif — ATP |
| 13 | 1.000 | Walker A / P-loop | GKDS motif — ATP |
| 48 | 0.959 | Substrate pocket | positions diphthine |
| 49 | 0.843 | Substrate pocket | positions diphthine |
| 50 | 0.789 | Substrate pocket | ammonia chemistry — Gln→His in archaea |
| 95 | 0.897 | Catalytic core | adjacent to conserved E97 base |
| 115 | 1.000 | C-terminal hairpin | substrate binding (YQ motif) |
| 116 | 1.000 | C-terminal hairpin | substrate binding (YQ motif) |
| 122 | 1.000 | C-terminal hairpin | substrate binding (GA motif) |
| 33 | 0.726 | Pocket loop | threshold-borderline (0.726) |
No experimental active site exists for EC 6.3.1.14 — M-CSA has no entry, and UniProt lists no binding sites for Q966L4. These 11 residues are therefore the first residue-level active-site hypothesis for the family, built purely from conservation ∩ pocket geometry.
D · can we trust the labels?
Of the 41 aligned members, only 15 carry any annotation; 26 are dead entries or blank. The question is whether the blank ones are the same enzyme. The test: do they keep the 11 catalytic-core residues? Read the two panels left→right — status, then the conservation check that resolves it.
Step 1 · what UniProt says about the 41
26 of 41 members (63%) arrive with no usable function. Annotation alone would leave most of the family dark.
Step 2 · how many of the 11 core residues each keeps
Every one of the 41 keeps at least 8 of the 11 core residues and matches the CATH domain. Because the catalytic core is intact, the blank entries are the same enzyme: 40 high-confidence, 1 conserved-variant. Conservation rescues what annotation left blank.
E · honest labelling
A predicted function is not a proven one. Every member (all 47) carries an evidence tier on a four-step scale — Swiss-Prot reviewed → TrEMBL auto-annotated → predicted from structure → cannot confirm. This FunFam has no Swiss-Prot entry, so it fills the lower three. Each card below shows what puts a member in that tier and a real example from the family.
The one cannot-confirm case (F6H602, grape) is instructive: its FF145 label is a homology-only automatic annotation (TrEMBL, protein-existence "Predicted", score 2/5), the catalytic P-loop is absent from the whole 775-aa protein, and its 63–191 region only partly resembles the 3.40.50.620 fold (TM 0.45 vs 0.95 for a genuine member) with CATH and TED carving it into different domains. It is flagged for manual follow-up in section B, not silently counted as a member.
F · the one thing homology would miss
Position 50 sits in the substrate pocket (scorecons 0.789) — conserved enough to matter, variable enough to split. The split falls exactly on kingdom, and exactly where the ammonia chemistry happens.
Why it could matter — the ammonia chemistry
The catch, at the second arrow: only free NH₃ can attack the intermediate, but at cellular pH almost all ammonia sits as NH₄⁺ (the NH₄⁺/NH₃ equilibrium is near pH 9.2). So the enzyme has to strip a proton — NH₄⁺ → NH₃ — right where position 50 sits. That is exactly where Gln and His behave differently:
a Gln side chain is a neutral amide. It can hydrogen-bond to orient the ammonia and the intermediate, but it cannot gain or lose a proton, so it cannot itself pull the proton off NH4+. Eukaryotic DPH6 must therefore rely on something else (water, phosphate, or another residue) for that proton transfer — and it clearly works, since human and yeast DPH6 are fully active with Gln.
a His imidazole has a pKa near 6, so at cellular pH it can exist in both protonated and neutral forms and shuttle a proton. A His at this substrate-facing position could act directly as the base that converts NH4+ to the attacking NH3, or stabilise the negatively-charged transition state as it forms. That is a catalytic role glutamine cannot play.
Literature check — is this a real split, or an annotation artefact?
The sceptic's worry: maybe the archaeal sequences were misfiled into this family — then "Gln vs His" would just be two different proteins, not a real signal. Two papers rule that out:
DPH6 (the DUF71/COG2102 family, a PP-loop ATPase) orthologues span nearly all sequenced Archaea and Eukaryotes — a distribution the authors read as origin in the two lineages' common ancestor. Either way, the archaeal His-carrying sequences are the SAME enzyme as the eukaryotic ones, not a different one.
Diphthamide itself is a modified histidine on elongation factor 2, conserved across eukaryotes AND archaea — confirming both domains of life run this pathway.
Literature confirms (1) archaea and eukaryotes share the same DPH6 enzyme (orthologues across both domains of life), and (2) its active site has never been experimentally studied. The Q50H split is therefore a genuinely novel, untested observation — exactly the kind of homology-blind detail that PMID 23013770 warns pure-homology annotation misses.
Both papers were selected from the family's 115-paper literature pool (EuropePMC keyword search, 93 papers · UniProt-cited references, 22 papers), the same auditable pool behind every member's evidence in Methods.
Other kingdom-split pocket positions
Three more positions split along the same eukaryote/archaea line, all in or near the pocket. Position 50 is the sharpest and sits closest to the ammonia step.
G · what these actually are
The family barely diversifies: 46 of the 47 members are the same enzyme. So instead of 46 near-identical entries, here is the one shared description — then the single exception, F6H602, whose real function is something else.
The shared function — all 46 members
Diphthine—ammonia ligase (DPH6), EC 6.3.1.14. Every one of the 46 catalyses the final, ATP-dependent amidation step of diphthamide biosynthesis on translation elongation factor 2 (eEF2): diphthine + NH₄⁺ + ATP → diphthamide + AMP + diphosphate PMID 23169644, 23468660. The family is the DUF71/COG2102 group, a PP-loop ATPase that spans Archaea and Eukaryotes PMID 23013770.
They are one enzyme, not a mixed bag, on the structural evidence: every scored member keeps the catalytic core intact — 32 of 41 keep 10–11 of the 11 core residues, all 41 keep ≥ 8, and each matches the CATH 3.40.50.620 fold. (The remaining 5 are inactive UniProt entries with no sequence left to score.)
How the label is earned splits the 46 in two: 17 carry EC 6.3.1.14 as a TrEMBL annotation; the other 29 have no UniProt functional annotation and are assigned from structure + conservation alone — the half of the family pure annotation would leave dark. The only within-function variation worth flagging is residue 50, Gln in eukaryotes vs His in archaea (section F).
The 47th — F6H602's likely real function (a hypothesis)
F6H602 (grape, Vitis vinifera) is annotated "diphthine—ammonia ligase" by homology, but it is not a functional DPH6. Current UniProt/InterPro shows a 775-aa multi-domain fusion: a DPH6-fold module (Gene3D 3.40.50.620 + 3.90.1490.10; InterPro DPH6/MJ0570 and Diphthamide-synthase signatures) fused to two RutC-like / YjgF-UK114 domains (Pfam PF01042, the RidA family).
The DPH6 module is catalytically dead: the family's PP-loop motif SGGKDS — the ATP-α-phosphate fingerprint present in every real member — is absent from the entire 775-aa sequence. The fold survives; the active site does not. So the "ligase" name is a fold-homology overreach.
Best guess at what it really does: its only catalytically-supported domains are the two RidA/YjgF deaminases, which elsewhere hydrolyse reactive enamine/imine metabolites (e.g. 2-aminoacrylate) to protect metabolism. So F6H602 most likely acts as a RidA-family reactive-intermediate deaminase that happens to carry a degenerate DPH6-like domain — not a diphthamide synthase. Caveat: it is a protein-existence-4 "Predicted" grape gene model, so a chimeric mis-prediction (two genes joined by a faulty model) cannot be excluded. Structural detail for this call is in section B.
H · dig in