CYP FUNCTIONAL TRIAGE · superfamily 1.10.630.10 · evidence rebuild

118 CYP candidates —
one function, or several?

Every claim below is re-derived from the primary data in this repository and shown as a figure: the alignment, the motif columns, the clustering, the structures, the substrate pocket, and the reactions with their UniProt/EC/PMID sources. Where the evidence contradicted the previous version of this report, the correction is stated explicitly rather than quietly fixed.

118
sequences
7
CYP families (given)
5
functional groups
0.992
clustering purity vs families
2
experimental structures
A · the input

Where the family names come from

Question: Where do the family names come from — the FASTA files, UniProt, or this analysis?
method
  • The 118 arrive already split into 7 per-family FASTA files (FF160_data/sequence/).
  • 720 / 725 / 947 / 90 / 88 / 974 / 723 = 67 / 31 / 8 / 5 / 4 / 2 / 1 = 118, matching ff160.fasta exactly.
  • The family label of a sequence = which file it came from. No family is guessed from a name.
CYP720 (group B)67CYP725 (group A)31CYP947 (group A)8CYP90 (group D)5CYP88 (group C)4CYP974 (group E)2CYP723 (group A)1
Fig A1. The 7 given families and their sizes (n=118). Bar length = share of the set. Colour = family, used consistently in every later figure.
B · the shared catalytic core

What "they are all P450s" actually means, residue by residue

Question: What exactly is the shared catalytic core — which motifs, located how, and how conserved is each one across the 118?
method
  • Locate 3 canonical P450 motifs, then map each back to its MSA column and read that column across all 118.
  • Haem loop: regex on the ungapped sequence, restricted to the C-terminal 40% where the motif belongs — matches 118/118.
  • I-helix: same approach, but the regex is strict and matches only 82/118; the alignment column is the evidence, and it carries Thr/Ser in 117/118.
  • ExxR: located by column scan (E at column c, R at c+3) — a raw-sequence regex also hits spurious ExxR elsewhere.
  • Every motif checked against the crystal structure 8X3E.
  • Conservation = scorecons.
heme-binding loop — FxxGxRxCxGMISL648P649F650G651GARQ652G653ATLG654R655LFIM656C657PVA658G659AWMN660EHD661col
Fig B1. Sequence logo over the heme-binding loop — FxxGxRxCxG, built from all 118 aligned sequences. Letter height = frequency of that residue in that column (non-gap rows); colour = residue chemistry. Columns 648–661 of the alignment; the shaded columns are the motif positions themselves, with scorecons: F650 (1.000), G653 (1.000), R655 (0.958), C657 (1.000), G659 (0.979).
K-helix salt bridge — ExxRIVA569DQNS570E571TGS572L573R574LMIV575GFAV576col
Fig B2. Sequence logo over the K-helix salt bridge — ExxR, built from all 118 aligned sequences. Letter height = frequency of that residue in that column (non-gap rows); colour = residue chemistry. Columns 569–576 of the alignment; the shaded columns are the motif positions themselves, with scorecons: E571 (0.990), R574 (1.000).
I-helix acid/alcohol pair — (A/G)Gx(E/D)TFHN485AG486GS487YNHF488ED489TS490TSV491ASVT492SKRT493col
Fig B3. Sequence logo over the I-helix acid/alcohol pair — (A/G)Gx(E/D)T, built from all 118 aligned sequences. Letter height = frequency of that residue in that column (non-gap rows); colour = residue chemistry. Columns 485–493 of the alignment; the shaded columns are the motif positions themselves, with scorecons: A486 (0.866), G487 (0.730), E489 (0.791), T490 (0.876).

The numbers, per motif

motifwhat it doesMSA columnresidues in that column (of 118)scorecons
haem loop FxxGxRxCxGits invariant Cys is the thiolate that ligates the haem iron — the residue that makes a P450 a P450657Cys 118/1181.000
K-helix ExxRGlu and Arg form an ion pair that locks the core fold — structural, not catalytic, hence two residues571 / 574E 117/118 (1 Lys) · R 118/1180.990 / 1.000
I-helix acid/alcoholthe Thr/Ser that delivers protons to split O₂490Thr 106 · Ser 11 · Asn 10.876

Verified in the crystal structure, not just in the alignment

The construct in 8X1W/8X3E starts at residue 43, so structure numbering = full-sequence numbering − 32. Both numbers are given here; the previous version quoted structure numbers only, which do not map onto the sequences without this offset.

rolesequence #structure #structural evidence
haem axial thiolate445413SG-Fe distance 2.69 A in 8X3E (coordination bond)
I-helix acid/alcohol Thr (O2 activation)310278THR at structure residue 278; alignment column 490 is Thr 106 / Ser 11 / Asn 1 of 118
K-helix ExxR salt bridge366/369334/337GLU334 + ARG337 in 8X3E; alignment columns 571/574 are E 117/118 and R 118/118
Fig B4. scorecons for every one of the 741 alignment columns.
jump to a motif residue:
scorecons occupancy (non-gap share) motif residue hover = read a column · drag = zoom · double-click = reset
How to read it. Each vertical bar is one alignment column; height = its scorecons score (1.0 = every one of the 118 sequences agrees). Hover any bar to see the exact score and which residues actually occur there, with counts. Red marks the four key residues of the three motifs: the haem-ligating Cys (1.000), the two halves of the ExxR salt bridge — Glu (0.990) and Arg (1.000), which pair with each other and so are two columns, not one — and the I-helix Thr (0.876). Across the whole alignment 133 columns score ≥0.7 and 42 score ≥0.9 (mean 0.388).
conclusion
All 118 carry the complete P450 catalytic machinery: haem Cys 118/118, ExxR 117/118 (one sequence has Lys for Glu), I-helix Thr/Ser 117/118. All 118 therefore run the same O₂-activation chemistry, and the catalytic core does not differentiate them. The following sections examine the alignment as a whole, then the substrate pocket.
C · sequence evidence

Hide the labels, and the sequences sort themselves back into the families

Question: What does the clustering look like, and does it agree with the family assignment?
method
  • Compare every sequence with every other one — all 6,903 pairs — and record their % identity from the MAFFT alignment.
  • That single table of pairwise identities is everything section C uses: summarised family-by-family (Fig C1), built into a tree (Fig C2), and as a distribution (Fig C3).
  • Building the tree (Fig C2): start with each sequence alone, repeatedly merge the two closest clumps, and record the identity at which each merge happened. “Closest” means the average identity across every pair spanning the two clumps (average linkage / UPGMA). 117 merges later everything is one tree.
  • Reading groups off the tree: draw a vertical line across it — whatever clumps remain to the left are the groups. The line is an identity threshold, so it is a number you can quote (the conventional CYP one is 40%).
  • The family labels are never given to any of these steps. They are used only afterwards, to check whether the result matches them.
720
n=67
725
n=31
947
n=8
90
n=5
88
n=4
974
n=2
723
n=1
CYP720 n=6743.527.827.433.526.733.828.2
CYP725 n=3127.861.630.828.027.726.562.2
CYP947 n=827.430.864.028.425.927.331.0
CYP90 n=533.528.028.434.927.333.826.4
CYP88 n=426.727.725.927.360.327.327.6
CYP974 n=233.826.527.333.827.373.326.7
CYP723 n=128.262.231.026.427.626.7–
25%40%60%74%outlined diagonal = within a family · hover a cell for the pair count
Fig C1. Median % identity between families — the same 6,903 pairwise comparisons, summarised one family against another. The outlined diagonal is how similar members of a family are to each other; every other cell is family versus family. The pattern is the whole argument in one table: within a family 35–73%, between families 25–34%. The single exception is CYP725 × CYP723 at 62.2% — both are taxane enzymes, and the lone CYP723 member is ~99% identical to a CYP725 by its own annotation. CYP723 has no diagonal value because a single sequence has no within-family pair.
100%80%60%40%CYP88V1 · CYP88A0A0B6VL99 · Lygodium japonicum · CYP88CYP88V2 · Ceratopteris richardii · CYP88CYP88V2 · Ceratopteris richardii · CYP88CYP90Z1 · Crocus sativus · CYP90CYP974B2 · Pteridium aquilinum · CYP974CYP974B7 · Alsophila spinulosa · CYP974CYP720B1 · A0A1D1XKS0 · Anthurium amnicola · CYP720A0A6J0L5E5 · Raphanus sativus · CYP720A0A397YRM2 · Brassica campestris · CYP720A0A816ZCX2 · Brassica napus · CYP720A0A3P6BA32 · Brassica campestris · CYP720M4DHM1 · Brassica campestris · CYP720A0A816QUK8 · Brassica napus · CYP720A0A0D3D0C8 · Brassica oleracea · CYP720A0A3P6HEF1 · Brassica oleracea · CYP720A0A087HI20 · Arabis alpina · CYP720A0A565B0E9 · Arabis nemorensis · CYP720A0A6D2K153 · Microthlaspi erraticum · CYP720R0GGB8 · Capsella rubella · CYP720D7KR97 · Arabidopsis lyrata · CYP720A0A178WHC7 · Arabidopsis thaliana · CYP720A0A654ENM9 · Arabidopsis thaliana · CYP720CYP720A1 · Thlaspi arvense · CYP720V4KKP9 · Eutrema salsugineum · CYP720A0A3S3MYW9 · Cinnamomum micranthum · CYP720A0A218XDD4 · Punica granatum · CYP720CYP720A1 · Eucalyptus grandis · CYP720A0A059A4T0 · Eucalyptus grandis · CYP720CYP720A1 · Sesamum indicum · CYP720A0A5C7H470 · Acer yangbiense · CYP720CYP720A1 · Citrus sinensis · CYP720A0A2H5QBX7 · Citrus unshiu · CYP720A0A7J7MET3 · Kingdonia uniflora · CYP720A0A6A1VJT6 · Morella rubra · CYP720A0A834WF25 · Senna tora · CYP720A0A498I5A9 · Malus domestica · CYP720A0A6N2LIX9 · Salix viminalis · CYP720CYP720A1 · Ampelopsis grossedentata · CYP720CYP720B1 · A0A438GCL4 · Vitis vinifera · CYP720D7T7D3 · Vitis vinifera · CYP720CYP720B2 · A0A076UAC7 · Pinus contorta · CYP720CYP720B2 · A0A076UAC3 · Pinus banksiana · CYP720CYP720B2 · Pinus taeda · CYP720CYP720B2 · Q50EK5 · Pinus taeda · CYP720CYP720B12 · Pinus taeda · CYP720CYP720B12 · A0A076U5L5 · Pinus contorta · CYP720CYP720PB12 · A0A076U7E2 · Pinus banksiana · CYP720CYP720B2 · A0A0G7ZNR4 · Picea glauca · CYP720CYP720B2 · E5FA69 · Picea sitchensis · CYP720CYP720B12 · A0A0G7ZP11 · Picea glauca · CYP720CYP720B12 · E5FA64 · Picea sitchensis · CYP720CYP720B12 · Picea sitchensis · CYP720CYP720B12 · Picea glauca · CYP720C0PR04 · Picea sitchensis · CYP720CYP720B12 · Picea glauca · CYP720CYP720B12 · A0A0G7ZNV7 · Picea glauca · CYP720CYP720PB11 · A0A076U3J6 · Pinus banksiana · CYP720CYP720B11 · A0A076U7E6 · Pinus contorta · CYP720CYP720B10 · Pinus taeda · CYP720CYP720B10 · A0A076U5E1 · Pinus contorta · CYP720CYP720B10 · A0A076U3K0 · Pinus contorta · CYP720CYP720B10 · A0A076U5D6 · Pinus banksiana · CYP720CYP720B68 · Taiwania cryptomerioides · CYP720CYP720B26 · Taxus cuspidata · CYP720CYP720B26 · A0A291FAZ1 · Taxus chinensis · CYP720CYP720B25 · A0A291FB22 · Taxus chinensis · CYP720CYP720B23 · Taxus cuspidata · CYP720CYP720B23 · A0A291FAU0 · Taxus chinensis · CYP720CYP720B35 · Taxus cuspidata · CYP720CYP720B24 · A0A291FB28 · Taxus chinensis · CYP720CYP720B36 · Taxus cuspidata · CYP720CYP720B30 · CYP720CYP720B69 · Taiwania cryptomerioides · CYP720CYP90L3 · Huperzia serrata · CYP90CYP90Y1 · Huperzia serrata · CYP90CYP90V2 · Ceratopteris richardii · CYP90A0A6I9R9G3 · Elaeis guineensis · CYP90CYP947A91 · CYP947CYP947A89 · Picea glauca · CYP947A0A0D6QVC1 · Araucaria cunninghamii · CYP947A0A0D6QWU3 · Araucaria cunninghamii · CYP947CYP947A90 · CYP947CYP947A56 · Taxus cuspidata · CYP947CYP947A88 · Taxus chinensis · CYP947CYP947A88 · Taxus cuspidata · CYP947CYP725A55 · Taxus chinensis · CYP725CYP725A55 · Taxus mairei · CYP725A0A0C9S3N6 · CYP725CYP725A4 · Taxus chinensis · CYP725H9BII6 · Taxus x · CYP725CYP725A4 · Taxus cuspidata · CYP725Q6WG30 · Taxus cuspidata · CYP725CYP725A3 · CYP725CYP725A19 · A0A291FAY5 · Taxus chinensis · CYP725CYP725A37 · Taxus cuspidata · CYP725CYP723A37 · Taxus mairei · CYP723CYP725A37 · Taxus chinensis · CYP725CYP725A39 · Taxus cuspidata · CYP725CYP725A18 · A0A291FAT7 · Taxus chinensis · CYP725CYP725A9 · Taxus chinensis · CYP725CYP725A9 · Taxus cuspidata · CYP725CYP725A38 · Taxus cuspidata · CYP725CYP725A11 · A0A291FAT3 · Taxus chinensis · CYP725CYP725A21 · Taxus cuspidata · CYP725CYP725A21 · A0A291FB05 · Taxus chinensis · CYP725CYP725A23 · Taxus chinensis · CYP725CYP725A23 · A0A291FB18 · Taxus chinensis · CYP725Q8H6A8 · Taxus chinensis · CYP725Q6SPR0 · Taxus x · CYP725CYP725A1 · Q9AXM6 · Taxus cuspidata · CYP725H9BII4 · Taxus x · CYP725CYP725A27 · CYP725CYP725A25 · CYP725CYP725A20 · Taxus cuspidata · CYP725CYP725A20 · A0A291FAZ6 · Taxus chinensis · CYP725CYP725A26 · CYP725CYP725A36 · Taxus cuspidata · CYP725CYP88 (4)CYP974 (2)CYP720 (67)CYP90 (4)CYP947 (8)CYP725 (10)CYP725 (21)% identity at which two branches join — read left (similar) to right (different)
Fig C2. The same comparisons drawn as a tree: every one of the 118 sequences is a tick on the left, and two branches join at the % identity shown on the scale along the top — so the horizontal position is a real, readable number, not an abstract coordinate. Sequences that join far to the left are near-identical; branches that only meet at the right-hand end are barely related. The coloured bars name the family of each block. Note that the families emerge as blocks on their own — the tree was built from identity alone and never saw the labels. The dashed line is the 40% identity used as the conventional CYP family boundary. Note how much empty space surrounds it: the family blocks finish joining up well to its left, and nothing else joins until well to its right. That gap is why the exact threshold does not matter — slide the line anywhere between roughly 35% and 41% and the same blocks come out.
11888592 groups: 71 of 118 sequences with their own family23 groups: 102 of 118 sequences with their own family34 groups: 110 of 118 sequences with their own family45 groups: 111 of 118 sequences with their own family56 groups: 113 of 118 sequences with their own family67 groups: 117 of 118 sequences with their own family78 groups: 117 of 118 sequences with their own family89 groups: 117 of 118 sequences with their own family910 groups: 117 of 118 sequences with their own family1011 groups: 117 of 118 sequences with their own family1112 groups: 117 of 118 sequences with their own family127 groups → 117 of 118number of groups the sequences were cut intosequences landing with their own family
Fig C3. A check that the grouping is not sensitive to where the line is drawn. The tree was cut into 2, 3, 4 … 12 groups in turn, and each time we counted how many of the 118 sequences ended up in a group made mostly of their own family. At 7 groups — the number of families — 117 of 118 land correctly, and the curve is flat after that: asking for more groups only shaves single sequences off, it does not rearrange anything.

Cutting at 40% identity

Cutting the tree at the conventional CYP-family threshold (40% identity) — rather than asking for a fixed number of groups — gives 9 clusters, and 117 of the 118 sequences sit in a cluster made of their own family. Every cluster and what it contains:

clustercomposition
#1CYP88 ×4
#2CYP974 ×2
#3CYP720 ×67
#4CYP90 ×2
#5CYP90 ×1
#6CYP90 ×1
#7CYP90 ×1
#8CYP947 ×8
#9CYP725 ×31, CYP723 ×1
why CYP90 fragments
Four of the nine clusters are CYP90 sequences sitting on their own. That is not a clustering failure — the five CYP90 members are genuinely that far apart. They are five different subfamilies (90A3, 90L3, 90V2, 90Y1, 90Z1) from plants separated by hundreds of millions of years: a clubmoss (Huperzia), a fern (Ceratopteris), saffron and oil palm. Their pairwise identities run 29.6–51.5%, and only the two from the same species (90L3 × 90Y1, 51.5%) clear the 40% line at all — so at a 40% cut they cannot group, by construction.
This is the one place where the family label and the sequence evidence disagree, and it is worth being explicit: CYP90 is held together by its name and its chemistry, not by sequence similarity within this set. The other evidence still places them together — the language-model method (Fig G1) recovers all five as one cluster, and they share the brassinosteroid pocket signature — but on identity alone they are five separate lineages.
conclusion
An algorithm that was never shown the family names recovers them almost perfectly — 117 of the 118 sequences end up grouped with their own family. Within-family identity (median 57.9%) and between-family identity (median 28.1%) form two clearly separated populations. The 7 families are therefore measurable groupings in sequence space, not only nomenclature.
D · structure evidence

One fold everywhere — shape barely separates the families

Question: The sequences differ this much — do the structures differ too?
method
  • Fetch an AlphaFold model for every family member carrying a UniProt accession, up to 10 per family, plus the experimental structure 8X1W.
  • That gives 25 structures — CYP725 11, CYP720 10, CYP947 2, CYP88 1, CYP90 1. CYP723 and CYP974 have no accession, so no model.
  • TM-align over all 300 pairs; each score = mean of the two chain-length normalisations.
725
11 struct
720
10 struct
947
2 struct
88
1 struct
90
1 struct
CYP725 11 struct0.93
55 pairs
0.88
110 pairs
0.89
22 pairs
0.85
11 pairs
0.88
11 pairs
CYP720 10 struct0.88
110 pairs
0.96
45 pairs
0.87
20 pairs
0.85
10 pairs
0.89
10 pairs
CYP947 2 struct0.89
22 pairs
0.87
20 pairs
0.90
1 pairs
0.84
2 pairs
0.86
2 pairs
CYP88 1 struct0.85
11 pairs
0.85
10 pairs
0.84
2 pairs
–0.83
1 pairs
CYP90 1 struct0.88
11 pairs
0.89
10 pairs
0.86
2 pairs
0.83
1 pairs
–
outlined diagonal = within a familyeverything else = between families– = that family has only one structure, so no within-family pair exists
Fig D1. TM-align summarised family against family — read exactly like Fig C1: the outlined diagonal is within a family, everything else is between families, and the small figure in each cell is how many structure pairs the median covers. Within a family the median TM is 0.943 (101 pairs); between families it is 0.875 (199 pairs). CYP88 and CYP90 contribute one structure each, so they have no within-family pair and their diagonal is empty.
conclusion
Every pair scores far above 0.5, so all 25 structures share one fold. Structure does separate the families slightly — median 0.943 within versus 0.875 between — but the ranges overlap heavily (0.877–0.995 against 0.828–0.915): the most dissimilar pair inside a family is less alike than the most similar pair across families. Contrast Fig C1, where within-family and between-family identity barely touch at all. Overall shape is a weak discriminator, so the functional difference has to be local. The substrate pocket is next.
E · the substrate pocket

The catalytic core is shared; the substrate pocket is not

Question: With the fold and the catalytic core shared, what physically makes one enzyme produce Taxol and another resin acid?
method
  • 1 · Find a structure with the substrate still in it. 8X3E is CYP725A4 crystallised as a holo complex — haem plus a bound taxadiene analogue (ligand A1L3H). That ligand marks where the substrate actually sits, so the pocket does not have to be predicted.
  • 2 · Take every residue that touches it. For each amino acid in the structure, measure the shortest distance from any of its atoms to any ligand atom. Keep those within 4.5 Å — van der Waals contact range. That gives 13 residues, listed with their real distances in the table below (2.35–4.19 Å).
  • 3 · Translate structure numbering into sequence numbering. The crystal construct starts at residue 43, so every structure number is 32 lower than the full-sequence number (Trp93 in the structure is Trp125 in the sequence).
  • 4 · Carry those positions onto all 118 sequences. The crystallised protein is sp|Q6WG30, which is itself one of the 118 — so its residues can be located in the MAFFT alignment directly, and the alignment column of each pocket residue is read off. No structural superposition and no homology modelling is involved in this step.
  • 5 · Read those 13 columns in every sequence. Each of the 118 now has a 13-letter pocket fingerprint. Per family, the consensus residue at each position is taken along with how much of the family agrees.
  • Assumption, stated: this maps one structure’s pocket onto every sequence, which holds only because all 118 share the same fold (Fig D1, TM 0.83–0.99) and the alignment is unambiguous at these columns.
pocket residue
(8X3E / seq)
dist to
substrate
CYP725
reference
7239477208890974how many of the 6
differ from it
TRP93 / 1252.49 ÅW 65%G 100%W 100%Y 93%W 100%Y 60%Y 100%4 of 6
PHE97 / 1293.63 ÅF 74%F 100%L 38%F 70%T 100%F 40%I 50%3 of 6
MET101 / 1332.74 ÅM 48%V 100%I 100%V 51%M 100%F 40%I 100%5 of 6
ALA107 / 1392.99 ÅA 42%V 100%V 50%I 51%T 100%L 75%T 100%6 of 6
LEU193 / 2254.19 ÅL 71%L 100%S 62%V 63%N 100%V 60%V 50%5 of 6
PHE197 / 2293.59 ÅL 39%L 100%L 100%L 93%R 100%L 40%L 100%1 of 6
SER270 / 3022.43 ÅA 32%F 100%G 88%N 49%M 100%F 40%G 100%6 of 6
LEU271 / 3033.93 ÅL 58%M 100%L 88%L 94%Y 100%L 60%I 50%3 of 6
HIS273 / 3052.35 ÅH 90%H 100%T 50%F 100%N 100%V 40%F 100%5 of 6
ALA274 / 3062.65 ÅA 81%A 100%G 75%A 100%A 100%A 80%A 100%1 of 6
THR278 / 3103.74 ÅT 81%T 100%T 100%T 100%S 100%T 60%T 100%1 of 6
VAL342 / 3743.45 ÅV 48%A 100%I 50%A 70%S 100%A 40%G 100%6 of 6
LEU451 / 4832.75 ÅF 45%F 100%F 62%S 49%H 100%T 40%G 100%4 of 6
Fig E1. The 13 residues that touch the bound substrate in holo 8X3E, and what each family has at that position (consensus residue + how much of the family agrees). CYP725 is the reference column — it is the taxane enzyme whose structure defined the pocket, so it is set apart rather than scored against itself. For the other six families: green = same residue as CYP725, amber = different. Blue-shaded rows are the three aromatic positions (W93 / F97 / F197) that pack against the taxane core in 8X3E. Note that these are not uniformly "taxane-only": CYP88 also carries Trp at 93, and even within CYP725 position 197 is predominantly Leu rather than Phe, so the cage is a tendency across the family rather than a strict signature.
conclusion
13 of 13 pocket positions differ between families. The two least variable positions belong to the catalytic machinery rather than the substrate cage: Thr278 is the I-helix threonine and Ala274 is a small floor residue. The residues that contact the substrate do change — for example Trp93 becomes Tyr in CYP720 (93%) and CYP90 (60%), and Phe97 becomes Thr in CYP88. The catalytic chemistry is shared while the substrate cage is family-specific, which provides a structural basis for the functional split.
E → F · how the five groups were drawn

The five pocket signatures, side by side

Question: Where exactly do the five groups come from?
short answer, before the figure
CYP720  67→ group B resin acidkept as it is
CYP88    4→ group C gibberellinkept as it is
CYP90    5→ group D brassinosteroidkept as it is
CYP974   2→ group E uncharacterisedkept as it is
CYP725  31
CYP723   1
CYP947   8
→ group A taxanethe only merge
7 families − 2 merges = 5 groups.
TRP93PHE97MET101ALA107LEU193PHE197SER270LEU271HIS273ALA274THR278VAL342LEU451A · taxane / Taxoln = 40WGAFLVSYIIMVLFVASILTLVSANLMFVGASFLVLMIFHTEAGSTSVIALSFLVB · resin acidn = 67YVFLTVIILSAVSLVNFLFFATAGTVSFLVC · gibberellinn = 4WTMTNRMYNASSHD · brassinosteroidn = 5YHLFIMFMALLPVLILFMFSGALFVVSFAAPTSATGVTLFRE · uncharacterisedn = 2YILITVFLGIVFATGGgroup consensusWFIVLLGLHATVFAYFVIVLNLFATASBWTMTNRMYNASSHCYFFLVLFLVATATDYIITVLGIFATGGE
Fig E2. One sequence logo per functional group, over the 13 pocket positions. Letter height = how much of that group carries that residue, so a single tall letter means the whole group agrees and a stack means it varies. Colour = amino-acid chemistry (violet aromatic, grey hydrophobic, green polar, blue basic, red acidic). The row underneath is each group’s consensus, shaded by how strongly the group holds it. Read down a column to compare the five groups at one position; read across a row to see one group’s whole pocket.

Read it down the columns. Each group holds its own pattern, and the differences are chemical, not cosmetic:

grouppocket consensuswhat stands out
A taxaneWFIVLLGLHATVFthe only group with Trp at the first position and His in the middle — the aromatic cage that grips the taxane ring
B resin acidYFVIVLFLFATASTrp→Tyr, His→Phe: a smaller, flatter, fully hydrophobic cavity
C gibberellinWTMTNRMYNASSHthe odd one out — Arg, Asn, His where every other group is hydrophobic: a polar, charged pocket for a carboxylic-acid substrate
D brassinosteroidYFFLVLFLVATATclosest to B, differing at the positions that set cavity length — consistent with a long sterol side chain
E fern CYPYIITVLGIFATGGGly where others carry bulky residues, so the widest cavity of the five; no characterised member, so no substrate is claimed
conclusion
Group C is separable on the pocket alone — a charged cavity among hydrophobic ones. Groups A and B differ consistently at the aromatic cage. D and E sit close to B and could not be told apart on these 13 residues alone, which is why they are held apart on family membership and on chemistry from the literature, not on the pocket. Everything up to here says “different substrates”; nothing up to here says which. That comes next, with sources.
F · the answer, with sources

Five substrate-defined functional groups

method
  • Reaction quoted from a reviewed (Swiss-Prot) entry where one exists, with EC number and Rhea id, plus EuropePMC literature.
  • Where the set holds only unreviewed (TrEMBL) entries, the reaction is taken from a characterised relative — and said so.
  • Where nothing is characterised, the card says so.
A · Taxane oxygenases · Taxol pathway
CYP725 + CYP723 + CYP947* · n=40
taxa-4(5),11(12)-diene
C20H32
O₂ + NADPH
⟶
5α-hydroxylation
taxa-4(20),11-dien-5alpha-ol
C20H32O
EC 1.14.14.176 · UniProt Q6WG30 (UniProtKB reviewed (Swiss-Prot), Taxus cuspidata) · Rhea RHEA:14049, RHEA-COMP:11964, RHEA-COMP:11965
Why this entry: This exact protein IS in the 118 (sp|Q6WG30) and IS the 8X1W/8X3E structure. Swiss-Prot reviewed.
Caveat: CYP947 is a Taxus/gymnosperm taxane-pathway candidate: it forms its own clean cluster (n=8) and has no characterised member, so its placement in group A is by pathway context, not by a measured reaction.
B · Diterpene resin acid (abietane) oxidases
CYP720 (720B + 720A) · n=67
abietadienol
C20H32O
O₂ + NADPH
⟶
C-18 oxidation (3 steps)
abietic acid
C20H30O2
no reviewed reaction available for this set
Caveat: No Swiss-Prot entry for the CYP720B members in this set (all TrEMBL, no EC). The reaction below is from the literature on characterised conifer CYP720B enzymes.
C · Gibberellin biosynthesis
CYP88 (fern CYP88V + KAO) · n=4
ent-kaurenoic acid
C20H30O2
O₂ + NADPH
⟶
C-7 oxidation → GA12 (3 steps)
gibberellin A12
C20H28O4
EC 1.14.14.107 · UniProt O23051 (UniProtKB reviewed (Swiss-Prot), Arabidopsis thaliana) · Rhea RHEA:33219, RHEA-COMP:11964, RHEA-COMP:11965
Why this entry: Swiss-Prot KAO1 (CYP88A3, Arabidopsis) - the characterised CYP88A ent-kaurenoic acid oxidase. The 118 contain LjKAO (A0A0B6VL99, TrEMBL) whose name matches this activity.
Caveat: The set's members are TrEMBL (LjKAO, A0A0B6VL99). Reaction taken from the characterised Arabidopsis KAO1.
D · Brassinosteroid biosynthesis
CYP90 (90A3/90L3/90V2/90Y1/90Z1) · n=5
22-hydroxycampesterol
C28H48O2
O₂ + NADPH
⟶
C-3 oxidation
(22S)-22-hydroxycampest-4-en-3-one
C28H46O2
EC 1.14.19.79 · UniProt Q42569 (UniProtKB reviewed (Swiss-Prot), Arabidopsis thaliana) · Rhea RHEA:69971, RHEA-COMP:11964, RHEA-COMP:11965
Why this entry: Swiss-Prot CPD/CYP90A1 (Arabidopsis). The 118 contain CYP90A3 (A0A6I9R9G3, TrEMBL, auto-annotated as 3beta,22alpha-dihydroxysteroid 3-dehydrogenase isoform) - same family, same chemistry class.
Caveat: Set members are TrEMBL; reaction from the characterised Arabidopsis CPD/CYP90A1. Note these are five different subfamilies (90A3/90L3/90V2/90Y1/90Z1) spanning clubmoss, fern, saffron and oil palm, only 29.6-51.5% identical to each other - grouped by name and chemistry, not by sequence similarity.
E · Uncharacterized fern CYP
CYP974 (Pteridium + Alsophila) · n=2
no reviewed reaction available for this set
Caveat: EuropePMC returns ZERO publications for CYP974. No substrate, no reaction, no EC. Flagged as unknown - the honest answer.
EuropePMC: 0 publications
source / provenance
UniProt REST (rest.uniprot.org) → reaction_sources.json; EuropePMC REST search → lit_records.json. Every accession and PMID above is a live link.
conclusion
Only one of the 118 sequences (sp|Q6WG30, the CYP725A4 that is also the crystal structure) has a reviewed, EC-numbered reaction of its own. The remainder are TrEMBL entries. The split itself rests on the measurements in sections C, D and E; the substrate assignments for groups B/C/D are inherited from characterised relatives, and group E carries no assignment.
G · independent cross-check

A method that never saw the alignment finds the same groups

Question: Everything so far derives from one alignment — would an alignment-free method find the same groups?
method
  • PLMView pipeline: mmseqs2 easy-cluster (118 → 14 anchors) → ankh-large per-residue embeddings → MULAN LightAttention weighting → embedding-space Needleman–Wunsch → SCUT spectral tree.
  • Uses neither the MSA %identity nor any 3D structure — an independent route to the same question.
clusterndominantcompositionfunctional group
#13CYP88CYP88 3/3C gibberellin
#26CYP947CYP947 6/6E uncharacterized
#327CYP725CYP725 25 + CYP723 1 + CYP720 1A taxane (725+723 together)
#42CYP720CYP720 2/2B resin-acid
#56CYP720CYP720 6/6B resin-acid
#67CYP90CYP90 5 + CYP974 2D brassinosteroid (+CYP974)
#733CYP720CYP720 33/33B resin-acid
Fig G1. Cutting the PLMView tree into 7 clusters, versus the given families. Scored on the 84/118 sequences that could be assigned a family from the header text alone (80 carry an explicit CYPxxx name, 4 more were resolved from the UniProt description); the remaining 34 were held out of the metric.
A0A0D3D0C8 · Brassica oleracea · CYP720 · group B (resin acid)CYP720B10 · A0A076U5E1 · Pinus contorta · CYP720 · group B (resin acid)CYP88V2 · Ceratopteris richardii · CYP88 · group C (gibberellin)CYP720B10 · A0A076U3K0 · Pinus contorta · CYP720 · group B (resin acid)CYP725A38 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)A0A834WF25 · Senna tora · CYP720 · group B (resin acid)CYP720B23 · Taxus cuspidata · CYP720 · group B (resin acid)CYP720B68 · Taiwania cryptomerioides · CYP720 · group B (resin acid)A0A3P6HEF1 · Brassica oleracea · CYP720 · group B (resin acid)CYP88V1 · CYP88 · group C (gibberellin)CYP90Y1 · Huperzia serrata · CYP90 · group D (brassinosteroid)CYP720B36 · Taxus cuspidata · CYP720 · group B (resin acid)CYP720B12 · Picea sitchensis · CYP720 · group B (resin acid)CYP720B10 · A0A076U5D6 · Pinus banksiana · CYP720 · group B (resin acid)CYP720B2 · Q50EK5 · Pinus taeda · CYP720 · group B (resin acid)CYP725A18 · A0A291FAT7 · Taxus chinensis · CYP725 · group A (taxane / Taxol)H9BII6 · Taxus x · CYP725 · group A (taxane / Taxol)CYP725A1 · Q9AXM6 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)Q6WG30 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)CYP720B12 · A0A0G7ZP11 · Picea glauca · CYP720 · group B (resin acid)CYP720A1 · Sesamum indicum · CYP720 · group B (resin acid)CYP88V2 · Ceratopteris richardii · CYP88 · group C (gibberellin)Q6SPR0 · Taxus x · CYP725 · group A (taxane / Taxol)CYP720B25 · A0A291FB22 · Taxus chinensis · CYP720 · group B (resin acid)CYP720B30 · CYP720 · group B (resin acid)CYP720B2 · Pinus taeda · CYP720 · group B (resin acid)CYP725A23 · A0A291FB18 · Taxus chinensis · CYP725 · group A (taxane / Taxol)A0A6I9R9G3 · Elaeis guineensis · CYP90 · group D (brassinosteroid)CYP720PB12 · A0A076U7E2 · Pinus banksiana · CYP720 · group B (resin acid)CYP947A91 · CYP947 · group A (taxane / Taxol)CYP720B69 · Taiwania cryptomerioides · CYP720 · group B (resin acid)CYP720B12 · E5FA64 · Picea sitchensis · CYP720 · group B (resin acid)C0PR04 · Picea sitchensis · CYP720 · group B (resin acid)D7KR97 · Arabidopsis lyrata · CYP720 · group B (resin acid)CYP725A36 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)A0A178WHC7 · Arabidopsis thaliana · CYP720 · group B (resin acid)A0A059A4T0 · Eucalyptus grandis · CYP720 · group B (resin acid)A0A6J0L5E5 · Raphanus sativus · CYP720 · group B (resin acid)A0A0D6QWU3 · Araucaria cunninghamii · CYP947 · group A (taxane / Taxol)CYP720B11 · A0A076U7E6 · Pinus contorta · CYP720 · group B (resin acid)A0A397YRM2 · Brassica campestris · CYP720 · group B (resin acid)CYP725A4 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)A0A0D6QVC1 · Araucaria cunninghamii · CYP947 · group A (taxane / Taxol)A0A3P6BA32 · Brassica campestris · CYP720 · group B (resin acid)CYP720A1 · Ampelopsis grossedentata · CYP720 · group B (resin acid)A0A6N2LIX9 · Salix viminalis · CYP720 · group B (resin acid)CYP720B2 · A0A076UAC7 · Pinus contorta · CYP720 · group B (resin acid)CYP947A90 · CYP947 · group A (taxane / Taxol)H9BII4 · Taxus x · CYP725 · group A (taxane / Taxol)A0A3S3MYW9 · Cinnamomum micranthum · CYP720 · group B (resin acid)Q8H6A8 · Taxus chinensis · CYP725 · group A (taxane / Taxol)A0A565B0E9 · Arabis nemorensis · CYP720 · group B (resin acid)CYP720B24 · A0A291FB28 · Taxus chinensis · CYP720 · group B (resin acid)CYP725A11 · A0A291FAT3 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP725A3 · CYP725 · group A (taxane / Taxol)CYP725A55 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP947A88 · Taxus cuspidata · CYP947 · group A (taxane / Taxol)CYP725A39 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)CYP725A20 · A0A291FAZ6 · Taxus chinensis · CYP725 · group A (taxane / Taxol)V4KKP9 · Eutrema salsugineum · CYP720 · group B (resin acid)A0A498I5A9 · Malus domestica · CYP720 · group B (resin acid)D7T7D3 · Vitis vinifera · CYP720 · group B (resin acid)CYP720B2 · A0A0G7ZNR4 · Picea glauca · CYP720 · group B (resin acid)CYP720B2 · E5FA69 · Picea sitchensis · CYP720 · group B (resin acid)CYP720A1 · Eucalyptus grandis · CYP720 · group B (resin acid)CYP720B10 · Pinus taeda · CYP720 · group B (resin acid)R0GGB8 · Capsella rubella · CYP720 · group B (resin acid)CYP947A56 · Taxus cuspidata · CYP947 · group A (taxane / Taxol)A0A2H5QBX7 · Citrus unshiu · CYP720 · group B (resin acid)CYP90L3 · Huperzia serrata · CYP90 · group D (brassinosteroid)CYP720B12 · Picea glauca · CYP720 · group B (resin acid)CYP725A21 · A0A291FB05 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP947A88 · Taxus chinensis · CYP947 · group A (taxane / Taxol)CYP947A89 · Picea glauca · CYP947 · group A (taxane / Taxol)CYP90Z1 · Crocus sativus · CYP90 · group D (brassinosteroid)CYP725A20 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)A0A087HI20 · Arabis alpina · CYP720 · group B (resin acid)CYP723A37 · Taxus mairei · CYP723 · group A (taxane / Taxol)CYP725A19 · A0A291FAY5 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP720B35 · Taxus cuspidata · CYP720 · group B (resin acid)CYP720B26 · A0A291FAZ1 · Taxus chinensis · CYP720 · group B (resin acid)CYP720B23 · A0A291FAU0 · Taxus chinensis · CYP720 · group B (resin acid)A0A218XDD4 · Punica granatum · CYP720 · group B (resin acid)CYP720B12 · Pinus taeda · CYP720 · group B (resin acid)A0A5C7H470 · Acer yangbiense · CYP720 · group B (resin acid)CYP720B12 · A0A0G7ZNV7 · Picea glauca · CYP720 · group B (resin acid)CYP725A26 · CYP725 · group A (taxane / Taxol)CYP725A9 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP720A1 · Thlaspi arvense · CYP720 · group B (resin acid)A0A6D2K153 · Microthlaspi erraticum · CYP720 · group B (resin acid)CYP725A21 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)A0A0B6VL99 · Lygodium japonicum · CYP88 · group C (gibberellin)CYP725A55 · Taxus mairei · CYP725 · group A (taxane / Taxol)CYP725A4 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP725A37 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)CYP725A23 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP725A9 · Taxus cuspidata · CYP725 · group A (taxane / Taxol)CYP720A1 · Citrus sinensis · CYP720 · group B (resin acid)CYP720PB11 · A0A076U3J6 · Pinus banksiana · CYP720 · group B (resin acid)CYP720B12 · Picea glauca · CYP720 · group B (resin acid)A0A654ENM9 · Arabidopsis thaliana · CYP720 · group B (resin acid)A0A816QUK8 · Brassica napus · CYP720 · group B (resin acid)CYP720B2 · A0A076UAC3 · Pinus banksiana · CYP720 · group B (resin acid)CYP720B1 · A0A438GCL4 · Vitis vinifera · CYP720 · group B (resin acid)CYP720B1 · A0A1D1XKS0 · Anthurium amnicola · CYP720 · group B (resin acid)A0A7J7MET3 · Kingdonia uniflora · CYP720 · group B (resin acid)CYP90V2 · Ceratopteris richardii · CYP90 · group D (brassinosteroid)CYP725A25 · CYP725 · group A (taxane / Taxol)CYP725A37 · Taxus chinensis · CYP725 · group A (taxane / Taxol)CYP974B7 · Alsophila spinulosa · CYP974 · group E (uncharacterised)CYP725A27 · CYP725 · group A (taxane / Taxol)M4DHM1 · Brassica campestris · CYP720 · group B (resin acid)CYP720B12 · A0A076U5L5 · Pinus contorta · CYP720 · group B (resin acid)A0A0C9S3N6 · CYP725 · group A (taxane / Taxol)CYP974B2 · Pteridium aquilinum · CYP974 · group E (uncharacterised)A0A6A1VJT6 · Morella rubra · CYP720 · group B (resin acid)CYP720B26 · Taxus cuspidata · CYP720 · group B (resin acid)A0A816ZCX2 · Brassica napus · CYP720 · group B (resin acid)A taxane / TaxolB resin acidC gibberellinD brassinosteroidE uncharacterised
Fig G2. The language model’s own view of the 118, coloured by the functional group they were assigned in section F. This uses PLMView’s embedding representation directly — no alignment, no structure, and the group labels were applied only after the coordinates existed. If the grouping were an artefact of the alignment, these colours would be scattered at random. They are not: each group occupies its own region.
conclusion
An orthogonal method reproduces the same clusters: CYP88 and CYP947 clean, CYP725+CYP723 together (both taxane), CYP90 with CYP974, and CYP720 splitting into sub-clades. The grouping is reproduced independently of the alignment and of any single algorithm.
NO — not one function · 5 groups

The 118 are all genuine plant cytochrome-P450 monooxygenases: one fold (TM 0.83–0.91), one catalytic core (haem Cys 118/118). They are not one enzyme. They divide into five substrate-defined groups: taxane / Taxol, diterpene resin acids, gibberellin, brassinosteroid, and one uncharacterised fern CYP.

Four independent lines agree: unsupervised clustering reproduces the given families (117 of 118 sequences grouped with their own); structure explicitly fails to separate them, which localises the difference; 13/13 substrate-pocket positions differ between families while the catalytic residues do not; and a protein-language-model tree that never saw the alignment recovers the same clusters (group purity 0.964).

Known limits: only 1 of the 118 has a reviewed EC-numbered reaction of its own; groups B/C/D inherit their substrate assignment from characterised relatives; CYP974 (n=2) has zero publications and no assignment; CYP90 is internally divergent (within-family median identity 34.9%); 14 near-identical pairs were kept in, medians used throughout.