Every claim below is re-derived from the primary data in this repository and shown as a figure:
the alignment, the motif columns, the clustering, the structures, the substrate pocket, and the reactions with their
UniProt/EC/PMID sources. Where the evidence contradicted the previous version of this report, the correction is
stated explicitly rather than quietly fixed.
118
sequences
7
CYP families (given)
5
functional groups
0.992
clustering purity vs families
2
experimental structures
A · the input
Where the family names come from
Question: Where do the family names come from — the FASTA files, UniProt, or this analysis?
method
The 118 arrive already split into 7 per-family FASTA files (FF160_data/sequence/).
The family label of a sequence = which file it came from. No family is guessed from a name.
Fig A1. The 7 given families and their sizes (n=118). Bar length = share of the set. Colour = family, used consistently in every later figure.
B · the shared catalytic core
What "they are all P450s" actually means, residue by residue
Question: What exactly is the shared catalytic core — which motifs, located how, and how conserved is each one across the 118?
method
Locate 3 canonical P450 motifs, then map each back to its MSA column and read that column across all 118.
Haem loop: regex on the ungapped sequence, restricted to the C-terminal 40% where the motif belongs — matches 118/118.
I-helix: same approach, but the regex is strict and matches only 82/118; the alignment column is the evidence, and it carries Thr/Ser in 117/118.
ExxR: located by column scan (E at column c, R at c+3) — a raw-sequence regex also hits spurious ExxR elsewhere.
Every motif checked against the crystal structure 8X3E.
Conservation = scorecons.
Fig B1. Sequence logo over the heme-binding loop — FxxGxRxCxG, built from all 118 aligned sequences. Letter height = frequency of that residue in that column (non-gap rows); colour = residue chemistry. Columns 648–661 of the alignment; the shaded columns are the motif positions themselves, with scorecons: F650 (1.000), G653 (1.000), R655 (0.958), C657 (1.000), G659 (0.979).
Fig B2. Sequence logo over the K-helix salt bridge — ExxR, built from all 118 aligned sequences. Letter height = frequency of that residue in that column (non-gap rows); colour = residue chemistry. Columns 569–576 of the alignment; the shaded columns are the motif positions themselves, with scorecons: E571 (0.990), R574 (1.000).
Fig B3. Sequence logo over the I-helix acid/alcohol pair — (A/G)Gx(E/D)T, built from all 118 aligned sequences. Letter height = frequency of that residue in that column (non-gap rows); colour = residue chemistry. Columns 485–493 of the alignment; the shaded columns are the motif positions themselves, with scorecons: A486 (0.866), G487 (0.730), E489 (0.791), T490 (0.876).
The numbers, per motif
motif
what it does
MSA column
residues in that column (of 118)
scorecons
haem loop FxxGxRxCxG
its invariant Cys is the thiolate that ligates the haem iron — the residue that makes a P450 a P450
657
Cys 118/118
1.000
K-helix ExxR
Glu and Arg form an ion pair that locks the core fold — structural, not catalytic, hence two residues
571 / 574
E 117/118 (1 Lys) · R 118/118
0.990 / 1.000
I-helix acid/alcohol
the Thr/Ser that delivers protons to split O₂
490
Thr 106 · Ser 11 · Asn 1
0.876
Verified in the crystal structure, not just in the alignment
The construct in 8X1W/8X3E starts at residue 43, so structure numbering = full-sequence numbering − 32.
Both numbers are given here; the previous version quoted structure numbers only, which do not map onto the sequences without this offset.
role
sequence #
structure #
structural evidence
haem axial thiolate
445
413
SG-Fe distance 2.69 A in 8X3E (coordination bond)
I-helix acid/alcohol Thr (O2 activation)
310
278
THR at structure residue 278; alignment column 490 is Thr 106 / Ser 11 / Asn 1 of 118
K-helix ExxR salt bridge
366/369
334/337
GLU334 + ARG337 in 8X3E; alignment columns 571/574 are E 117/118 and R 118/118
Fig B4. scorecons for every one of the 741 alignment columns.
How to read it. Each vertical bar is one alignment column; height = its scorecons score
(1.0 = every one of the 118 sequences agrees). Hover any bar to see the exact score and which residues
actually occur there, with counts. Red marks the four key residues of the three motifs: the haem-ligating
Cys (1.000), the two halves of the ExxR salt bridge
— Glu (0.990) and Arg (1.000),
which pair with each other and so are two columns, not one — and the I-helix Thr
(0.876).
Across the whole alignment 133 columns score ≥0.7 and 42 score ≥0.9
(mean 0.388).
conclusion
All 118 carry the complete P450 catalytic machinery: haem Cys 118/118, ExxR 117/118 (one sequence has Lys for Glu), I-helix Thr/Ser 117/118. All 118 therefore run the same O₂-activation chemistry, and the catalytic core does not differentiate them. The following sections examine the alignment as a whole, then the substrate pocket.
C · sequence evidence
Hide the labels, and the sequences sort themselves back into the families
Question: What does the clustering look like, and does it agree with the family assignment?
method
Compare every sequence with every other one — all 6,903 pairs — and record their % identity from the MAFFT alignment.
That single table of pairwise identities is everything section C uses: summarised family-by-family (Fig C1), built into a tree (Fig C2), and as a distribution (Fig C3).
Building the tree (Fig C2): start with each sequence alone, repeatedly merge the two closest clumps, and record the identity at which each merge happened. “Closest” means the average identity across every pair spanning the two clumps (average linkage / UPGMA). 117 merges later everything is one tree.
Reading groups off the tree: draw a vertical line across it — whatever clumps remain to the left are the groups. The line is an identity threshold, so it is a number you can quote (the conventional CYP one is 40%).
The family labels are never given to any of these steps. They are used only afterwards, to check whether the result matches them.
720
n=67
725
n=31
947
n=8
90
n=5
88
n=4
974
n=2
723
n=1
CYP720 n=67
43.5
27.8
27.4
33.5
26.7
33.8
28.2
CYP725 n=31
27.8
61.6
30.8
28.0
27.7
26.5
62.2
CYP947 n=8
27.4
30.8
64.0
28.4
25.9
27.3
31.0
CYP90 n=5
33.5
28.0
28.4
34.9
27.3
33.8
26.4
CYP88 n=4
26.7
27.7
25.9
27.3
60.3
27.3
27.6
CYP974 n=2
33.8
26.5
27.3
33.8
27.3
73.3
26.7
CYP723 n=1
28.2
62.2
31.0
26.4
27.6
26.7
–
25%40%60%74%outlined diagonal = within a family · hover a cell for the pair count
Fig C1. Median % identity between families — the same 6,903 pairwise comparisons, summarised one family against another. The outlined diagonal is how similar members of a family are to each other; every other cell is family versus family. The pattern is the whole argument in one table: within a family 35–73%, between families 25–34%. The single exception is CYP725 × CYP723 at 62.2% — both are taxane enzymes, and the lone CYP723 member is ~99% identical to a CYP725 by its own annotation. CYP723 has no diagonal value because a single sequence has no within-family pair.
Fig C2. The same comparisons drawn as a tree: every one of the 118 sequences is a tick on the left, and two branches join at the % identity shown on the scale along the top — so the horizontal position is a real, readable number, not an abstract coordinate. Sequences that join far to the left are near-identical; branches that only meet at the right-hand end are barely related. The coloured bars name the family of each block. Note that the families emerge as blocks on their own — the tree was built from identity alone and never saw the labels. The dashed line is the 40% identity used as the conventional CYP family boundary. Note how much empty space surrounds it: the family blocks finish joining up well to its left, and nothing else joins until well to its right. That gap is why the exact threshold does not matter — slide the line anywhere between roughly 35% and 41% and the same blocks come out.
Fig C3. A check that the grouping is not sensitive to where the line is drawn. The tree was cut into 2, 3, 4 … 12 groups in turn, and each time we counted how many of the 118 sequences ended up in a group made mostly of their own family. At 7 groups — the number of families — 117 of 118 land correctly, and the curve is flat after that: asking for more groups only shaves single sequences off, it does not rearrange anything.
Cutting at 40% identity
Cutting the tree at the conventional CYP-family threshold (40% identity) — rather than asking for a fixed number of
groups — gives 9 clusters, and 117 of the 118 sequences sit in a cluster
made of their own family. Every cluster and what it contains:
cluster
composition
#1
CYP88 ×4
#2
CYP974 ×2
#3
CYP720 ×67
#4
CYP90 ×2
#5
CYP90 ×1
#6
CYP90 ×1
#7
CYP90 ×1
#8
CYP947 ×8
#9
CYP725 ×31, CYP723 ×1
why CYP90 fragments
Four of the nine clusters are CYP90 sequences sitting on their own. That is not a clustering failure — the five CYP90
members are genuinely that far apart. They are five different subfamilies (90A3, 90L3, 90V2, 90Y1, 90Z1) from plants
separated by hundreds of millions of years: a clubmoss (Huperzia), a fern (Ceratopteris), saffron and oil palm.
Their pairwise identities run 29.6–51.5%, and only the two from the same species (90L3 × 90Y1, 51.5%) clear
the 40% line at all — so at a 40% cut they cannot group, by construction.
This is the one place where the family label and the sequence evidence disagree, and it is worth being explicit:
CYP90 is held together by its name and its chemistry, not by sequence similarity within this set. The other evidence still
places them together — the language-model method (Fig G1) recovers all five as one cluster, and they share the
brassinosteroid pocket signature — but on identity alone they are five separate lineages.
conclusion
An algorithm that was never shown the family names recovers them almost perfectly — 117 of the 118 sequences end up grouped with their own family. Within-family identity (median 57.9%) and between-family identity (median 28.1%) form two clearly separated populations. The 7 families are therefore measurable groupings in sequence space, not only nomenclature.
D · structure evidence
One fold everywhere — shape barely separates the families
Question: The sequences differ this much — do the structures differ too?
method
Fetch an AlphaFold model for every family member carrying a UniProt accession, up to 10 per family, plus the experimental structure 8X1W.
That gives 25 structures — CYP725 11, CYP720 10, CYP947 2, CYP88 1, CYP90 1. CYP723 and CYP974 have no accession, so no model.
TM-align over all 300 pairs; each score = mean of the two chain-length normalisations.
725
11 struct
720
10 struct
947
2 struct
88
1 struct
90
1 struct
CYP725 11 struct
0.93
55 pairs
0.88
110 pairs
0.89
22 pairs
0.85
11 pairs
0.88
11 pairs
CYP720 10 struct
0.88
110 pairs
0.96
45 pairs
0.87
20 pairs
0.85
10 pairs
0.89
10 pairs
CYP947 2 struct
0.89
22 pairs
0.87
20 pairs
0.90
1 pairs
0.84
2 pairs
0.86
2 pairs
CYP88 1 struct
0.85
11 pairs
0.85
10 pairs
0.84
2 pairs
–
0.83
1 pairs
CYP90 1 struct
0.88
11 pairs
0.89
10 pairs
0.86
2 pairs
0.83
1 pairs
–
outlined diagonal = within a familyeverything else = between families– = that family has only one structure, so no within-family pair exists
Fig D1. TM-align summarised family against family — read exactly like Fig C1: the outlined diagonal is within a family, everything else is between families, and the small figure in each cell is how many structure pairs the median covers. Within a family the median TM is 0.943 (101 pairs); between families it is 0.875 (199 pairs). CYP88 and CYP90 contribute one structure each, so they have no within-family pair and their diagonal is empty.
conclusion
Every pair scores far above 0.5, so all 25 structures share one fold. Structure does separate the families slightly — median 0.943 within versus 0.875 between — but the ranges overlap heavily (0.877–0.995 against 0.828–0.915): the most dissimilar pair inside a family is less alike than the most similar pair across families. Contrast Fig C1, where within-family and between-family identity barely touch at all. Overall shape is a weak discriminator, so the functional difference has to be local. The substrate pocket is next.
E · the substrate pocket
The catalytic core is shared; the substrate pocket is not
Question: With the fold and the catalytic core shared, what physically makes one enzyme produce Taxol and another resin acid?
method
1 · Find a structure with the substrate still in it.8X3E is CYP725A4 crystallised as a holo complex — haem plus a bound taxadiene analogue (ligand A1L3H). That ligand marks where the substrate actually sits, so the pocket does not have to be predicted.
2 · Take every residue that touches it. For each amino acid in the structure, measure the shortest distance from any of its atoms to any ligand atom. Keep those within 4.5 Å — van der Waals contact range. That gives 13 residues, listed with their real distances in the table below (2.35–4.19 Å).
3 · Translate structure numbering into sequence numbering. The crystal construct starts at residue 43, so every structure number is 32 lower than the full-sequence number (Trp93 in the structure is Trp125 in the sequence).
4 · Carry those positions onto all 118 sequences. The crystallised protein is sp|Q6WG30, which is itself one of the 118 — so its residues can be located in the MAFFT alignment directly, and the alignment column of each pocket residue is read off. No structural superposition and no homology modelling is involved in this step.
5 · Read those 13 columns in every sequence. Each of the 118 now has a 13-letter pocket fingerprint. Per family, the consensus residue at each position is taken along with how much of the family agrees.
Assumption, stated: this maps one structure’s pocket onto every sequence, which holds only because all 118 share the same fold (Fig D1, TM 0.83–0.99) and the alignment is unambiguous at these columns.
pocket residue (8X3E / seq)
dist to substrate
CYP725 reference
723
947
720
88
90
974
how many of the 6 differ from it
TRP93 / 125
2.49 Å
W 65%
G 100%
W 100%
Y 93%
W 100%
Y 60%
Y 100%
4 of 6
PHE97 / 129
3.63 Å
F 74%
F 100%
L 38%
F 70%
T 100%
F 40%
I 50%
3 of 6
MET101 / 133
2.74 Å
M 48%
V 100%
I 100%
V 51%
M 100%
F 40%
I 100%
5 of 6
ALA107 / 139
2.99 Å
A 42%
V 100%
V 50%
I 51%
T 100%
L 75%
T 100%
6 of 6
LEU193 / 225
4.19 Å
L 71%
L 100%
S 62%
V 63%
N 100%
V 60%
V 50%
5 of 6
PHE197 / 229
3.59 Å
L 39%
L 100%
L 100%
L 93%
R 100%
L 40%
L 100%
1 of 6
SER270 / 302
2.43 Å
A 32%
F 100%
G 88%
N 49%
M 100%
F 40%
G 100%
6 of 6
LEU271 / 303
3.93 Å
L 58%
M 100%
L 88%
L 94%
Y 100%
L 60%
I 50%
3 of 6
HIS273 / 305
2.35 Å
H 90%
H 100%
T 50%
F 100%
N 100%
V 40%
F 100%
5 of 6
ALA274 / 306
2.65 Å
A 81%
A 100%
G 75%
A 100%
A 100%
A 80%
A 100%
1 of 6
THR278 / 310
3.74 Å
T 81%
T 100%
T 100%
T 100%
S 100%
T 60%
T 100%
1 of 6
VAL342 / 374
3.45 Å
V 48%
A 100%
I 50%
A 70%
S 100%
A 40%
G 100%
6 of 6
LEU451 / 483
2.75 Å
F 45%
F 100%
F 62%
S 49%
H 100%
T 40%
G 100%
4 of 6
Fig E1. The 13 residues that touch the bound substrate in holo 8X3E, and what each family has at that position (consensus residue + how much of the family agrees). CYP725 is the reference column — it is the taxane enzyme whose structure defined the pocket, so it is set apart rather than scored against itself. For the other six families: green = same residue as CYP725, amber = different. Blue-shaded rows are the three aromatic positions (W93 / F97 / F197) that pack against the taxane core in 8X3E. Note that these are not uniformly "taxane-only": CYP88 also carries Trp at 93, and even within CYP725 position 197 is predominantly Leu rather than Phe, so the cage is a tendency across the family rather than a strict signature.
conclusion
13 of 13 pocket positions differ between families. The two least variable positions belong to the catalytic machinery rather than the substrate cage: Thr278 is the I-helix threonine and Ala274 is a small floor residue. The residues that contact the substrate do change — for example Trp93 becomes Tyr in CYP720 (93%) and CYP90 (60%), and Phe97 becomes Thr in CYP88. The catalytic chemistry is shared while the substrate cage is family-specific, which provides a structural basis for the functional split.
E → F · how the five groups were drawn
The five pocket signatures, side by side
Question: Where exactly do the five groups come from?
short answer, before the figure
CYP720 67
→ group B resin acid
kept as it is
CYP88 4
→ group C gibberellin
kept as it is
CYP90 5
→ group D brassinosteroid
kept as it is
CYP974 2
→ group E uncharacterised
kept as it is
CYP725 31 CYP723 1 CYP947 8
→ group A taxane
the only merge
7 families − 2 merges = 5 groups.
Fig E2. One sequence logo per functional group, over the 13 pocket positions. Letter height = how much of that group carries that residue, so a single tall letter means the whole group agrees and a stack means it varies. Colour = amino-acid chemistry (violet aromatic, grey hydrophobic, green polar, blue basic, red acidic). The row underneath is each group’s consensus, shaded by how strongly the group holds it. Read down a column to compare the five groups at one position; read across a row to see one group’s whole pocket.
Read it down the columns. Each group holds its own pattern, and the differences are chemical, not cosmetic:
group
pocket consensus
what stands out
A taxane
WFIVLLGLHATVF
the only group with Trp at the first position and His in the middle — the aromatic cage that grips the taxane ring
B resin acid
YFVIVLFLFATAS
Trp→Tyr, His→Phe: a smaller, flatter, fully hydrophobic cavity
C gibberellin
WTMTNRMYNASSH
the odd one out — Arg, Asn, His where every other group is hydrophobic: a polar, charged pocket for a carboxylic-acid substrate
D brassinosteroid
YFFLVLFLVATAT
closest to B, differing at the positions that set cavity length — consistent with a long sterol side chain
E fern CYP
YIITVLGIFATGG
Gly where others carry bulky residues, so the widest cavity of the five; no characterised member, so no substrate is claimed
conclusion
Group C is separable on the pocket alone — a charged cavity among hydrophobic ones. Groups A and B differ
consistently at the aromatic cage. D and E sit close to B and could not be told apart on these 13 residues alone, which is why they are held apart on family membership and on chemistry from the literature, not on the pocket.
Everything up to here says “different substrates”; nothing up to here says which. That comes next, with sources.
F · the answer, with sources
Five substrate-defined functional groups
method
Reaction quoted from a reviewed (Swiss-Prot) entry where one exists, with EC number and Rhea id, plus EuropePMC literature.
Where the set holds only unreviewed (TrEMBL) entries, the reaction is taken from a characterised relative — and said so.
Why this entry: This exact protein IS in the 118 (sp|Q6WG30) and IS the 8X1W/8X3E structure. Swiss-Prot reviewed.
Caveat: CYP947 is a Taxus/gymnosperm taxane-pathway candidate: it forms its own clean cluster (n=8) and has no characterised member, so its placement in group A is by pathway context, not by a measured reaction.
Caveat: No Swiss-Prot entry for the CYP720B members in this set (all TrEMBL, no EC). The reaction below is from the literature on characterised conifer CYP720B enzymes.
Why this entry: Swiss-Prot KAO1 (CYP88A3, Arabidopsis) - the characterised CYP88A ent-kaurenoic acid oxidase. The 118 contain LjKAO (A0A0B6VL99, TrEMBL) whose name matches this activity.
Caveat: The set's members are TrEMBL (LjKAO, A0A0B6VL99). Reaction taken from the characterised Arabidopsis KAO1.
Why this entry: Swiss-Prot CPD/CYP90A1 (Arabidopsis). The 118 contain CYP90A3 (A0A6I9R9G3, TrEMBL, auto-annotated as 3beta,22alpha-dihydroxysteroid 3-dehydrogenase isoform) - same family, same chemistry class.
Caveat: Set members are TrEMBL; reaction from the characterised Arabidopsis CPD/CYP90A1. Note these are five different subfamilies (90A3/90L3/90V2/90Y1/90Z1) spanning clubmoss, fern, saffron and oil palm, only 29.6-51.5% identical to each other - grouped by name and chemistry, not by sequence similarity.
Caveat: EuropePMC returns ZERO publications for CYP974. No substrate, no reaction, no EC. Flagged as unknown - the honest answer.
EuropePMC: 0 publications
source / provenance
UniProt REST (rest.uniprot.org) → reaction_sources.json; EuropePMC REST search → lit_records.json. Every accession and PMID above is a live link.
conclusion
Only one of the 118 sequences (sp|Q6WG30, the CYP725A4 that is also the crystal structure) has a reviewed, EC-numbered reaction of its own. The remainder are TrEMBL entries. The split itself rests on the measurements in sections C, D and E; the substrate assignments for groups B/C/D are inherited from characterised relatives, and group E carries no assignment.
G · independent cross-check
A method that never saw the alignment finds the same groups
Question: Everything so far derives from one alignment — would an alignment-free method find the same groups?
Uses neither the MSA %identity nor any 3D structure — an independent route to the same question.
cluster
n
dominant
composition
functional group
#1
3
CYP88
CYP88 3/3
C gibberellin
#2
6
CYP947
CYP947 6/6
E uncharacterized
#3
27
CYP725
CYP725 25 + CYP723 1 + CYP720 1
A taxane (725+723 together)
#4
2
CYP720
CYP720 2/2
B resin-acid
#5
6
CYP720
CYP720 6/6
B resin-acid
#6
7
CYP90
CYP90 5 + CYP974 2
D brassinosteroid (+CYP974)
#7
33
CYP720
CYP720 33/33
B resin-acid
Fig G1. Cutting the PLMView tree into 7 clusters, versus the given families. Scored on the 84/118 sequences that could be assigned a family from the header text alone (80 carry an explicit CYPxxx name, 4 more were resolved from the UniProt description); the remaining 34 were held out of the metric.
Fig G2. The language model’s own view of the 118, coloured by the functional group they were assigned in section F. This uses PLMView’s embedding representation directly — no alignment, no structure, and the group labels were applied only after the coordinates existed. If the grouping were an artefact of the alignment, these colours would be scattered at random. They are not: each group occupies its own region.
conclusion
An orthogonal method reproduces the same clusters: CYP88 and CYP947 clean, CYP725+CYP723 together (both taxane), CYP90 with CYP974, and CYP720 splitting into sub-clades. The grouping is reproduced independently of the alignment and of any single algorithm.
NO — not one function · 5 groups
The 118 are all genuine plant cytochrome-P450 monooxygenases: one fold (TM 0.83–0.91),
one catalytic core (haem Cys 118/118). They are not one enzyme.
They divide into five substrate-defined groups: taxane / Taxol,
diterpene resin acids, gibberellin,
brassinosteroid, and one uncharacterised fern CYP.
Four independent lines agree: unsupervised clustering reproduces the given families (117 of 118 sequences grouped with their own);
structure explicitly fails to separate them, which localises the difference; 13/13 substrate-pocket
positions differ between families while the catalytic residues do not; and a protein-language-model tree that never saw the alignment
recovers the same clusters (group purity 0.964).
Known limits: only 1 of the 118 has a reviewed EC-numbered reaction of its own; groups B/C/D inherit their
substrate assignment from characterised relatives; CYP974 (n=2) has zero publications and no assignment; CYP90 is internally divergent
(within-family median identity 34.9%); 14 near-identical pairs were kept in, medians used throughout.