Lesson 11 of 20 · 8 min
EMBL-EBI, UniProt and Structures
By the end of this short overview you will know the European half of the database world — five resources hosted around EMBL-EBI and its partners — and how each one takes the same BRCA1 you met on the NCBI side and describes it from a fresh angle.
The other hub, not a rival
In The Two Hubs: NCBI Entrez and EMBL-EBI you met EMBL-EBI (the European Molecular Biology Laboratory's European Bioinformatics Institute) as NCBI's counterpart. This module opens it up. The key idea to carry in: NCBI and EBI are not competitors racing to store the same thing twice. For raw sequence they are mirrors — they share submissions daily through the INSDC agreement you saw in Project and Sample Accessions Across INSDC, so a read deposited in Europe appears in America and vice versa. For everything built on top of that sequence, they diverge, and that is where EBI's distinct resources earn their place.
Deposition versus knowledgebase, one more time
The single most useful distinction for reading any of these five resources is the one from Primary vs Derived Databases. A deposition (archival) database stores exactly what a submitter sent, forever, unchanged — it answers "what evidence exists?". A knowledgebase is expert-curated and rewritten as understanding improves — it answers "what do we currently believe is true?". Every EBI resource sits on one side of that line, and knowing which side tells you how much to trust a given field.
| Resource | Type | Answers about BRCA1 |
|---|---|---|
| ENA | deposition | what nucleotide sequence and reads were submitted? |
| Ensembl | derived | where does the gene sit on the genome, and its transcripts? |
| UniProt | knowledgebase | what does the protein do? |
| PDB | deposition | what 3D structures were solved? |
| InterPro / Pfam | derived | what domains and families is it built from? |
Following BRCA1 across the five
Each of these resources holds one BRCA1 record you can look up today. ENA is Europe's copy of the nucleotide archive. Ensembl carries the gene as ENSG00000012048, whose canonical transcript ENST00000357654 is sequence-identical to RefSeq's NM_007294 — the two institutes agreeing on one reference transcript per gene even though they store and annotate it separately. UniProt is where the gene becomes a protein: entry P38398 (BRCA1_HUMAN), Swiss-Prot reviewed. PDB holds pieces of that protein solved in atomic detail — 1JNX is the C-terminal BRCT repeat region, 1JM7 the BRCA1/BARD1 RING-domain heterodimer. InterPro and Pfam name the reusable parts inside P38398: the RING finger and the tandem BRCT domains.
Try it yourself: open UniProt entry P38398 and scroll to its cross-references section. In one record you will see it point out to Ensembl (ENSG00000012048), to PDB (1JNX, 1JM7), and to InterPro domain signatures — the whole module, wired together from a single page. Not every database links back the same way, though: RefSeq's own protein record for this gene mentions UniProt only in a prose annotation, not as a structured pointer, so always check the exact form a cross-reference takes before you script against it.
Five resources, one BRCA1, two institutes that agree on the sequence and specialise everywhere else. That is the shape of the European side — and the rest of this module walks each stop in turn.