This page documents the full pipeline behind every value in SurfPotDB, and how to reuse the downloadable files. For the overview and the dataset's scope, see the About page.

What each structure ships

For each computed structure we report four electrostatic energy contributions (in kT), the range of the molecular-surface potential (in kT/e), and two downloadable files: the surface potential mesh (.vtp, for ParaView/PyVista) and the protonated structure with charges and radii (.pqr, for PyMOL/VMD).

How structures were selected

Solving the Poisson–Boltzmann equation on the entire PDB would repeat work on near-identical entries. We reduced the PDB to a non-redundant set of representatives through the following funnel:

1. Download Entire RCSB PDB, keeping only structures with fewer than 1000 protein residues.
2. Filter Removed entries released before year 2010 and membrane proteins (via the RCSB API).
3. Cluster Grouped by sequence identity using RCSB's pre-computed clusters at 90% identity: 126,139 proteins → 32,861 clusters.
4. Representative One representative per cluster, chosen by best resolution (ties broken by larger residue count).
5. This release Of the cluster representatives, computed the ligand- and ion-free ones (protein-only): 8,571 structures, standing for 19,017 PDB entries. Where a representative also carried nucleic acid or other non-protein chains, only the protein was kept for the calculation.

Calculations run on the representatives only, and a query for any member of a cluster is resolved to its representative (see Cluster resolution below), so a single computed structure answers for its whole cluster. This first release is deliberately restricted to protein-only inputs: every computed structure is free of bound ligands and ions, so the electrostatics rest on a single, well-defined assignment of charges and radii, with no ambiguous parametrization of cofactors. Representatives that carry ligands or ions are deferred to forthcoming releases, once a robust parametrization of their charges and radii is in place.

How the PDBs were protonated

Poisson–Boltzmann electrostatics depends critically on the protonation state of titratable residues, which crystal structures do not provide. We assigned protonation with MCCE4 (Multi-Conformation Continuum Electrostatics):

This protonated PQR is the input to the final electrostatics calculation and is the .pqr file offered for download. Because it carries the per-atom charges and radii assigned by MCCE, the downloaded .pqr lets you recover the exact protonation state (which titratable groups are protonated/deprotonated, and in which conformer) that MCCE determined at pH 7.0, the ground-truth protonation behind every number on this site.

Electrostatics calculation (NextGenPB)

Each protonated structure is solved with NextGenPB as a standalone linearized Poisson–Boltzmann calculation, using a single, documented parameter set: solute dielectric 2, solvent dielectric 80, ionic strength 0.145 M, temperature 298.15 K. The run produces the molecular-surface potential map (.vtp) and the energy decomposition reported here (polarization, ionic, Coulombic, total).

The four energy contributions are given in kT. To convert to kcal/mol, multiply by 0.593 (1 kT at 298.15 K = 0.593 kcal/mol).

Potential anywhere inside the protein

The published .vtp meshes are enriched: for each surface node they store not only the potential phi and the normal, but also the polarization charge q_pol (the flux of the dielectric displacement through the surface). That extra field is what makes the surface solution reusable: combined with the atomic charges from the .pqr, it lets you reconstruct the electrostatic potential (and field) at any point inside the structure, not just on the surface.

Concretely, a single published .vtp reconstructs the potential anywhere from the files alone, including at the atom centers, where each atom's own (infinite) Coulomb self-term is automatically excluded so the values stay physically meaningful right at the nuclei.

Values are in reduced units kT/e (≈ 25.7 mV at 298.15 K), matching the phi array in the .vtp.

A small standalone tool, ngpb_potential.py (documentation), does exactly this: given a structure's .vtp and .pqr alone, it reconstructs the electrostatic potential and field anywhere inside the protein, on a grid, along a path, at arbitrary points, or at the atom centers (with the Coulomb self-term excluded). No solver re-run needed.

# potential at atom centers (self-energy excluded), from the published files
python ngpb_potential.py 7kr0.vtp 7kr0.pqr

How to use the data

Every structure ships two files, plus a dataset-wide index. Search a PDB code on the Database page to get its download links.

Cluster resolution

When you search a PDB code on the Database page:

Dataset diversity in detail

The 8,571 computed representatives are non-redundant by construction (one per 90%-identity sequence cluster). Annotations pulled from the RCSB PDB API show the set spans a broad swath of the known protein universe.

Functional classification (Enzyme Commission)

1,405 representatives (16.4%) carry an enzyme (EC) annotation, covering all seven top-level EC classes:

Oxidoreductases (EC 1)1481.7%
Transferases (EC 2)5126.0%
Hydrolases (EC 3)5336.2%
Lyases (EC 4)831.0%
Isomerases (EC 5)1011.2%
Ligases (EC 6)770.9%
Translocases (EC 7)50.1%

Per-class counts sum to slightly more than 1,405 because a few multi-function enzymes carry EC numbers in more than one top-level class and are counted in each.

Functional classification beyond enzymes (GO molecular function)

EC numbers only describe enzymes, so they leave the ~84% non-enzyme majority undifferentiated. To break that down we roll each structure's Gene Ontology molecular-function annotations up to their top-level GO categories. This labels non-enzymes too — receptors, transporters, transcription factors, structural proteins — in one scheme that also recovers more enzymes than carry an explicit EC number (catalytic activity ≈ 30% vs 16% with an EC).

Enzyme (catalytic activity)2,58730.2%
Binding only1,13613.3%
Regulator / inhibitor98011.4%
Structural molecule5536.5%
Transcription regulator3904.6%
Molecular adaptor3894.5%
Transporter3063.6%
Receptor (signaling)2262.6%
Toxin2142.5%
Other molecular function500.6%
Antioxidant490.6%
Electron transfer270.3%
Translation regulator260.3%
No molecular-function annotation2,96834.6%

A structure can carry more than one molecular function (e.g. an enzyme that also binds nucleic acids), so the categories overlap and the percentages do not sum to 100%. The generic "binding" category is shown only for structures with no more specific function. These same categories drive the Protein class filter on the database page.

Structural classification (CATH)

The annotated structures cover 1,409 distinct CATH superfamilies across every major fold class:

Mainly Alpha1,01611.9%
Mainly Beta1,30315.2%
Alpha/Beta1,72620.1%
Few secondary structures440.5%
Special750.9%

Taxonomic diversity

Source organisms span all four superkingdoms of cellular and viral life:

Eukaryota4,49752.5%
Bacteria2,83033.0%
Viruses8259.6%
Archaea1972.3%
Unassigned / synthetic5716.7%
Functional diversity by GO molecular-function class
Functional diversity by GO molecular function — non-enzymes (receptors, transporters, transcription factors, …) resolved alongside enzymes.
Taxonomic diversity by superkingdom
Taxonomic diversity by superkingdom.

CATH and Pfam do not yet annotate many recent entries (notably cryo-EM structures), so dataset breadth is best read from the counts of distinct families (2,644 Pfam, 1,409 CATH) rather than from per-class percentages. Non-redundancy is guaranteed by the sequence clustering (one representative per cluster).

← Back to About  ·  Go to the database →