PDB vs mmCIF: Is the Legacy PDB File Format Becoming Obsolete?
Understand why PDBx/mmCIF is replacing legacy PDB as the complete archive format, what changes in 2027, how AlphaFold fits, and when conversion remains appropriate.
The question "Is PDB obsolete?" mixes three different things:
the global Protein Data Bank archive, a four-character PDB entry
identifier, and the legacy text format whose records begin with
labels such as ATOM and HETATM. The archive
remains foundational. The identifier scheme and file format are
the layers changing.
This distinction matters now because structural pipelines are
encountering more large complexes, richer metadata, extended PDB
identifiers, and predicted structures distributed as CIF files.
Treating every .cif as an inconvenience to convert can
discard information or create identifiers that no longer map
cleanly to the source.
PDB is a database, an identifier, and a file format
Separate the terms before comparing formats
| Term | What it means | Current status |
|---|---|---|
| Protein Data Bank | The international archive of experimentally determined macromolecular structures managed through the wwPDB partnership. | Active and expanding; not made obsolete by a format transition. |
| Legacy PDB format | A fixed-column, record-oriented text representation defined by the frozen version 3.30 specification. | Still widely read, but no longer extended to support new archive content [2]. |
| PDBx/mmCIF | A dictionary-backed tabular and key-value representation for coordinates, entities, experiments, identifiers, and relationships. | Primary PDB archive standard since 2014 [2]. |
| CIF | The broader Crystallographic Information Framework syntax family. | In macromolecular PDB and modern AlphaFold output, a .cif file commonly contains PDBx/mmCIF data; verify the dictionary and source. |
| PDB entry ID | The identifier for an archive entry, historically four characters. | Transitions to a 12-character extended form for newly issued IDs on July 21, 2027 [1]. |
Why the legacy PDB format reached its limits
Fixed fields versus extensible data
| Requirement | Legacy PDB | PDBx/mmCIF |
|---|---|---|
| Large structures | Cannot fully represent entries with more than 62 chains or 99,999 ATOM records in one standard file [2,3]. | Does not impose those fixed atom, residue, or chain-count limits [2]. |
| Identifiers | Column widths constrain chain, atom, residue, and entry identifiers. | Whitespace-delimited values and named data items support longer identifiers. |
| Metadata | New content must fit frozen record types and columns; the format is no longer extended [2]. | Dictionary-defined categories can represent richer experimental, biological, and validation information. |
| Relationships | Many relationships must be inferred from record conventions and special cases. | Categories and data items have explicit, machine-readable definitions and relationships [2]. |
| Parsing | Column positions and historical exceptions require format-specific logic. | Key-value and loop tables follow a defined grammar and dictionary. |
| Archive completeness | Some large entries are unavailable as one complete legacy PDB file or are distributed as best-effort bundles [2,3]. | One archival file can preserve the complete entry. |
PDBx/mmCIF does not make coordinates intrinsically more accurate. Its advantage is representational: it can preserve more of the identifiers, relationships, metadata, and complete structure in one extensible record.
What changes on July 21, 2027
The transition timeline
| Period | Archive behavior | Workflow implication |
|---|---|---|
| Since 2014 | PDBx/mmCIF is the standard PDB archive format and the basis for wwPDB data processing and annotation [2]. | Treat mmCIF support as a current requirement, not a future experiment. |
| Before July 21, 2027 | Legacy four-character IDs and many corresponding PDB-format files remain available alongside PDBx/mmCIF. | Use this period to test parsers, filenames, databases, reports, and downstream tools with extended IDs and mmCIF. |
| From July 21, 2027 | OneDep issues 12-character IDs for new entries; those entries are released in PDBx/mmCIF and PDBML, not legacy PDB [1]. | A PDB-only ingestion pipeline will be unable to consume every newly released entry. |
| Existing entries | Older entries receive an extended-ID representation that prepends pdb_0000 to the historical four-character ID; their PDB DOIs remain persistent [1]. | Preserve both source identifier and DOI rather than deriving identity only from a four-character filename. |
RCSB PDB has also published a one-year transition notice urging users and software developers to adopt the reorganized archive, extended identifiers, and PDBx/mmCIF before the deadline [4].
Why AlphaFold makes the transition more visible
AlphaFold output is not one universal format
| Source | Documented coordinate output | Practical consequence |
|---|---|---|
| AlphaFold 2 workflow | EMBL-EBI training documents predicted structures saved in PDB and mmCIF, with confidence information associated with the outputs [6]. | Preserve the confidence files and source metadata even if a PDB copy is convenient. |
| AlphaFold 3 source code | The official output documentation states that the top-ranked model is written as <job_name>_model.cif in mmCIF and that PDB output is not provided [5]. | Downstream software must read mmCIF or use an explicit, validated conversion. |
| Archive or prediction service | Packaging, confidence fields, naming, and companion files depend on the provider and version. | Do not infer provenance, experimental method, or confidence semantics from the extension alone. |
An AlphaFold .cif output does not imply that the
legacy PDB extension is invalid for all predicted models. It shows
that modern structural tools increasingly choose a representation
that can carry larger systems, richer entity definitions, and
more explicit identifiers.
Use mmCIF as the master and PDB as a compatibility export
A practical format policy
| Workflow stage | Preferred record | Reason |
|---|---|---|
| Acquisition | Original PDBx/mmCIF plus validation and companion files. | Preserves archive identifiers, entities, assemblies, metadata, and complete large structures. |
| Internal storage | Versioned original file with checksum, source URL, entry ID, DOI, retrieval date, assembly and model selection. | Separates source evidence from later transformations. |
| Analysis | Use a parser that understands the relevant PDBx/mmCIF dictionary and identifier namespaces. | A text reader that extracts coordinates alone may lose entity or author-versus-label mappings. |
| Legacy-tool handoff | Generate PDB only when the receiving application requires it. | Limits conversion to a visible compatibility boundary. |
| Reporting | Record original and derived filenames, tool and version, conversion command, warnings, identifier map, and checksum. | Makes the last-mile transformation reviewable and repeatable. |
A compatibility PDB file can be perfectly reasonable for a small receptor and a well-understood downstream tool. The mistake is treating the conversion as lossless without checking whether the source exceeds PDB field widths, uses long or numerous chain IDs, depends on label and author identifier mappings, contains multiple models, or carries assembly and experimental metadata that the receiving tool will not preserve.
Keep receptor selection separate from format selection. The receptor-selection guide addresses biological state, pocket evidence, and experimental versus predicted structures; choosing mmCIF does not resolve those scientific questions automatically.
Validate every conversion instead of trusting the extension
Conversion checks that prevent silent structural drift
| Check | Compare before and after | Failure signal |
|---|---|---|
| Completeness | Models, chains, residues, atoms, alt locations, waters, ligands, ions, and cofactors. | Counts change without a declared filter or the structure is split unexpectedly. |
| Identifiers | Entry, entity, label and author chain IDs, residue numbers, insertion codes, atom names, and ligand IDs. | Truncation, collision, renumbering, or ambiguous one-character chain mapping. |
| Coordinates | Matched-atom coordinates, occupancy, B factors or confidence-field use, and model number. | Atoms cannot be mapped one to one or coordinates change beyond formatting tolerance. |
| Assembly | Asymmetric unit versus biological assembly and applied transformations. | The converted file represents a different biological object. |
| Connectivity and chemistry | Polymer breaks, disulfides, metals, modified residues, covalent links, and ligand identity. | Required bonds or components disappear or must be guessed downstream. |
| Provenance | Original checksum, conversion tool and version, command, warnings, and derived checksum. | The PDB file cannot be traced back to one source record and transformation. |
Where MolNexus fits
MolNexus accepts receptor input as PDB, CIF/mmCIF, or a direct RCSB PDB identifier. This lets a Windows researcher bring a modern archive or predicted-structure file into the same local receptor-preparation, interaction-box, AutoDock Vina, pose review, history, and export workflow without first creating an unmanaged PDB copy solely for ingestion.
MolNexus does not determine the correct biological assembly, validate an experimental structure, certify an AlphaFold binding pocket, guarantee lossless conversion, or make docking evidence equivalent to binding evidence. The source file, structure selection, receptor-preparation decisions, and validation record remain the user's scientific responsibility.
MolNexus 0.1.1 is an available Windows 10/11 64-bit desktop application. A free Windows trial limited to two ligand docking runs and the US$499 one-time perpetual license are available now.
A migration checklist for structural pipelines
- Inventory: find readers, writers, filename rules, database fields, regular expressions, and reports that assume four-character IDs or
.pdb. - Add native mmCIF tests: include ordinary entries, large assemblies, multiple models, long chain IDs, modified residues, ligands, and extended-ID examples.
- Preserve namespaces: decide how label and author identifiers are represented internally and in exports.
- Separate source from derivative: retain the original PDBx/mmCIF and write converted PDB files to a documented derived layer.
- Validate compatibility exports: compare counts, identifiers, coordinates, assembly, chemistry, and warnings.
- Remove hidden assumptions: do not slice entry IDs, chain IDs, or residue fields to legacy widths in databases or user interfaces.
- Rehearse before 2027: run the complete acquisition-to-analysis workflow against the beta archive and current mmCIF-only structures [1,4].
The safest policy is simple: preserve the richest authoritative source, make compatibility transformations explicit, and test the scientific object after every transformation.
Frequently asked questions
References
- Worldwide Protein Data Bank. Transition to Extended IDs and the PDB Beta Archive wwPDB documentation (2026) Official July 21, 2027 transition date, extended-ID scheme, archive reorganization, and new-entry format policy.
- Worldwide Protein Data Bank. PDBx/mmCIF General FAQ wwPDB PDBx/mmCIF documentation (2026) Official comparison of the frozen legacy format with the dictionary-backed PDBx/mmCIF archive standard and its limits.
- RCSB Protein Data Bank. Structures Without Legacy PDB Format Files RCSB PDB help documentation (2026) Official examples and conditions under which an entry cannot be represented as one legacy PDB-format file.
- RCSB Protein Data Bank. Big PDB Changes Are One Year Away: Extended PDB IDs, the PDB Beta Archive, and PDBx/mmCIF RCSB PDB News (2026) Current official transition notice for users and software developers preparing for the 2027 archive change.
- Google DeepMind. AlphaFold 3 Output Official AlphaFold 3 source-code documentation (2026) Official mmCIF output structure, naming, and explicit absence of PDB output from the AlphaFold 3 source-code pipeline.
- EMBL-European Bioinformatics Institute. AlphaFold2 Inputs and Outputs: Recap EMBL-EBI Training (2026) Authoritative training resource documenting AlphaFold 2 PDB and mmCIF coordinate outputs and associated confidence information.