Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON) - #4712

Open
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif
Open

Implementing gemmi-based mmcif reader (with easy extension to PDB/PDBx and mmJSON)#4712
marinegor wants to merge 154 commits into
MDAnalysis:developfrom
marinegor:feature/mmcif

Conversation

@marinegor

@marinegormarinegor commented Sep 20, 2024

Copy link
Copy Markdown
Contributor

Fixes#2367 and also extends #4303 and solves #5089

Changes made in this Pull Request:

  • uses gemmi library (link) to parse mmcif files
  • adds a class MMCIFReader(base.SingleFrameReaderBase) and class MMCIFParser(TopologyReaderBase) classes for that

As a bonus, this implementation would potentially allow to read any of the gemmi-supported formats (source):

  • mmCIF (PDBx/mmCIF),
  • PDB (with popular extensions),
  • mmJSON

Also, this (with slight modifications) also would allow reading mmcif with multiple models sharing the same topology, as well as more feature-rich parsing of PDBs (the same code without changes can be used for parsing altlocs, charges, etc, from all of these formats).

However, I'm slightly lost on what's to be done next for this PR to be merged, so I'm asking if someone could help me navigate here (tagging @richardjgowers here as author of original PDBx implementation 4303).

PR Checklist

  • Tests?
  • Docs?
  • CHANGELOG updated?
  • Issue raised/referenced?

Developers certificate of origin


📚 Documentation preview 📚: https://mdanalysis--4712.org.readthedocs.build/en/4712/

@pep8speaks

pep8speaks commented Sep 20, 2024

Copy link
Copy Markdown

Hello @marinegor! Thanks for updating this PR. We checked the lines you've touched for PEP 8 issues, and found:

Line 28:80: E501 line too long (84 > 79 characters)
Line 41:80: E501 line too long (85 > 79 characters)
Line 42:80: E501 line too long (93 > 79 characters)
Line 61:80: E501 line too long (104 > 79 characters)
Line 65:80: E501 line too long (87 > 79 characters)
Line 67:80: E501 line too long (107 > 79 characters)

Line 2:24: W291 trailing whitespace
Line 60:80: E501 line too long (111 > 79 characters)
Line 72:80: E501 line too long (123 > 79 characters)
Line 82:80: E501 line too long (122 > 79 characters)
Line 106:80: E501 line too long (108 > 79 characters)
Line 113:80: E501 line too long (80 > 79 characters)
Line 128:80: E501 line too long (91 > 79 characters)
Line 175:80: E501 line too long (126 > 79 characters)
Line 185:80: E501 line too long (125 > 79 characters)
Line 224:80: E501 line too long (126 > 79 characters)
Line 242:80: E501 line too long (140 > 79 characters)
Line 281:80: E501 line too long (87 > 79 characters)
Line 292:80: E501 line too long (90 > 79 characters)

Line 56:80: E501 line too long (80 > 79 characters)
Line 57:80: E501 line too long (84 > 79 characters)

Line 335:26: W292 no newline at end of file

Line 48:80: E501 line too long (103 > 79 characters)
Line 81:80: E501 line too long (80 > 79 characters)
Line 97:80: E501 line too long (86 > 79 characters)
Line 271:80: E501 line too long (90 > 79 characters)
Line 340:80: E501 line too long (104 > 79 characters)
Line 387:80: E501 line too long (83 > 79 characters)
Line 436:80: E501 line too long (80 > 79 characters)
Line 463:80: E501 line too long (80 > 79 characters)
Line 481:80: E501 line too long (80 > 79 characters)
Line 493:80: E501 line too long (80 > 79 characters)
Line 494:80: E501 line too long (80 > 79 characters)
Line 497:80: E501 line too long (83 > 79 characters)
Line 498:80: E501 line too long (86 > 79 characters)
Line 546:80: E501 line too long (82 > 79 characters)
Line 547:80: E501 line too long (82 > 79 characters)
Line 549:80: E501 line too long (88 > 79 characters)
Line 551:80: E501 line too long (88 > 79 characters)
Line 552:80: E501 line too long (81 > 79 characters)
Line 777:80: E501 line too long (81 > 79 characters)
Line 778:80: E501 line too long (87 > 79 characters)
Line 779:80: E501 line too long (84 > 79 characters)
Line 780:80: E501 line too long (85 > 79 characters)
Line 781:80: E501 line too long (83 > 79 characters)

Comment last updated at 2024-10-25 11:17:29 UTC

@github-actions

github-actionsBot commented Sep 20, 2024

Copy link
Copy Markdown

Linter Bot Results:

Hi @marinegor! Thanks for making this PR. We linted your code and found the following:

Some issues were found with the formatting of your code.

Code LocationOutcome
main package⚠️ Possible failure
testsuite⚠️ Possible failure

Please have a look at the darker-main-code and darker-test-code steps here for more details: https://github.com/MDAnalysis/mdanalysis/actions/runs/11148966346/job/30986736623


Please note: The black linter is purely informational, you can safely ignore these outcomes if there are no flake8 failures!

@richardjgowersrichardjgowers left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good so far, will require a small test file to check reader/parser halves.

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
Comment threadpackage/pyproject.toml Outdated
@IAlibay

Copy link
Copy Markdown
Member

@IAlibay your comment had been addressed and the gemmi reader is also tested in Windows on azure. Could you updated your review, please? Otherwise I'll dismiss it in a day in an effort to merge this important new feature.

(I am not sure what the linter wants, I'll try to make it happy.)

@orbeckst can you give me until Tuesday please? I agree it's an important feature and I do want to review it, however there's some other high priority items within MDAnalysis that needs addressing first.

Either way, I'm not going to be releasing a be version of MDAnalysis until after the 17th - so it not being merged until then isn't an issue.

@orbeckst

Copy link
Copy Markdown
Member

Ok

@orbeckstorbeckst mentioned this pull request Jul 15, 2026
5 tasks

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the very brief review, mostly cleaning things to do.

Comment threadpackage/pyproject.toml Outdated
Comment threadpackage/pyproject.toml Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP.cif
Comment threadtestsuite/MDAnalysisTests/data/mmcif/1YJP_invalid.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif Outdated
Comment threadtestsuite/MDAnalysisTests/data/mmcif/7ETN.cif.gz
Comment threadtestsuite/MDAnalysisTests/data/mmcif/multimodel_warning.cif Outdated
Comment threadpackage/MDAnalysis/topology/__init__.py
@ianmkenney

ianmkenney commented Jul 22, 2026

Copy link
Copy Markdown
Member

Atoms without altLocs in the MMCIF will have null bytes assigned to their altLoc attribute. When writing the structure out to PDB, it includes those null bytes, in violation of the PDB spec (not that this have stopped anyone before). Not fully sure of the consequences but figured I'd raise it before it's merged.

Altered tail of 1BD2.cif to include non-null altLocs:

HETATM 6376 O O . HOH I 6 . ? -10.858 27.299 2.616 1.00 63.18 ? 257 HOH E O 1 HETATM 6377 O O . HOH I 6 . ? -11.731 43.369 -12.769 1.00 34.61 ? 258 HOH E O 1 HETATM 6378 O O A HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
HETATM 6379 O O B HOH I 6 . ? 3.361 39.263 -10.471 1.00 56.34 ? 259 HOH E O 1
#

Tail of reconstructed PDB (but \0 is actually a null byte)

HETATM 6376 O \0HOH E 257 -10.858 27.299 2.616 1.00 63.18 E O HETATM 6377 O \0HOH E 258 -11.731 43.369 -12.769 1.00 34.61 E O HETATM 6378 O AHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O HETATM 6379 O BHOH E 259 3.361 39.263 -10.471 1.00 56.34 E O ENDMDL
END

You can recreate this with the following:

https://gist.github.com/ianmkenney/624300e0395bcf2c918c2fa8b92fd86e

Just clone MDA and check out the PR head next to the Makefile and run make test.

Comment threadpackage/MDAnalysis/topology/MMCIFParser.py Outdated
@BradyAJohnston

Copy link
Copy Markdown
Member

Everything should now be addressed - but I can't figure out why codecov is reporting such low test coverage because it's got 100% coverage when I run it locally.

@orbeckst

Copy link
Copy Markdown
Member

Current coverage here is sensible: project — 93.88% and patch — 97.02% of diff hit, so I do not see any issues. Perhaps it was just slow to update?

Now just waiting for @IAlibay .

@IAlibay

Copy link
Copy Markdown
Member

Thanks for the ping (sorry I didn't see the earlier changes, travelling for conferences).

I'll aim to review by end of the week.

@orbeckstorbeckst left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very briefly skimmed and I might well be overlooking something so, just as comments:

Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py
Comment threadpackage/MDAnalysis/coordinates/MMCIF.py
Comment threadpackage/MDAnalysis/topology/MMCIFParser.py
Comment threadpackage/MDAnalysis/coordinates/__init__.py

@IAlibayIAlibay left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's unfortunately quite a few things that still need addressing.

Please feel free to open follow-up issues for some of these.

Comment thread.github/actions/setup-deps/action.yaml
Comment thread.github/actions/setup-deps/action.yaml Outdated
[tool.setuptools.packages.find]
namespaces = false

[tool.setuptools.package-data]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing data/mmcif/*.gz and data/mmcif/*.cif entries.

Comment on lines 138 to 139
| MDAnalysis/coordinates/MMCIF\.py
| MDAnalysis/topology/MMCIFParser\.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should these be excluded from black, is the black to un-exclude these in a follow-up PR?

DSSP = (_data_ref / "dssp").as_posix()

# MMCIF data: valid structures from RCSB and generated by Biopython
MMCIF = (_data_ref / "mmcif").as_posix()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks the convention we use everywhere else for datafiles (and I notice that the DSSP one does too...).
It's fine if this ships as-is, but I would like this to be raised as an issue to be fixed in a follow-up PR please.

Comment threadpackage/MDAnalysis/topology/PDBParser.py
Comment threadpackage/CHANGELOG
"pyedr>=0.7.0",
"pytng>=0.2.3",
"gsd>3.0.0",
"gemmi>=0.7.3", # for mmcif format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"gemmi>=0.7.3", # for mmcif format
"gemmi>=0.7.3", # for mmcif format

[nit] 2 spaces to keep with standard python formatting, either that or remove the comment - we don't do this for any other file format.

serials.append(atom.serial)
names.append(atom.name)
chainids.append(chain.name)
elements.append(atom.element.name)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDB parser, we:

  • Check if the capitalized element is in SYMB2Z
  • Store the capitalized element

Do we need to do either of these here?

AltLocs(altlocs),
Atomids(serials),
Atomnames(names),
Atomtypes(names),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the PDBParser, we use elements for AtomType, not names. Should this not be doing the same thing?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PDBx/mmCIF Reader/Topology Reader

11 participants

@marinegor@pep8speaks@orbeckst@yuxuanzhuang@BradyAJohnston@hmacdope@ljwoods2@ianmkenney@IAlibay@richardjgowers@PardhavMaradani