Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

History

113 Commits

Repository files navigation

BioJava-Spark

Algorithms that are built around BioJava and are running on Apache Spark

Build StatusLicenseStatusVersion

Starting up

Some initial instructions can be found on the mmtf-spark project

https://github.com/sbl-sdsc/mmtf-spark

First download and untar a Hadoop sequence file of the PDB (~7 GB download)

wget http://mmtf.rcsb.org/v1.0/hadoopfiles/full.tar
tar -xvf full.tar

Or you can get a C-alpha, phosphate, ligand only version (~800 Mb download)

wget http://mmtf.rcsb.org/v1.0/hadoopfiles/reduced.tar
tar -xvf reduced.tar

Second add the biojava-spark dependecy to your pom

<dependency>
<groupId>org.biojava</groupId>
<artifactId>biojava-spark</artifactId>
<version>0.2.1</version>
</dependency>

Extra Biojava examples

Do some simple quality filtering

floatmaxResolution = 3.0f;
floatmaxRfree = 0.3f;
StructureDataRDDstructureData = newStructureDataRDD("/path/to/file")
.filterResolution(maxResolution)
.filterRfree(maxRfree);

Summarsing the elements in the PDB

Map<String, Long> elementCountMap = BiojavaSparkUtils.findAtoms(structureData).countByElement();

Finding inter-atomic contacts from the PDB

Doublemean = BiojavaSparkUtils.findContacts(structureData,
newAtomSelectObject()
.groupNameList(newString[] {"PRO","LYS"})
.elementNameList(newString[] {"C"})
.atomNameList(newString[] {"CA"}),
cutoff)
.getDistanceDistOfAtomInts("CA", "CA")
.mean();
System.out.println("\nMean PRO-LYS CA-CA distance: " + mean);

About

💥 Algorithms that are built around BioJava and run on Apache Spark

Resources

Stars

8 stars

Watchers

9 watching

Forks

Releases

Packages

Used by

Contributors

Languages