Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - StabRise/spark-pdf: PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it · GitHub
Skip to content

Repository files navigation


Spark Pdf

Open In Colab Qick StartTestMaven Central VersionLicenseCodacy Badge

Share on XShare on LinkedInShare on Reddit


⭐ Star us on GitHub — it motivates us a lot!

Source Code: https://github.com/StabRise/spark-pdf

Quick Start Jupyter Notebook Spark 3.5.x on Databricks: PdfDataSourceDatabricks.ipynb

Quick Start Jupyter Notebook Spark 3.x.x: PdfDataSource.ipynb

Quick Start Jupyter Notebook Spark 4.0.x: PdfDataSourceSpark4.ipynb

With Spark Connect: PdfDataSourceSparkConnect.ipynb


Welcome to the Spark PDF

The project provides a custom data source for the Apache Spark that allows you to read PDF files into the Spark DataFrame.

If you found useful this project, please give a star to the repository.

👉 Works on Databricks now. See the Databricks example. Solved issue with read from the volume (Unity Catalog).

Key features:

  • Read PDF documents to the Spark DataFrame
  • Support efficient read PDF files lazy per page
  • Support big files, up to 10k pages
  • Support scanned PDF files (call OCR for text recognition from the images)
  • No need to install Tesseract OCR, it's included in the package
  • 👉 Compatible with ScaleDP, an Open-Source Library for Processing Documents using AI/ML in Apache Spark.
  • Works with Spark Connect

Requirements

  • Java 8, 11, 17
  • Apache Spark 3.3.2, 3.4.1, 3.5.0, 4.0.0
  • Ghostscript 9.50 or later (only for the GhostScript reader)

Spark 4.0.0 is supported in the version 0.1.11 and later (need Java 17 and Scala 2.13).

Installation

Binary package is available in the Maven Central Repository.

  • Spark 3.5.*: com.stabrise:spark-pdf-spark35_2.12:0.1.17
  • Spark 3.4.*: com.stabrise:spark-pdf-spark34_2.12:0.1.11 (issue with publishing fresh version)
  • Spark 3.3.*: com.stabrise:spark-pdf-spark33_2.12:0.1.17
  • Spark 4.0.*: com.stabrise:spark-pdf-spark40_2.13:0.1.17

Options for the data source:

  • imageType: Oputput image type. Can be: "BINARY", "GREY", "RGB". Default: "RGB".
  • resolution: Resolution for rendering PDF page to the image. Default: "300" dpi.
  • pagePerPartition: Number pages per partition in Spark DataFrame. Default: "5".
  • reader: Supports: pdfBox - based on PdfBox java lib, gs - based on GhostScript (need installation GhostScipt to the system)
  • ocrConfig: Tesseract OCR configuration. Default: "psm=3". For more information see Tesseract OCR Params
  • password: Password for protected PDF files

Output Columns in the DataFrame:

The DataFrame contains the following columns:

  • path: path to the file
  • page_number: page number of the document
  • text: extracted text from the text layer of the PDF page
  • image: image representation of the page
  • document: the OCR-extracted text from the rendered image (calls Tesseract OCR)
  • partition_number: partition number

Output Schema:

root
|-- path: string (nullable = true)
|-- filename: string (nullable = true)
|-- page_number: integer (nullable = true)
|-- partition_number: integer (nullable = true)
|-- text: string (nullable = true)
|-- image: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- resolution: integer (nullable = true)
| |-- data: binary (nullable = true)
| |-- imageType: string (nullable = true)
| |-- exception: string (nullable = true)
| |-- height: integer (nullable = true)
| |-- width: integer (nullable = true)
|-- document: struct (nullable = true)
| |-- path: string (nullable = true)
| |-- text: string (nullable = true)
| |-- outputType: string (nullable = true)
| |-- bBoxes: array (nullable = true)
| | |-- element: struct (containsNull = true)
| | | |-- text: string (nullable = true)
| | | |-- score: float (nullable = true)
| | | |-- x: integer (nullable = true)
| | | |-- y: integer (nullable = true)
| | | |-- width: integer (nullable = true)
| | | |-- height: integer (nullable = true)
| |-- exception: string (nullable = true)

Example of usage

Scala

importorg.apache.spark.sql.SparkSessionvalspark=SparkSession.builder()
.appName("Spark PDF Example")
.master("local[*]")
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16")
.getOrCreate()
valdf= spark.read.format("pdf")
.option("imageType", "BINARY")
.option("resolution", "200")
.option("pagePerPartition", "2")
.option("reader", "pdfBox")
.option("ocrConfig", "psm=11")
.option("password", "pdf_password")
.load("path to the pdf file(s)")
df.select("path", "document").show()

Python

frompyspark.sqlimportSparkSessionspark=SparkSession.builder \
.master("local[*]") \
.appName("SparkPdf") \
.config("spark.jars.packages", "com.stabrise:spark-pdf-spark35_2.12:0.1.16") \
.getOrCreate()
df=spark.read.format("pdf") \
.option("imageType", "BINARY") \
.option("resolution", "200") \
.option("pagePerPartition", "2") \
.option("reader", "pdfBox") \
.option("ocrConfig", "psm=11") \
.option("password", "pdf_password") \
.load("path to the pdf file(s)")
df.select("path", "document").show()

Disclaimer

This project is not affiliated with, endorsed by, or connected to the Apache Software Foundation or Apache Spark.

Releases

Packages

Used by

Contributors

Languages