Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Document Classification Application

This Python application is a document classification tool that utilizes Large Language Models (LLM) on Amazon Bedrock to classify various documents through in-context learning. The application allows users to upload PDF or image files and classify the content of these files into predefined labels or categories.

The classifier.ipynb notebook walks you through this solution. Alternatively, there is a Strealit App that for a better user experience to classify documents.

Features

  • File Upload: Users can upload PDF or image files through the Streamlit interface.
  • Document Processing: The application handles the upload and processing of PDF and image files, including converting PDF files into individual image pages (to be used with Claude3 Vision). You can select between Amazon Textract or Claude 3 Vision for document processing.
  • Text Extraction: For PDF and image files, the application uses Amazon Textract to extract the text content from the documents. The extracted text is cached in an Amazon S3 bucket for future use.
  • Document Classification: Users can provide a manifest file containing a list of possible labels or categories for the documents. The application prompts the selected Claude language model with the extracted text and the list of possible labels, and the model generates a response classifying the document content into one or more of the provided labels.
  • Model Selection: The application supports various Claude models, including claude-3-sonnet, claude-3-haiku, claude-instant-v1, claude-v2, and claude-v2:1. Users can select the desired model through the Streamlit sidebar.
  • Cost Calculation: The application calculates and displays the cost of using the selected model based on input and output token pricing defined in the pricing.json file.
  • Caching and Persistence: The application caches the extracted text and Textract results in the S3 bucket to avoid redundant processing.

To run this Streamlit App on Sagemaker Studio follow the steps below:

  • Set up the necessary AWS resources:

Configuration

The application's behavior can be customized by modifying the config.json file. Here are the available options:

  • Bucket_Name: The name of the S3 bucket used for caching documents and extracted text.
  • max-output-token: The maximum number of output tokens allowed for the AI assistant's response.
  • bedrock-region: The AWS region where the Bedrock runtime is deployed.
  • s3_path_prefix: The S3 bucket path where uploaded documents are stored (without the trailing foward slash).
  • textract_output: The S3 bucket path where the extracted content by Textract are stored (without the trailing foward slash).

If You have a sagemaker Studio Domain already set up, ignore the first item, however, item 2 is required.

  • Set Up SageMaker Studio
  • SageMaker execution role should have access to interact with Bedrock, Textract
  • Launch SageMaker Studio
  • Clone this git repo into studio
  • Open a system terminal by clicking on Amazon SageMaker Studio and then System Terminal as shown in the diagram below
  • Navigate into the cloned repository directory using the cd command and run the command pip install -r req.txt to install the needed python libraries
  • Run command python3 -m streamlit run classifier.py --server.enableXsrfProtection false --server.enableCORS false to start the Streamlit server. Do not use the links generated by the command as they won't work in studio.
  • To enter the Streamlit app, open and run the cell in the StreamlitLink.ipynb notebook. This will generate the appropiate link to enter your Streamlit app from SageMaker studio. Click on the link to enter your Streamlit app.
  • ⚠ Note: If you rerun the Streamlit server it may use a different port. Take not of the port used (port number is the last 4 digit number after the last : (colon)) and modify the port variable in the StreamlitLink.ipynb notebook to get the correct link.

To run this Streamlit App on AWS EC2 (I tested this on the Ubuntu Image)

  • Create a new ec2 instance
  • Expose TCP port range 8500-8510 on Inbound connections of the attached Security group to the ec2 instance. TCP port 8501 is needed for Streamlit to work. See image below
  • Connect to your ec2 instance
  • Run the appropiate commands to update the ec2 instance (sudo apt update and sudo apt upgrade -for Ubuntu)
  • Clone this git repo git clone [github_link]
  • Install python3 and pip if not already installed
  • EC2 instance profile role has the required permissions to access the services used by this application mentioned above.
  • Install the dependencies in the requirements.txt file by running the command sudo pip install -r req.txt
  • Run command tmux new -s mysession. Then in the new session created cd into the ChatBot dir and run python3 -m streamlit run classifier.py to start the streamlit app. This allows you to run the Streamlit application in the background and keep it running even if you disconnect from the terminal session.
  • Copy the External URL link generated and paste in a new browser tab. To stop the tmux session, in your ec2 terminal Press Ctrl+b, then d to detach. to kill the session, run tmux kill-session -t mysession

Usage

  1. Launch the Streamlit application in your web browser.

  2. In the sidebar, select the desired Claude model and specify whether you want to classify labels per page or for the entire document.

  3. Upload a manifest file containing the list of possible labels and their descriptions. Label names, a colon and the description. Each label-description pair per line saved in a .txt file. For Example:

    drivers license : This is a US drivers license W2 : This is a tax reporting form Bank Statement : This is personal bank document PayStub : This is an individual's pay info

  4. Upload the PDF or image files you want to classify.

  5. Click the "Classify" button to initiate the classification process.

  6. The application will display the classification results and the cost associated with using the selected model.

Contributing

Contributions to this project are welcome. If you find any issues or have suggestions for improvements, please open an issue or submit a pull request.

License

This project is licensed under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages