This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
This repository was archived by the owner on Sep 18, 2023. It is now read-only.

Repository files navigation

Age Restriction Analysis of Movie Scripts

The goal of our project is to automate and extend the age rating of movies based on the script of the movie. The criteria for our evaluation will be primarily based on the official rating rules used by the Motion Picture Association.

1. General information

Team Members:

Existing code fragments

Utilized libraries

  • Elastic search
  • React
  • Fastapi
  • Spacy
  • Sklearn Check requirements.txt in sub-modules of project for detailed information.

Contributions

See "Project log" section.

2. Project State

Planning State

One of our high level milestones for November could not be achivet yet. Find a detailed list of goals achived and not achived.

Achived Goals

  • Basic project setup
  • Setup elastic search hosting on linux server
  • Data crawling
  • Data understanding
  • Data analysis
  • Start of data preprocessing
  • Get baseline for PG-Ratings

Open Goals

  • Data preprocessing pipeline
  • Decision which statistical method to use for age ratings
  • First PG ratings on films

Future Planning

Currently, the project is behind schedule with respect to the initial milestone plan. This should be made up by the lecture break from 22.12.22 to 07.01.23. One reason for the delay is the change in data sourcing. The operator of a platform for film scripts had unexpectedly stopped responding.

High-level Architecture Description

High level application architecture

High level processing architecture

todo: add descriptions to preprocessing steps and knowledge from data analysis

Experiments

First experiment to find the official age ratings of movies through a TMDB API. For this we wrote a python script and used a selection of movie titles to get the age ratings. Results can be found here

Data analysis, explocation and description. Results can be found in next section.

3. Data Analysis

This section can be found in this jupyter-notebook:

Project log

Davit

  • POC dataset crawling
  • Linux server for elasticsearch instance
  • Data understanding, exploration and preprocessing
  • Labels scraping
  • Model fine tuning
  • Random Forest Implementation

Jakob

  • Research papers and materials
  • Research for baseline dataset
  • Implementation of finding age ratings by movie titles with TMDB API
  • Implementation of RNN based on paper
  • Preprocessing
  • Creation of training and test datasets

Leon

  • Setup git project (react, fastapi, elasticsearch, dockerfiles)
  • Setup git-hooks integration
  • Run elasticsearch instance on linux server
  • Frontend development
  • Fastapi development
  • SVM Classifier
  • Elasticsearch setup and upload

Project structure

Find a raw overview how the project is structured:

assets

Documentation files

data

Training data and model results

data_exploration

Data exploration notebook

data_gathering

  • Notebooks for scraping and collecting relevant data (e.g. age ratings)
  • Elasticsearch upload script

data_preprocessing

  • Notebooks to preprocess data

elasticsearch

  • Nginx config
  • Documentation
  • Elasticsearch was set up on our server. If you want to access the elasticsearch instance, please contact us.

fastapi

  • API project

own_model

  • SVM and Random Forest model implementation

react_ageflix

  • Frontend project

severity_model

  • Model from referenced paper

Presentation

Presentation can be found in the assets folder.

About

Text analytics research project for classifying music tracks into genres by analysing their lyrics.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages